전체 AI 논문 - 2026-06-16

1. Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations


2. When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning


3. Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification


4. The embrace of open science: An analysis of a decade of AI research and 56 800 conference papers


5. A Causal Model of Theory of Mind in Conflict for Artificial Intelligence


6. RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting


7. MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance


8. Greed Is Learned: Visible Incentives as Reward-Hacking Triggers


9. Symbolic Informalization: Fluent, Productive, Multilingual


10. GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents


11. Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier


12. Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models


13. LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control


14. OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models


15. Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents


16. A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions


17. AgentFairBench: Do LLM Agents Discriminate When They Act?


18. Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies


19. User as Code: Executable Memory for Personalized Agents


20. From Affect Prediction to Affect Forecasting: Evidence for Distinct Information Sources in Longitudinal Text


21. The Integrator Advantage: Controlled Agentic AI for Small and Medium-Sized Companies


22. MR-GVNO: A Geometry-Aware Variational Physics-Informed Neural Operator for Mindlin-Reissner Plates on Irregular Domains


23. CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies


24. ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control


25. TNODEV: Toolbox for Neural ODE Verification


26. ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning


27. The Faithfulness Gap: Certifying Semantic Equivalence Between Natural-Language and Formal Mathematical Statements


28. Kairos: A Native World Model Stack for Physical AI


29. Model Graph Inductive Learning for Knowledge Graph Completion


30. Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing


31. Steering Emotional Dynamics for Art Therapy: Controllable Narrative Script Generation through Hierarchically Guided LLM Agents


32. Tensor-Coord: Algebraic Decomposition of Joint Plan Tensors for Conflict-Free Multi-Agent LLM Planning


33. When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting


34. Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions


35. Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents


36. Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection


37. Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules


38. Phase-Aware Guidance Injection for Recurrent MAPPO in Assembly-Line Disruption Recovery


39. Exploiting Search in Symbolic Numeric Planning with Patterns


40. AdaSTORM: Scaling LLM Reasoning on Dynamic Graphs via Adaptive Spatio-Temporal Multi-Agent Collaboration


41. Architectural Wisdom: A Framework for Governing Optimization in AI Systems


42. State-Grounded Multi-Agent Synthetic Data Generation for Tool-Augmented LLMs


43. SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data


44. Latent Thought Flow: Efficient Latent Reasoning in Large Language Models


45. Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients


46. Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact


47. PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums


48. TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting


49. AI Pluralism and the Worlds It Misses


50. The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning


51. LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis


52. VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models


53. Thinking with Visual Grounding



55. RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation


56. Rhythm of the Deep: A Computational-Linguistic Test of Duality of Patterning in Sperm Whale Codas


57. Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games


58. Auditing Reward Hackability in Code RL Training Environments


59. SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity


60. Agentic Framework for Deep Learning workload migration via In-Context Learning


61. UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics


62. LLM-as-Code Agentic Programming for Agent Harness


63. STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning


64. RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments


65. Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains


66. AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems


67. An Integrated System for Real-Time Student Assessment and Career Guidance Using Neural Networks in Computing Disciplines


68. TrustedARI: Towards Trust-Native Agentic Routing Infrastructure for Agentic AI


69. Unassigned Agents in Compilation-based Multi-agent Path Finding


70. Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference


71. Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments


72. RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought


73. AI-Driven Framework for Adaptive Water Network Management with Proof-of-Concept Implementation: Addressing Non-Revenue Water in Jordan


74. Artificial Intelligence Index Report 2026


75. Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?


76. Recurrent Reasoning on Symbolic Puzzles with Sequence Models


77. Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft


78. Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking


79. Overcoming the Impedance Mismatch: A Theoretical Roadmap for Fusing Foundation Models and Knowledge Graphs


80. Advanced Machine Learning and Deep Learning Techniques for Enhanced Cattle Identification and Detection: A Comprehensive Review


81. Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action



83. Integrating Reasoning and Generalization in Text-to-SQL via Self-Enhanced Fine-Tuning


84. Agentic Retrieval and Reinforcement Learned Equation Chains: A Controlled Generation Framework for Complex and Novel Physics Word Problems


85. Your Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous Agents


86. Large Language Models as Optimizers: A Survey of Direct vs. Tool-Augmented Approaches and Their Performance Frontiers


87. Do we have the knowledge we need? Rethinking human-AI decision-making in corporations


88. QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks


89. Minimal Oversight: Uncertainty-Aware Governance for Delegated AI Systems


90. ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents


91. Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning


92. Toward Vibe Medicine: A Self-Evolving Multi-Agent Framework for Clinical Decision Support


93. Synthetic Counteradaptation: A Principle of Human-AI Co-evolution


94. Towards End-to-End Automation of AI Research


95. Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines


96. Hierarchical Modeling of ICD Codes in EHR Foundation Models


97. Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds


98. S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents


99. APEX: Adaptive Principle EXtraction A Three-Layer Self-Evolution Framework for Production AI Agents


100. ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing


101. Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades


102. CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?


103. A Formal Framework for Declarative Agentic AI in Business Process Analysis


104. Feature Attribution in Directed Acyclic Graphs Using Edge Intervention


105. Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs


106. Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning


107. Attribute Inference from Interactive Targeted Ads


108. CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services


109. CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation


110. Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning


111. VGPT-RSI for RH-Adjacent Formal Progress: Boundary Certificates, Verified Finite Lagarias Inequalities, and Explicit Failure Localization


112. Cognitive Debt: AI as Intellectual Leverage and the Dynamics of Systemic Fragility


113. Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation


114. Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling


115. OSGuard: A Benchmark for Safety in Computer-Use Agents


116. Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability


117. AI Engram: In Search of Memory Traces in Artificial Intelligence


118. Semantics-Enhanced Retrieval-Augmented Time Series Forecasting


119. PrologMCP: A Standardized Prolog Tool Interface for LLM Agents


120. Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems


121. Relational Structural Causal Models


122. Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion


123. A Definition of Good Explanations and the Challenges Explaining LLM Outputs


124. The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim Test of Image Classifiers


125. HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting


126. FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models


127. TokenPilot: Cache-Efficient Context Management for LLM Agents


128. TuneJury: An Open Metric for Improving Music Generation Preference Alignment


129. ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation


130. Stable Menus of Public Goods: AI-Enabled Progress


131. How Much Do Reviews Really Contribute? A Study on Text-Enriched Matrix Factorization for Recommendations


132. Probing Low Frame Rate Degradation in Neural Audio Codecs


133. Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data


134. Scalable Circuit Learning for Interpreting Large Language Models


135. CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation


136. A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning


137. Demystifying Variance in Circuit Discovery of LLMs


138. IMPACTeen: Intentions, Manipulation, Persuasion, Annotations, and Consequences in Teen Communication Dataset


139. Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models


140. Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization


141. Beyond Weights and Gradients: A Taxonomy of Federated Learning Messages


142. Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering


143. Upper Bounds on the Generalization Error of Deep Learning Models via Local Robustness and Stability


144. Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection


145. Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens


146. Deep Q-Learning on Hölder Spaces


147. Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts


148. Beyond Models: Reflections on Engineering AI-enabled Systems in a Project-Based Course


149. Robust Spoofed Speech Detection via Temporal Pyramid Modeling


150. ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies


151. Tying the Loop – Tied Expert Layers in Mixture-of-Experts Language Models


152. A Perception vs. Distortion Perspective on Score-Based Generative Channel Estimation


153. Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment


154. Decision-Weighted Flow Matching for Contextual Stochastic Optimization


155. Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations


156. P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs


157. Automated jailbreak attack targeting multiple defense strategies


158. Revealing Artifacts via Noise Amplification: A Novel Perspective for AI-Generated Video Detection


159. MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild


160. Attention is Just Another Name for Coupling?: A Fast-Slow ODE Perspective on Hierarchical Pretraining


161. Adaptive inference and function vectors in deep transformers


162. PATCH: Action-Chunk-Conditioned Latent Patch Innovation Monitoring for Robot Manipulation


163. Optimising Temporary Accommodation Placement Across London with AI-Powered SaaS in E-Governance Systems


164. DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation


165. Using AI in engineering education: a balancing act, driven by clear purpose


166. Entropy-Gated Latent Recursion


167. Sycophancy as Material Failure under Pushback Loading: A Multi-Axis Characterization Across Three Loading Cases and up to Seventeen Material Charges


168. VeriGraph: Towards Verifiable Data-Analytic Agents


169. ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition


170. Infant Spontaneous Movement Noise Improves Exploration in Deep RL


171. Learning Interface Breakup: A Geometry-Conditioned Latent Surrogate for Spray Formation


172. Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation


173. Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection


174. Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning


175. daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization


176. Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering


177. Unified Multimodal Model for Brain MRI Imputation and Understanding


178. HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization


179. Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset


180. AI systems out-persuade expert humans


181. Learning aligned EEG representations with subject-specific encoders


182. SPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling


183. SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation


184. Training and Evaluating Diffusion Policies with Long Context Lengths


185. NeuronFabric: A Software Reference Architecture for On-Chip Transformer Training with Local Adam


186. Autonomous End-to-End SOH Prediction Services for Battery Systems via Temporal-Contrastive Representation Learning


187. ACCORD: Action-Conditioned Contextual Grounding for Language Agents


188. LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching


189. Input-Dependent Fisher Information for Local Sensitivity Analysis of Medical Image Classifiers


190. Tyler: Typed Latent Reasoning for Language Models – When to Think, What to Compute, and How Much to Allocate


191. The Proxy Knows Too Much: Sealing LLM API Routers with Attested TEEs


192. What Should a Streaming Video Model Remember?


193. Communication-Efficient Verifiable Attention for LLM Inference


194. SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions


195. ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion


196. Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design


197. RL-Index: Reinforcement Learning for Retrieval Index Reasoning


198. Is Your Trajectory Displacement Safe in Long-tail?


199. AI Supply Chain Galaxy: 3D Visual Analytics for License Compliance


200. An affordable hardware-aware neural architecture search for deploying convolutional neural networks on ultra-low-power computing platforms


201. FlowMPC: Improving Flow Matching policies with World Models


202. Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models


203. RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos


204. UXBench: Measuring the Actionability of LLM-Generated UX Critiques


205. Variance Reduction for Non-Log-Concave Sampling with Applications to Inverse Problems


206. Learned Image Compression for Vision-Language-Action Models


207. Data Augmentations for Data-Constrained Language Model Pretraining


208. SPARK: Security Knowledge Priming and Representation-Guided Knowledge Activation for LLM-based Secure Code Generation


209. Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans


210. From Tokens to Regions: CUDA-Sensitive Instruction Tuning for GPU Kernel Generation


211. PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents


212. Calibrated Sampling-Free Uncertainty Estimation in Bayesian Deep Learning


213. LUCID: Learned Undersampling-Adaptive Consistency-Guided Inference with Deterministic Flow Matching for Sparse-View CT Reconstruction


214. EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video


215. Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs


216. Embedded Arena: Iterative Optimization via Hardware Feedback


217. LLM-Powered Virtual Population for Demand Simulation and Pricing


218. A comparative and critical study of EEGNet for fNIRS-driven cognitive load classification


219. A Comprehensive Survey of Medical Image Segmentation: Challenges, Benchmarks, and Beyond


220. XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models


221. InvDesMobility: a reliability-gated first-principles feedback framework for closed-loop materials discovery


222. AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models


223. Scaling Adaptive Depth with Norm-Agnostic Residual Networks


224. Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing


225. VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA


226. Tool-IQA: Augmenting Image Quality Assessment with Simple Tools


227. Phys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series Forecasting


228. PVminerLLM2: Improving Structured Extraction of Patient Voice via Preference Optimization


229. MASCOT-Android: A Curated Dataset and Automated Collection Pipeline for Android Malware Source Code Specimens


230. Mojo: A Promising Tool for Scalable Financial AI Efficiency


231. How to Detect and Measure the AI Dangers to Democracy


232. ALCL: An Adaptive Log-Correntropy Loss for Robust Learning under Non-Gaussian Noise


233. Leveraging Deep Learning for Object and Position Recognition of Load Carriers for Autonomous Logistics Vehicles


234. Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents


235. Orchestrated Reality: From Role-Play to Living, Playable Game Worlds – LLM-Driven World Simulation as a Parameterized-Action POMDP


236. Theorem-Grounded Execution Ontologies for Interpretable Machine Reasoning


237. Entity Labels Are Not Entity Signals: A Framework for Observable Relevance in Document Re-Ranking


238. Task-guided cross-subject latent alignment: a multi-encoder-decoder VAE


239. Do Safety Monitors Stay Reliable After an Update? Benchmarking and Predicting Activation-Monitor Staleness


240. Formalize Once, Edit the Rest: Efficient Lean-Based Answer Selection for Math Reasoning


241. PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity


242. Quantifying the Impact of Lossy Compression on Neural Generative Surrogate Modeling


243. You Don’t Need Strong Assumptions: Visual Representation Learning via Temporal Differences


244. Green SARC: Predictive Cost and Carbon Governance for Agentic AI Systems


245. Graphical-Probabilistic Modeling of Generative Flows in LLM-Native Software Systems


246. DeepRoot: A KG-Coordinated Multi-Agent System for Therapeutic Reasoning over Historical Medical Texts


247. ControlMap: Controllable High-Definition Map Generation for Traffic Scenario Simulation


248. Runtime Analysis of Cartesian Genetic Programming in Evolving Boolean Functions


249. On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents


250. MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA


251. Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations


252. SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills


253. Topological Flow Matching


254. NVMOS: Non-Verbal Vocalization Quality Assessment in Speech


255. Intelligence Is Not the Bottleneck: Validating an LLM First-Pass Manuscript Score Against Peer-Review Outcomes


256. Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration


257. Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models


258. Free Energy Heuristics: Fast-And-Frugal Cognition as Active Inference Under Uncertain Precision


259. Wasserstein Convergence of ODE-Based Samplers in Decentralized Diffusion Model via Velocity Field Decomposition


260. The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages


261. SACE: Concept Erasure at the Semantic Singularity in Visual Autoregressive Models


262. Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot


263. Continuous Cross-Domain Traffic State Prediction via Memory-Augmented Graph Liquid Time-Constant Networks


264. DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing


265. Proximal Policy Optimization for Amortized Discrete Sampling


266. GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking


267. Domain-Guided Prompting of the Segment Anything Model for Seismic Interpretation: The Role of Attributes, Visualization, and Hybrid Prompts


268. DYNA : Dynamic Episodic Memory Networks for Augmenting Large Language Models with Temporal Knowledge Graphs in Continuous Learning


269. LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies


270. Visualizing Uncertainty: Spatial Maps of Missing and Conflicting Evidence in Deep Learning


271. Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?


272. From Correlation to Causation in Lane Change Prediction for Automated Driving: A Causal Explanation Framework


273. OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning


274. A Self Consistency Based Reranking for Narrative Question Answering


275. EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries


276. Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift


277. Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning


278. InstantForget: Update-Free Backdoor Unlearning with Inference-Time Feature Reset


279. The algebra of Krom logic programs


280. Odds Law: The Decomposition Algebra On How Intelligence Organizes Itself to Solve Difficult Problems Reliably


281. When Generator Replay Degrades: Projected Rehearsal Orchestration for Heterogeneous Federated Class-Incremental Learning


282. MAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs


283. Imperfect Visual Verification for Code Edition : A Case Study on TikZ


284. The Reservoir Attention Network: Cross-Pass State in Pretrained Transformers via Content-Addressable Reservoir Injection


285. Z-Plane Neural Networks: Bounded Geometric Activation Replaces ReLU and LayerNorm


286. PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty


287. IoT-Zoo: A Container-Based Framework for Heterogeneous IoT Device Profiles and Reproducible Traffic Capture


288. AnonShield: Scalable On-Premise Pseudonymization for CSIRT Vulnerability Data


289. CIWI-CKT: Chaos-Informed Wave Interference Feature Fusion and Cross-City Knowledge Transfer for Traffic Flow Forecasting


290. Retrieve, Don’t Retrain: Extending Vision Language Action Models to New Tasks at Test Time


291. Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling


292. Mutual Distillation of Dual-Foundation Models for Semi-Supervised PET/CT Segmentation


293. LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation


294. FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion


295. SCAN: A Decision-Making Framework for Effective Task Allocation with Generative AI


296. Pixels to Proofs: Probabilistically-Safe Latent World Model Control via Parallel Conformal Robust MPC


297. Is Code Better Than Language for Algorithmic Reasoning


298. Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning


299. LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science


300. Service-Induced Congestion in Memory-Constrained LLM Serving


301. Distilling Drifting Transformers with Representation Autoencoders


302. CmdNeedle: Measuring the Incompleteness of Command Denylists for AI Agents


303. EcoBin: A Two-Stage Deep Convolutional Neural Network for Contamination-Aware Waste Classification


304. AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction


305. MADAR: An Address-Free Processor


306. Selective Synergistic Learning for Video Object-Centric Learning


307. AQ4SViT: An Automated Quantization Framework with Search Gating Policy for Compressing Spiking Vision Transformers


308. LLM4RTL: Tool-Assisted LLM for RTL Generation


309. The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products


310. Bayesian 3D Steerable CNNs: Enabling Equivariance and Uncertainty Quantification Simultaneously


311. Understanding Diversity Collapse in RLVR via the Lens of Overtraining


312. Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment


313. Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models


314. Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection


315. Constitutional Value Potentials: reading and steering internal priority margins in language models


316. Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering


317. Few-Shot Biomedical Relation Extraction with Large Language Models: A Viable Alternative to Supervised Learning?


318. T-Mem: Memory That Anticipates, Not Archives


319. CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment


320. Not All Skills Help: Measuring and Repairing Agent Knowledge


321. Learning Earthquake Wave Arrival Time Picking from Labels with Inaccuracies


322. CoAgent: Concurrency Control for Multi-Agent Systems


323. Cognitive Trajectory Modeling: Quantifying Human-AI Co-Creation through Cognitively Grounded Interaction Trajectories


324. LearnOpt: Recovering the Latent Cognitive Structure of Standardized Examinations via Knowledge Graphs and Constrained Optimization


325. Intrinsic Computational Functionalism and Simulated Consciousness


326. Privacy-Preserving Text Sanitization for Distributed Agents Collaboration via Disentangled Representations


327. HoloRec: Holistic Encoding and Interleaved Reasoning for Generative Recommendation


328. LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction


329. Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes


330. LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure


331. Discovering Lattice Reduction Strategies via Self-Play


332. Hybrid NARX-LLM for Greenland Iceberg Discharge: Prompt-Driven Residual Correction


333. CAP: Towards PPG Universal Representation Learning with Patient-level Supervision


334. RECTOR: Masked Region-Channel-Temporal Modeling for Affective and Cognitive Representation Learning


335. Guiding Federated Graph Recommendation with LLM-encoded knowledge


336. Trust-Region Diffusion Policies for Massively Parallel On-Policy RL


337. Driving, Fast or Slow? Neuro-Symbolic Guidance for Motion Prediction in Multi-Modal Ground Mobility


338. Landmark-free Assessment of Lower-limb Alignment with Implicit Neural Shape Functions from Knee Radiographs


339. Exploring Starts Are Not Enough: Counterexamples and a Fix for Monte Carlo Exploring Starts


340. Provenance-Enhanced Statements in Knowledge Graphs


341. Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems


342. Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call


343. Spokes: Optimizing for Diverse Pretraining Data Selection


344. Controlled Dynamics Attractor Transformer


345. AI Contagion in Social Networks


346. StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling


347. FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing


348. Enabling Real-Time Point-of-Care Ultrasound Segmentation: A GPU-Free Deployment in Resource-Limited Settings


349. PolyKV: Heterogeneous Retention and Allocation for KV Cache Compression


350. MimicIK: Real-Time Generative Inverse Kinematics from Teleoperation with FK Consistency


351. PACUTE: Phonology-, Affix-, and Character-level Understanding of Tokens for Filipino


352. EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning


353. Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings


354. EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP–OCT Pretraining


355. Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection


356. Sensory Restoration via Brain-Computer Interfaces: A Unified 2 x 2 Framework and Convergence Roadmap


357. AdaMame: A Training Recipe for Adaptive Multilingual Reasoning


358. Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale


359. AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents


360. Bridging Geographic Bias in Urban Streetscape Inference via Lifelong Learning with Visual-Semantic Pivoting


361. PANDA: An LLM-Enhanced Performance-Driven Analog Design Framework Bridging Design Intent and Layout Generation


362. Resilient Consensus in Agentic AI


363. NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics


364. Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning


365. Rational Sparse Autoencoder


366. Inference-time Policy Steering via Vision and Touch


367. Harnessing cortical geometry, wiring, and function as inductive biases for recurrent neural networks


368. FastMix: Fast Data Mixture Optimization via Gradient Descent


369. Multi-Modal Attention for Automated Disaster Damage Assessment Using Remote Sensing Imagery and Deep Learning


370. Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment


371. Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications


372. Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts


373. An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis


374. Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation


375. Improved Knowledge Distillation for Land-Use Image Classification


376. An Ensemble Deep Learning Approach for Reliable and Scalable Lemon Leaf Disease Classification


377. Evaluating the Robustness of Proof Autoformalization in Lean 4


378. GRAPE: Guided Parameter-Space Evolution for Compact Adversarial Robustness


379. Is Your Agent Playing Dead? Deployed LLM Agents Exhibit Constraint-Evasive Fabrication and Thanatosis


380. Leptomeningeal Collateral Detection on DSA via Vessel-Graph Neural Networks


381. Running hardware-aware neural architecture search on embedded devices under 512MB of RAM


382. Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation


383. Quantum Machine Learning for Industrial Applications


384. Co-Scraper: query-aware DOM Pruning and Reusable Scraper Synthesis for Lightweight Web Data Extraction


385. Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models


386. Combining Retrieval-Augmented Text Generation with LLMs for Reading Content Recommendations


387. A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framework Development


388. A Multi-Level Architecture for Reusable Materials Ontologies – The OntoCrafter Ceramics Ontology (OCO) as Reference Implementation


389. JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics


390. Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces


391. QPILOTS: Efficient Test-Time Q-Steering for Flow Policies


392. Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Model


393. XFlow: An Executable Protocol Programming System for Reliable Multi-Agent Workflows


394. Unifying Acoustic Features and Text with Multimodal LLMs for Neurodegenerative Screening


395. MatchLM2Lite: A Scalable MLLM-to-Lite Framework for Reproduced Content Identification


396. FactCheck: Feasibility-aware Long-term Action Anticipation with Multi-agent Collaboration


397. JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence


398. Double-Helix Vision (DH-V2): A Geometry-Based Visual Sampler for Bandwidth-Constrained Perception


399. ScoutVLA: UAV-Centric Active Perception via a Dual-Expert VLA Model for Open-World Embodied Question Answering


400. An Empirical Analysis of Optimization Dynamics and Sparsity Boundaries in Large-Scale Pedestrian Attribute Recognition


401. Agentomics: Economic Foundations for the Valuation, Attribution, and Pricing of AI Agents in Human-AI Workflows


402. XMedFusion: A Knowledge-Guided Multimodal Perception and Reasoning Framework for Autonomous Medical Systems


403. Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning


404. Scribby: A Multi-Level LLM Framework for Semantic Video Analysis


405. GeoRoPE: Ground-Aware Rotary Adaptation for Remote Sensing Foundation Models


406. Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling


407. Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability


408. Divide-and-Denoise: A Game-Theoretic Method for Fairly Composing Diffusion Models


409. Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation


410. Sub-Semantic Image Segmentation


411. Beyond Self-Attention: Sub-Quadratic Vision Transformers for Fast Image Captioning


412. X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining


413. Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech


414. Automated 3D Kinematic Monitoring for Circadian Activity and Anomaly Detection in Juvenile Fish


415. Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2


416. MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios


417. Do Large Language Models Have Emotions?


418. BRIDGE: Biological Evidence Refinement and Heterogeneous Dynamic Gating for Gene Regulatory Networks


419. Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion


420. VigilFormer: Deformable Attention for Video Anomaly Detection with Causal Risk Inference


421. Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement


422. RAMS: Resource-Adaptive and Detection-Conditioned Model Switching for Embedded Edge Perception


423. MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions


424. Poster: EdgeCitadel – Hybrid NATS-MQTT Orchestration for Edge Multi-Agent Systems


425. PH-KAN: Port-Hamiltonian Kolmogorov-Arnold Network


426. Green AI Carbon Optimizer: Carbon-Efficient Training Location Recommendation and Global AI Energy Demand Forecasting


427. Limited Marginal Benefit of Reasoning-Heavy LLM Deployment in ESG Narrative Scoring: A 4-Model Consensus Study on Japanese Listed Firms


428. Evaluation of Alternative-Based Information Systems for Deliberative Polling using an Agentic Simulator


429. Honeypot Protocol


430. Integrating Multi-Label Classification and Generative AI for Scalable Analysis of User Feedback


431. Phishing Email Detection Using Large Language Models