전체 AI 논문 - 2026-09-17

1. Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments


2. Flag Game: A Toy Model for Mechanistic Swarm Interpretability


3. MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education


4. Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization


5. Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning


6. Function Lives Where Variance Doesn’t: Task-Weighted Charts of a Language Model’s Computation


7. Suppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing


8. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data


9. Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows


10. CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents


11. Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale


12. Clueing up LLMs with Tool-Augmented Deductive Reasoning


13. Which LLM is Best for Translating Natural Language Goals to PDDL


14. Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning



16. Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection


17. Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making


18. TRIPROBE: Probing Task Separability Beyond Classification for XAI


19. AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution


20. Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models


21. Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs


22. First Token Matters: Understanding Safety Collapse in Large Reasoning Models


23. Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning


24. Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery


25. The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models


26. Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving


27. WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories


28. HPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition


29. Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland


30. Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts


31. Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents


32. Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition


33. Visual Compliance via Executable Safety Rule Entailment


34. What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models


35. Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum


36. Building Trust in Artificial Intelligence: A Necessity for Railway Applications


37. Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI


38. BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs


39. REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement


40. Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment


41. WFM: Wiki Foundation Model for Complex Agentic Reasoning


42. Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting


43. Symbolic Temporal Supervision of LLM Agents Using Contracts


44. Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost


45. AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines


46. When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation


47. Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition


48. Teaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools


49. The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction


50. Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning


51. Missing Bridges: Composition-Aware Active Imitation Learning


52. Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation


53. RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents


54. TuiML: Machine Learning for AI Agents


55. Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits


56. When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI


57. Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI


58. Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations


59. Collaborative Memory for Multi-Agent VLM Systems


60. OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning


61. ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software


62. Do Frontier Models Seek Safety Evidence Before Acting?


63. The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?


64. SNOMED CT Concept Recommendation from Masked Clinical Context


65. Learning Heterogeneous Preferences


66. A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning


67. FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment


68. SAGE: Governed Artifact Generation from Enterprise Guidelines


69. Imitation Learning for Autonomous Driving in CARLA


70. A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products


71. NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation


72. GVD: Governed Versioning and Deduplication for Document Repositories


73. GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents


74. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video


75. What You Can’t See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization


76. Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees


77. One Color Preprocessing Improves DSATUR


78. EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents


79. Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records


80. Objective vs. Search: Decomposing What Makes a Good Tokeniser


81. A Zeroth-Order Paradigm for LLM Preference Alignment


82. Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation


83. Affora: A Design System for Agent-Friendly Interfaces


84. rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference


85. Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria


86. Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation


87. Securing quantum error correction against misleading advice from AI agents


88. Probabilistic Linear Explanations


89. Double descent is the principle of least action


90. RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control


91. MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents


92. TalkMatrix: Generating Character Dialogue that is Both Consistent and Diverse


93. Tabular Deep Learning vs Classical Machine Learning for Urban Land Cover Classification


94. WordPolo: Evaluating Language Models Through Iterative Semantic Feedback


95. BadQubits: An LLM-Based Framework for Static Pre-Execution Detection of Structurally Harmful Quantum Circuits


96. Transcribe, Then Reason: Two-Pass Decomposition for Multimodal Review


97. Automated Dental Caries Segmentation in Panoramic Radiographs Using Dual-Stage Deep Learning


98. StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction


99. Dose-Aware Cold Diffusion with Physics Consistency for Generalizable Low-Dose CT Reconstruction


100. Social Laws for Multi-agent Coordination in Stochastic Environments


101. Higher-order pruning of experts in mixture-of-experts language models


102. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking


103. NeuroECG: ECGFounder-Based Deep ECG Representation for EEG-Free Neurological Prognostication After Cardiac Arrest


104. ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions


105. Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection


106. Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN


107. GrainSpeech: Less Context, More Detail for Compact Speech Synthesis


108. Ask the Tool, Don’t Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It


109. Using OCR Heads to Verbalize Image Semantics


110. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks


111. A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages


112. Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening


113. Echo: Learning-based Matching Decompilation using Trusted Back Translation


114. Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging


115. Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification


116. CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning


117. GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media


118. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?


119. Online Robust Reinforcement Learning Through Monte-Carlo Planning


120. Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents


121. Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces


122. On-the-Fly Homographies Calibration for Multi-Camera Tracking


123. Interpretable Patch-Based Deep Learning for Wildfire Spread Prediction from Ensemble Simulations


124. VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval


125. MiST: Mid-Training LLMs for Cybersecurity


126. ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models


127. CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models


128. Multitask Reinforcement Learning for Assisting Choice Model Specification


129. TERN: A Delta-rule Memory with a Seasonal Reference and Online Adaptation for Epidemic Forecasting


130. A Non-Linear Neuron Based Detection of Isolated Pixels in Binary and Grayscale Images using Contrast Sensitive Receptive Fields


131. Reliable Virtual Sensing: A Multi-Domain Benchmark for Robustness Under Sensor Failures


132. GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models


133. Semantic CSI Feedback for Beam Selection: When Task-Aware Embeddings from Sparse Pilots Outperform Full-Bandwidth Reconstruction


134. Autonomy in Check: Governor-Mediated Adaptive Security at the Edge


135. Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR


136. Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers



138. A Study of the Reliability of Agentic AI-Generated Programs


139. I code or AI code: A comparative evaluation of AI-rated scores in classroom observations


140. ${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models


141. Quanta: A Self-Contained Python Library for Hybrid Retrieval over Quantised Embeddings, Lexical Indexes, and Knowledge Graphs


142. Remembering Solomon Marcus


143. APGEM: Adaptive Policy-Guided Error Mitigation for Quantum Reinforcement Learning on a Real-World CVRP Case Study


144. CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling


145. A Lightweight CNN Integrated Compact Convolutional Transformer for Multi-Scale Feature Learning and reducing computational complexity for breast cancer mammography image detection and classification


146. CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026


147. Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers


148. MoRE: Mixture of Reused Experts


149. DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning


150. Rethinking How We Evaluate Methodological Progress in Health AI


151. PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs


152. A Comprehensive Review of Generative Physical Artificial Intelligence


153. Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment


154. Agora: Git as Shared Memory for Collective AutoResearch


155. Mask 2D-3D: Adaptive Dual-Masked Autoencoder Network for Image-to-Point Cloud Registration


156. From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale


157. An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks


158. Physics-Informed Neural Networks for Fast Multilayer Spectral Inversion of Hα 6562.8 A and Ca II 8542.1 A Spectra


159. Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations


160. The Attention Within: Consensus Dynamics in Selective State Space Models


161. Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders


162. EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing


163. Walking the Score Manifold: Continuous-time Generative Dynamics on Learned Data Manifolds


164. Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming


165. Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels


166. AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content


167. PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research


168. RoboVAD: A Large Cross-Domain Evaluation Benchmark for Anomaly Detection in Robotic Arm Manipulation Videos


169. Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents


170. Learning Nuclear Structure with AI: Radii and Collectivity


171. Adaptive hybrid coupling with operator inference, the overlapping Schwarz alternating method and reinforcement learning


172. Procedural Pretraining for Molecular Property Prediction


173. Learning Multi-Humanoid Pickup and Transport via Decentralized Object-Centric Control


174. Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks


175. The Free Inference Dimension: Complexity Measure for Zero-Collision Navigation under Hypothesis Mixtures


176. Principled Koopman Representations with Kalman Inference for Efficient Time-Series Prediction


177. QiT: Quantum-Inspired Transformer for Visual Recognition Task


178. SAiFE-gym: Model-based Environments for Automated Market Making with Concentrated Liquidity


179. AI and Human Approaches to Mathematical Problem Solving


180. Information Set Emulation: Causal Certificates for AI Derived EHR Features


181. When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments


182. HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models


183. Is Luke the Author of a Gospel and the Acts of the Apostles?


184. CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors


185. Evolution of US Oral Political Language


186. REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff


187. One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG


188. Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents


189. Accelerating Diffusion Sampling via Speculative Draft Trees


190. The Missing “I Don’t Know”: Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention


191. Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents


192. Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models


193. Scaling Articulated Rationales for MLLM-based Recommendation


194. Lecture notes on Physics Informed Neural Networks, Neural Operators, and their applications


195. Decentralized Optimal Equilibrium Learning Over Dynamic Networks


196. Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models


197. Evolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory


198. Where Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch


199. Independence-System Realisations in Single-Source Unsplittable Flow


200. BLADE: ReliaBle Dynamic Hardware-Aware SNN-ANN Boundary SeLection for Event-BAseD Object DEtection


201. Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds


202. WARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI


203. REQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration


204. The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models