LLM 관련 주요 논문 - 2026-09-29

1. Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models


2. Reinforcing Agentic Creativity in Scientific Ideation with Night Science


3. Report: Progressive Disclosure of Agent Skills


4. Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control


5. PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents


6. Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization


7. Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts


8. TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science


9. Signatures of semantic search in the activations of large language models


10. IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing


11. Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents


12. Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning


13. From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining


14. RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis


15. Continuous Context Management


16. SRHarness: A Harness for Agentic Symbolic Regression


17. Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning


18. Self-Adapting Group of Experts for Multi-Agent Reasoning


19. Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits


20. TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving


21. Narrowing the Horizon: Quantifying Topic Saliency Shifts in Generative Monoculture


22. Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework


23. Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale


24. Imprint Reader: From Weight-Update Readout to Behavioral Intervention


25. Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses


26. EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents


27. Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference


28. Tool Mediation Alters Refusal Mechanisms in Large Language Models


29. Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling


30. Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research


31. When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents


32. What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages


33. Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation


34. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?


35. TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation


36. From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing


37. Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces


38. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias


39. ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization


40. VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection


41. PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents


42. Fewer Assumptions by Design: A Reusable Skill for LLM-Assisted Verus Verification


43. One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair


44. On the Limits of Metacognitive Monitoring in LLMs


45. Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models


46. UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning



48. Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs


49. When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety


50. SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing


51. RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents


52. FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management


53. Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents


54. OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming


55. Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses


56. MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics


57. LLMs for Executable Multi-Agent System Specification Generation


58. TULIP: Targeted LLM Unlearning at Layers Identified Per-Input


59. SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences


60. Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence


61. PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations


62. FlowState: Execution State as Memory for Long-Horizon LLM Agents


63. SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning


64. Remember Before You’re Asked: MemDream for Self-Probing Memory Evolution


65. APOLO: Automatic Prompt Optimization for Ontology Learning


66. The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions


67. Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets


68. PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems


69. When Does Structured Knowledge Help Neural Theorem Proving?


70. CORTEX: Learning to Share and Specialize in Dense Language Models


71. OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue


72. SkillFocus: Evolving Agent Skills via Capability Decomposition


73. Org-Agent: Beyond Personal Assistants Towards Organizational Agents


74. PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents


75. Improving Large Language Models for Code through Runtime Program-State Reasoning


76. SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing


77. Test-Time Scaling via Budgeted Multi-Attribute Verification


78. ControlScope: Workflow Revision and Reliability in LLM Agents


79. BIABench: Evaluating AI agents on real-world bioimage analysis tasks


80. QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models


81. Evolving Support Priorities in Empathetic Reinforcement Learning


82. AdaGuard: An Adaptive Guard Model with User-defined Policies


83. When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model


84. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures


85. Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning


86. PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models


87. ReplayLens: Auditing Agents’ Use of Outcomes


88. TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora


89. Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms


90. StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents


91. From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents


92. K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings


93. GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation


94. PhysFieldBench: Can Multimodal Models Understand Physical Fields?


95. Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization


96. Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation


97. Jev in Medicine: A Benchmark Evaluation. Preliminary Results


98. EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events


99. Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents


100. HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents



102. Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges


103. When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents


104. R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution


105. How code helps different tasks? A decompositional lens on LLM post-training


106. Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory


107. Dual-Vocabulary Language Model for Cross-Tokenizer Distillation


108. Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong


109. HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration


110. BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation


111. Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse


112. SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications


113. One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion


114. Auditing Agent Actions through Query-Conditioned Attribution


115. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents


116. Trajectory Unlearning on LLM-based Agents


117. ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments


118. JustQuant: You Don’t Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization


119. Reasoning on the Simplex: Geometric Fixed-Point Models


120. PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment


121. LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs


122. When Does the Concept of “Dog” Emerge in an Audio LLM?


123. Raven: The Harness of Harnesses for Composable Agentic Intelligence


124. MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses


125. COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement


126. CoViST: Visual Token Compression via Composable States


127. Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models


128. Agentic Multi-Turn Reasoning: A Fairness Approach


129. The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization


130. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces


131. Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets


132. Feedback Makes Perfect: A Closed-Loop Framework for NL-to-STL Translation


133. LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models


134. ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents


135. CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents


136. Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries


137. Unlocking Latent Personalization in LLMs


138. SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought


139. Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning


140. LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks


141. On Device Agentic Operation Caches – Classifier-Centric NL-to-Action Generation


142. Modular Discovery of General Game-Playing Algorithms with Large Language Models


143. LLM sequential decision making under uncertainty in biochemical domains


144. Large Language Models Substantially Compress Well-Being Inequality but Largely Preserve Its Socioeconomic Structure


145. Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States


146. The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents


147. X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization


148. The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining


149. Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence


150. Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows


151. Logical subspace in LLMs


152. Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models


153. StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills


154. FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors


155. Improving LLM Collaboration via Multi-Agent Preference Learning


156. The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces


157. Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs


158. When Can First-Order Models of Fine-Tuning Bound Forgetting?


159. Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support


160. Overwhelmed by Choice: Studying LLM Decision Making at Scale


161. Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs


162. Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret


163. Nutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutrition


164. AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks


165. $T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training


166. Forecasting Intraday USD/CAD Exchange Rate with News-Derived Monetary-Policy Signals


167. Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models


168. Adaptive Consistency Graph for Long-Horizon Agents


169. SkillVine: Agent Skill Evolution via Branching Exploration


170. CAIRN: Dynamic Fact-Intent DAGs for Multi-Agent Exploration


171. World Agent: Can Language Models Keep a World Running?


172. PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks


173. What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation


174. Contract Memory Compiler: Resolve, Then Traverse


175. From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving


176. Business Compromise Detection with Agentic AI and LLM-driven Knowledge Discovery


177. Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study


178. “You’re Right, Let Me Fix It”: How LLM Agents Damage Correct Work When Falsely Accused


179. EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval


180. ProTTT: Learning to Learn Semantic User Memory with Test-Time Training


181. Artificial intelligences and human scientists exhibit complementary strengths in theory building


182. Porimon: An LLM-Based Pokémon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented Generation


183. LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses


184. Fail Loudly: An Auditable Runtime for Agentic Data Analysis


185. MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents


186. From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis


187. Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents


188. DAAF: From Failure Localization to Editable System Assets in LLM Agents


189. Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs


190. RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems


191. When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models


192. Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling


193. VPEvolve: A Self-Evolving Virtual Process Engineer for Computational Lithography


194. ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation


195. From Latents to Wires: Surgical Post-Editing on Large Language Models


196. Authorization Closure Graph: Minimal Repair for LLM Agents with Evolving User Instructions


197. PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins


198. Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict


199. Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization


200. AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority


201. ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs


202. Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context


203. HyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop Reasoning


204. Delayed Supervision for Test-Time Language Models


205. GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents


206. Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs


207. LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound


208. Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents


209. RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation


210. A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases


211. Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments


212. Noisy Test-Time Reinforcement Learning for Code LLMs


213. PastForward: Faster On-Device GUI Agents via Computational Experience Reuse


214. Residual Streams Read, Recurrent States Remember: The Global Workspace in Mamba Models


215. Toward Interactive Understanding of Code APIs


216. EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory


217. A Benchmark for LLM’s Understanding of Middle School and High School Science Topics


218. SenseAgent: An LLM Agent for Adaptive Cross-Domain IMU Sensing


219. CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing


220. Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination


221. BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering


222. Improving Medical Calculation of LLMs with Embedded Coding


223. EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks


224. Choir: An Open Protocol for Distributed Multi-Agent Autoformalization


225. COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents


226. IndustryLLM: Failure-Driven LLM Training for Industrial Procurement


227. LLM Judge Validation Under Sparse Overlap: From Inference to Design


228. CP-Agent: A Harness-Engineered Agent for Crystal Plasticity Simulation Workflows


229. Witeness Overlap: Directional Provenance Inside Open-Weight Model Families


230. Telescopic Language Models


231. TokenCast: Forecasting Token Consumption During LLM Agent Execution


232. KV-streams for Efficient Compaction in Agentic Reinforcement Learning


233. Distillation Defenses Easily Break After Reinforcement Learning


234. Behavioral Foundation Models for Quality Diversity


235. Twist, Don’t Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models


236. SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents


237. QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks


238. FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models


239. Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability


240. Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery


241. Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs


242. Frontier Learning: Training LLM Reasoners at the Edge of Capability


243. Semantic Prefix Oracles for LLM Decoding: Contracts and Differential Validation


244. AwarenessBench: Assessing Cognitive Capabilities of Language Models


245. “Nothing to See Here’’: Unintended Disclosure through Revision Traces of LLM Deliverables