LLM 관련 주요 논문 - 2026-09-02

1. Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers


2. Can LLMs Discover Scientific Laws in Real and Parallel Worlds?


3. EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation


4. When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation


5. Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement


6. Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers


7. EdiTikZ: Scientific Figure Editing from Revision Trajectories


8. EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems


9. SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding


10. LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting


11. Automated Event Log Generation from Unstructured Text Using Finetuned LLMs


12. A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation


13. Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations


14. Prompt-Robust Language Models: Which Training Strategies Work?


15. H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning


16. Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs


17. WorldBench: Culturally Grounded Benchmark for Multilingual Agents


18. AgentFactory: Towards Automated Agentic System Design and Optimization


19. Data-Driven Persona-Conditioned Agents for A/B Test Simulation


20. Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees


21. CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins


22. VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don’t Mean Preferences


23. RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation


24. In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?


25. CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training


26. Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems


27. Towards Generalizable Visually Grounded Exploration of Household Devices


28. AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation


29. One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning


30. Towards a Reliable and Practical Eval Pipeline


31. Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks


32. S^3martCirc: Self-supervised Smart Circuit Discovery


33. ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents


34. Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs


35. Agentic Empirical Asset Pricing: Methodological Foundations


36. SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification


37. ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything


38. Value Over Language Model: Detecting Original Contribution in Writing


39. Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs


40. Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets


41. SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task



43. Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs


44. Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts


45. Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random


46. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs


47. VoiceLongMemEval: Do Assistants Remember How You Sounded?


48. ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation


49. Validity-Aware Jailbreak Evaluation for Large Language Models


50. The Privacy-Hallucination Tradeoff in Differentially Private Language Models


51. EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models


52. Towards a Belief-Based World Model for LLM Agents


53. Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations


54. SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents


55. RestoreBench: Can AI Agents Restore Power Flow Convergence?


56. Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models


57. Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective


58. The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems


59. Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching


60. The Answer Is Not the Argument


61. Hypotheses-Guided Self Distillation for Continual Personalization


62. Authority Bias in Conversational Search Engines for Academic Paper Recommendation


63. Invalidation Contracts for Cross-Episode Agent Memory


64. Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems


65. LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark


66. Asymmetries in Spontaneous and Instructed Deception


67. MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts


68. SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces


69. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets


70. Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls


71. Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models


72. CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?


73. The Rise of Verbal Reinforcement Learning


74. Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs


75. From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification


76. Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories


77. Can LLMs Design Video Coding Tools? A Case Study on Planar Mode


78. LatentPress: Context Compression Beyond Text and Vision


79. GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions


80. When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning



82. Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching


83. Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity


84. CHARM: Character Hallucination for Multicultural Role Play Benchmark


85. Probing Factual Knowledge Transfer with Training Data Interventions


86. Bandits in Prod: Hyperparameter Optimization at Inference Time


87. MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval


88. Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models


89. The Constitutional Coverage Trilemma in AI Governance


90. Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents


91. From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs


92. REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs


93. Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges


94. Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion


95. EDRAC: Benchmarking Arabic Dialect Reading Comprehension


96. StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions


97. SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models


98. Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close


99. Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages


100. Disclosure-Gated User Simulation for Companion-Agent Evaluation


101. Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling


102. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding


103. Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches


104. Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO


105. Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation


106. Replacing Training with Memory: Listwise Selection for Text-to-SQL


107. Visual Attention Faithfulness in Vision-Language Models is Heterogeneous


108. Agentic programs: an emerging form of scientific software in computational materials science


109. Instella-MoE Technical Report


110. Solaris: Towards Interfaces That Are Generated, Not Coded


111. Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures


112. Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models


113. Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity


114. Restrict, Don’t Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation



116. Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning


117. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems


118. Predicting Program Exit Code with LLMs and Programming Language Semantics


119. WiseSpec: Requirements-Driven Agents for Code Generation


120. EM^2Mem: Event-Centric Multimodal Memory for Large Language Models


121. The Safeguard Worked. Is the LLM System Safer?


122. The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space


123. RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces


124. Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models


125. EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities


126. Exploring Collaboration between a language and a non-language agent


127. HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference


128. Capability-Gated Language Models: Security Composes, Utility Does Not


129. (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement


130. FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos


131. Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning


132. Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning


133. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts


134. Workload Identification with Physical Side Channels for AI Governance


135. Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer’s Disease Detection


136. Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems


137. Don’t Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes


138. QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization


139. Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation


140. Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models


141. Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs


142. Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy


143. Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding


144. Commit-first LLM judging inherits the judge’s own errors


145. Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation


146. AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis


147. Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy


148. Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning


149. Medical Causal Hypothesis Verification with Large Language Models


150. RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving


151. CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language


152. ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation


153. Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment


154. From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling


155. REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent


156. trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories


157. Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning


158. InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information