LLM 관련 주요 논문 - 2026-06-23

1. Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles


2. AI Exposure Scores: what they measure, what they miss, and what comes next


3. Causal Discovery in the Era of Agents


4. SPIRAL: Learning to Search and Aggregate


5. The Topology of Ill-Posed Questions: Persistent Homology for Detection and Steering in LLMs


6. POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation


7. CADRE: Stable, Parameter Efficient Adaptation of Medical Vision Language Models with Bounded Forgetting and Prior Drift


8. Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems


9. Abstract representational geometry supports inference in large language models


10. GIF: Locally Sound Geometric Information Flow Control for LLMs


11. Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation


12. IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation – the Case of the SpaceX (SPCX) IPO


13. A Stackelberg Framework for Resource-Aware LLM Agents: Learning, Repair, and Conditional Guarantees


14. When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models


15. Plans Don’t Persist: Why Context Management Is Load Bearing for LLM Agents


16. When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents


17. ThermoLLM: Thermodynamics-Aware HVAC Control with Spatial-Semantic Knowledge Graph


18. Agent-as-a-Router: Agentic Model Routing for Coding Tasks


19. CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents


20. RaMem: Contextual Reinstatement for Long-term Agentic Memory


21. MINCE: Shrinking LLM Evaluation Datasets via Few-Model Monte Carlo Calibration


22. Measuring Behavior Portability in Large Language Models


23. A Formula-Driven Survey and Research Agenda for On-Policy Distillation


24. The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models


25. GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation


26. Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards


27. Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task Risk Fusion


28. VISTA Architect: A graph database-oriented health AI system demonstrated in multidisciplinary tumor boards


29. Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations


30. AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent


31. Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs


32. SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment


33. PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement


34. Text2DSL: LLM-Based Code Generation for Domain-Specific Languages


35. Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents


36. SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments


37. VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows


38. PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs


39. Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment


40. PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems


41. MetaPS: Adaptive Programmatic Strategy Selection for Market Agents


42. ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery


43. Hypothesis-Driven Skill Optimization for LLM Agents


44. Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice


45. When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR


46. CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents


47. Holmes: Multimodal Agentic Diagnosis for Mixed-Language Mobile Crashes at Industrial Scale


48. Learning the ARTS of Search for Automated Discovery


49. ForEx: A Formal Verification Framework for Explainable Reasoning in Logical Fallacy Detection and Annotation


50. AgentCAT: Simulating Computerized Adaptive Testing via Multi-Agent Large Language Models


51. Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents


52. Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems


53. Counsel: A Meta-Evaluation Dataset for Agentic Tasks


54. Composing Verifiable Conceptual Models via Building Blocks: Towards Design-Time Verification of Agentic AI Workflows


55. AI Alignment From Social Choice Perspectives


56. AutoRAS: Learning Robust Agentic Systems with Primitive Representations


57. Don’t Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents


58. Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention


59. ARCO: Adaptive Rubric with Co-Evolution for Multi-Step LLM-Based Agents


60. Trip+: Benchmarking Agents in Personalized Interactive Travel Planning


61. Learning Burst-Aware Early Warning Models for Capacity Stress under AI Workload Surges in Hyperscale Data Centers


62. Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models


63. Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training


64. Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines


65. Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning


66. Agentic Time Machine as an Infrastructure for Future-Event Forecasting


67. How Should Agents Read Demonstrations? Hierarchical Structure Beats Flat Action Logs


68. AutoACSL: Synthesizing ACSL Specifications by Integrating LLMs with CPG-Based Static Analysis


69. Neurosymbolic Clinical Trial Matching via LLM-Driven Abduction and Logical Verification


70. When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study


71. Process-Reward Tactic Evolution for Long-Horizon Bioinformatics Workflows


72. FairTutor: Equity-Aware Pedagogical LLM Routing for Budget-Constrained AI Tutoring


73. Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes


74. From Question Answering to Task Completion: A Survey on Agent System and Harness Design


75. DrugBench: Evaluating AI Control Protocols for Medication Harm Mitigation


76. From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents


77. Skill Coverage: A Test Adequacy Metric for Agent Skills


78. SPARC: A Multi-Agent System for Electrical Circuit Question Answering


79. An LLM-Explainable DRL Framework for Passenger-Directed Autonomous Driving


80. RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents


81. Harnessing Agent Skills: Architectural Patterns and a Reference Architecture for Skill-Mediated LLM Agents


82. AlphaMemo: Structured Search-Process Memory for Self-Evolving Alpha Mining Agents


83. In LLM Reasoning, there is Irrationality on top of Value Misalignment


84. PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate


85. The New Associationism: Lessons from Deep Learning


86. Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies


87. Semantic Browsing: Controllable Diversity for Image Generation


88. AIR: Adaptive Interleaved Reasoning with Code in MLLMs


89. Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?


90. Tapered Language Models


91. Data Selection Through Iterative Self-Filtering for Vision-Language Settings


92. Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers


93. Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models


94. What Does a Chemical Language Model Know About Molecules?


95. GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation


96. Detecting Malicious Agent Skills in the Wild using Attention


97. HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models


98. Automated Semantic Fault Localization in SysML v2: A Human-in-the-Loop Framework Using Knowledge-Graph Augmented LLMs


99. Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting


100. Energy-Based Transformers as Predictors of Reading Difficulty


101. Exposing the Illusion of Erasure in Knowledge Editing for LLMs


102. P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture


103. MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations


104. When Does Intrinsic Self-Correction Help? A Task-Sensitive Analysis


105. Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory


106. LLM-Aided A* Search in Non-Geometric Network Graphs


107. PRIDE: Privileged Information-enhanced Distillation for Empathetic Dialogue Generation


108. ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation


109. Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies


110. Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs


111. From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection


112. The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel


113. EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning


114. Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning


115. StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs


116. Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently



118. From Fragments to Paths: Task-Level Context Recovery for Large Industrial Codebases


119. Priority-Aware Learning-Unlearning Correction for Dynamic Decentralized LoRA Fine-Tuning


120. VideoLatent: Video-Language Learning via Latent Self-Forcing


121. Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis


122. AI Fiction in the Wild


123. Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking


124. Libretto: Giving LLM Agents a Sense of Musical Structure


125. The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs


126. Orthogonal Representation Editing: Decoupling Semantic Entanglement in Batch Knowledge Editing of LLMs


127. Context-Aware Distillation and Ablation for Text2DSL


128. Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do


129. Training-Free Semantic Correction for Autoregressive Visual Models


130. Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing


131. An LLM-Orchestrated Agent for Directional-Coupler Design with Self-Consistent Eigenmode and FDTD Validation


132. All Green, Still Broken: Real-Flow Verification Lessons from an LLM-Integrated, Multi-Market Web Application


133. Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation


134. CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories


135. Words as Difference Makers: How Large Language Models Determine Causal Structure in Text


136. Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding


137. Reinforcement learning to improve large language model-based automated code compliance systems


138. Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset


139. First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers


140. On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning


141. BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories


142. Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model


143. Leveraging Large Language Models to Obscure Code Stylometry: A Comparative Study of GPT-3.5 and GPT-4


144. SCENIC: Semantic-Conditioned Edge-Aware Neural Framework for Structured IoT Command Generation


145. From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa


146. MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation


147. Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability


148. Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases


149. On the Expressive Power of Weight Quantization in Large Language Models


150. When Is Emergent Consensus Real? A Measured Coupling Gain and a Validity Diagnostic for LLM Agent Societies


151. L20-Edu-135M: An Auditable Single-GPU Study of Data-Efficient Small Language Modeling


152. $π$-RAG: Oblivious Retrieval via Semantic Quantization and Transcendental Addressing for Large Language Models


153. BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language


154. TraceView: Interactive Visualization of Agentic Program Repair Trajectories


155. CodeTeam: An LLM-Powered Multi-Agent Framework for Repository-Level Code Generation


156. Channel Location Constrains the Auditability of Subliminal Learning


157. Old Fictions, New Skins: Evaluating the Manipulative Capabilities of LLMs in Healthcare


158. Fine-Tuning Large Language Models for Quantum Reasoning


159. From RAN Control to Agentic Intelligence: Architecture and Vision for Energy Efficient AI-RAN


160. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning


161. Latent Confidence Alignment for LLM Self-Assessment


162. Scaling Performance and Low-Resource Annotation with Many-Shot In-Context Learning for Named Entity Recognition


163. Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead


164. Protein contacts are already in the attention: a single-forward-pass alternative to the Categorical Jacobian


165. The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference


166. Harness-MU: A Safe, Governed, and Effective Harness for Multi-User LLM Agents


167. UniRank: Unified Rank Allocation for Low-Rank LLM Compression


168. AgentDSE: Reasoning-Augmented Architectural Design Space Exploration


169. CNnotator: LLM-Guided Memory Safety Annotation Synthesis


170. Generating Public Health Responses using Survey-Augmented Large Language Models


171. CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks


172. HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning


173. Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents


174. Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reasoning


175. PrivacyAlign: Contextual Privacy Alignment for LLM Agents


176. Clinical Term Extraction using Open-Source Small Language Models


177. TACO: Task-Aware Column Description Generation Using LLMs


178. Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers


179. When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model


180. The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection


181. Decoupling the Declarative from the Procedural in Vision-Language-Action Models


182. Evaluation of Small Language Models for Arabic Language Processing


183. CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents


184. Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study


185. MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning


186. Evaluating LLMs for Real-World Web Vulnerability Detection


187. An Empirical Study of OpenPangu Quantization on Ascend NPUs


188. Recency/Frequency Adaptive KV Caching for Large Language Model Serving


189. FleetAgent: Teleoperation Assistant for Autonomous Fleets via Vectorized V2N Messages


190. When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs


191. Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders


192. Beyond Hooking Onto the World: Referential Profiles and the Numerical Structure of LLM Grounding


193. MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models


194. An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game


195. AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?


196. AdaMem: Learning What to Remember for Personalized Long-Horizon LLM Agents


197. AgentMeter: Evaluating Model-CLI Matching for CLI-Based Local Task-Solving Agents


198. LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations


199. Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer


200. CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays


201. The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence


202. Text-to-Image Generative AI for Modeling and Simulation: Methods, Opportunities, and Applications


203. Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs


204. Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning


205. Comparing Transformers and Hybrid Models at the Token Level



207. PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs


208. Latent Personal Memory: Represent personal memory as dynamic soft prompts


209. Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents


210. The Token Tax of Epistemic Accuracy: Comparing RAG and Long-Context Architectures for Document-Grounded Generative AI Applications


211. PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality


212. Can LLMs Reason About Brand Ownership? An Empirical Study of Domain Attribution Intelligence


213. Translating Inference-Time Control to Radiology Vision-Language Models: Activation Steering for Pneumonia Classification on Chest X-rays


214. Formally Verified Code Synthesis for Structured Data Translation in a Medical Internet of Things


215. Beyond ‘One Language, One Script’: Quantifying Orthographic Bias in Multilingual VLMs with PuMVR