LLM 관련 주요 논문 - 2026-08-04

1. AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies


2. Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions


3. Real-Time Detection and Repair of LLM Agent Failures


4. ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision


5. Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks


6. MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models


7. Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training


8. SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents


9. MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4


10. Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories


11. PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs


12. From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents


13. Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints


14. Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning


15. Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering


16. MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents


17. Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation


18. HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning


19. EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers


20. Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG


21. Evolving in the Agent Jungle via History-Informed Opponent Awareness


22. Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study


23. CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents


24. ReasonCast: Towards Explainable Time Series Forecasting with Reasoning


25. EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning


26. Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models


27. FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling


28. PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning


29. Rewriting or Reweighting? A Geometric Account in Language Models


30. SearchMaster: Grounded and Regulated Self-Play for Search Agents


31. CockpitHAT: Dependency-Graph-Driven Hierarchical Attribution for Embodied Multi-Agent Cockpits


32. Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction


33. REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models


34. FRAMES: Guarded and Dual-Objective Skill Evolution for Agents in Policy-Governed Enterprise Workflows


35. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs


36. MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents


37. CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM Agents


38. LaCache: Robust Semantic Caching for LLM Serving


39. Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness


40. Beyond Single-Use Tokens: Durable Authorization State for Replay-Resistant LLM Agent Actions


41. GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks


42. When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary


43. TCPO: Turn-Level Credit Policy Optimization


44. GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks


45. Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw


46. When Memory Updates but Behavior Does Not: Repairing Implicit Stale Dependencies in Personalized Agent Responses


47. Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning


48. Securing Agentic AI: From Per-Action Checks to Trajectory Assurance


49. Emergence Invariance: From Symbolized Thought to Interface Refinement


50. V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory


51. Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning


52. Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics


53. Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning


54. CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories


55. CRAFTS: Collaborative Role-Adaptive Fine-Tuning of LLM Agents for Chemical Process Simulation


56. High-Stakes Decisions with Language Models: Insights from Emergency Triage


57. G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution


58. Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models


59. Learning What to Remember and What to Internalize in LLM Self-Evolution via Adaptive Memory-Parameter Coordination


60. CT-PrepAgent: Bounded Policy and Controlled Execution for Adaptive CT Data Preparation


61. Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races


62. The Graph Language: How Knowledge Graphs Speak to Large Language Models


63. PATH-Bench: Path-Dependent Evaluation of Lifelong Agents


64. Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating


65. Don’t Offer What Can’t Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale


66. Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models


67. Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets


68. SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling


69. Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel


70. Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question Answering


71. PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent


72. TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents


73. Modeling Social Dynamics with an LLM-Enabled Agent Based Network-Dynamic (LAND) Model


74. CADIR: A Cross-Backend Editable Intermediate Representation for Agentic CAD Generation


75. Isotropy Cliffs: The Geometric Signature of Decision-Making in Large Language Models


76. Large language models improve physician accuracy but lead to false reliance


77. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation


78. FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction


79. Behavioral Grammar: Detecting Adaptive Malware via Tiny Language Model Priors and Second-Order Temporal Analysis


80. Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations


81. DGA$_2$D: Directed Graph-Guided Automated Algorithm Design with Large Language Models


82. When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty


83. DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization


84. Escaping Confidence Trap: Evolutionary Decoding for Mathematical Reasoning in Diffusion LLMs


85. Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations


86. CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding


87. F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Language Models


88. Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LLM Serving Traces


89. TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LLMs


90. SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs


91. Where did the ambiguity go? Examining how multimodal models interpret polysemous words


92. Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates


93. CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization


94. More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness


95. Personalizing Large Language Model Agents with Small Policy Models


96. TRACE-TS: Attribution-Grounded and Traceable Sensor-Language Reasoning for Human Activity Understanding


97. AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?


98. Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce


99. Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale


100. RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models


101. Linguistic Context Recodes Visual Representations in Vision-Language Models


102. SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented and Memory-Grounded LLM Systems


103. Motif-Mamba: network motif improved mamba for long-range sequence modeling


104. Request-Level Energy Attribution for Batched LLM Serving


105. Memory Reward Inflation in Self-Improving LLM Agents


106. Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process


107. CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection


108. Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware


109. Enhancing LLMs with Context-Specific Knowledge for Mitigating Misinformation in SMEs: A RAG-based Modeling and Analysis


110. AutoFOAM: The Self-Refining Autonomous OpenFOAM Agent


111. Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework


112. Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection


113. Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment


114. Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning


115. Antares: Foundation Models for Agentic Vulnerability Localization


116. Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification


117. HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts


118. PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs


119. Self-Improving Large Language Models via Progressive Experience Evolution


120. How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models


121. Geometry-Guided Layerwise FFN Width Allocation in Transformers


122. CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship


123. TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving


124. AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning


125. Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection


126. Effective and Efficient Context Retrieval via Partial Dependency Graph for Repository-Level Code Generation


127. Energy-Efficient LLM Serving via Disaggregated Attention–FFN and Flexible Frequency Scaling


128. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models


129. LAB-Tab: LLM-Augmented Bayesian Network Adaptation for Few-Shot Tabular Generation


130. DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents


131. LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation


132. Illuminating Visual Identity in Universal Multimodal Embeddings


133. PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation


134. EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming


135. Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit


136. X-KGRank: A Knowledge Graph RAG Framework for Explainable Recommendations via Pattern Mining and LLM Re-Ranking


137. TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics


138. MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication


139. Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch


140. Rethinking Generative AI Literacy: An Integrative, Developmental, and Dialectical Framework for K-12 Teacher Education


141. ProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions


142. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation


143. SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction


144. CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models


145. RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection


146. Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge


147. PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters’ Lack of Knowledge


148. HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning


149. Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard


150. Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning


151. Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics


152. Same violence, different answer: how AI responds to coercive control against women across languages


153. When Replanning Becomes the Bottleneck: Budgeted Replanning for Embodied Agents


154. Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+


155. Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety


156. Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution


157. LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning


158. Context Compaction Theory


159. Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents


160. Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design


161. Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere


162. It’s the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling


163. SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs


164. Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization


165. Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception


166. DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text


167. Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks


168. What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents


169. CallScreenBench: Benchmarking On-Device Models as Phone Secretaries


170. Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking


171. Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy


172. Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning


173. MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models


174. Hierarchical Solomonoff Induction: An Unbounded Machine Learning Model


175. Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms


176. Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Ecosystems


177. RefactorAssist: Agentic Refinement for Reliable Code Refactoring


178. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems


179. Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures


180. Coverage-Driven Adaptive Keyframe Selection for Video Understanding


181. From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness


182. Element-Aware Group Learning for E-Commerce Image Generation


183. AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications


184. A Context-Aware Cultural Heritage Guide Powered by LLMs


185. Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages


186. Auditable Release Control for Pedagogical Leakage in LLM Tutors


187. CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings


188. CeQe: Grounding Lexical Retrieval in Semantic Evidence


189. Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks


190. AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction


191. Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments


192. Verifiable Checks for Business Rule Consistency


193. Artificial Intelligence and Modeling & Simulation: An Overview


194. ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression


195. Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct


196. Hybrid Attention Estimation Pipeline for Adaptive HRI Using an Expressive Robotic Head


197. Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind


198. Inference-Time Policy Alignment for Fair Reinforcement Learning


199. A Synthetically-accessible Universe of Chemically Recyclable Polymers


200. DiffusionGemma Technical Report


201. Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity


202. A Fortran General-Purpose Transpiler: Proof of Concept


203. LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations


204. Counting the Cost of War Under Satellite Embargo: Zero-Shot Estimation of Impacted Infrastructure


205. SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models


206. LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology


207. Logographic Character Visual Pretraining via Semantic-based Contrastive Learning


208. Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs


209. Neural Circuit Function Inference with LLMs


210. Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study


211. SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach


212. Width, Memory, and Delay: A Resource Accounting for the Limits of Flat Multi-Agent Systems


213. Role Steering of Language Models for Social Simulations


214. Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models


215. CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs


216. What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs


217. Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams


218. DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis


219. AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents


220. MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents


221. RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review


222. Cost-Effective Automated Judging of Natural-Language Mathematical Proofs


223. Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results