LLM 관련 주요 논문 - 2026-08-14

1. QuoteBench: How Matched Scores Can Hide Command-Path Failures


2. MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination


3. RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level


4. Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes


5. LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning


6. LLM-Guided Graph Generation for Structure-Based Local Improvement Methods


7. StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems


8. vToken: Token-Level Virtualization for Reclaimable KV Caches


9. TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems


10. Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents


11. SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents


12. Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement


13. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback


14. SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference


15. Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds


16. Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)


17. DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition


18. Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence


19. Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence


20. Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents


21. PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs


22. Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies


23. The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis


24. Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs


25. Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy


26. On the Expressive Power of Transformers


27. Designing AI Pipelines for Decision-Ready ITSM Intelligence


28. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries


29. Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence


30. Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues


31. Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces


32. Trie Automata for Constrained Decoding over Large Finite Sets


33. $\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution


34. Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents


35. Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction


36. Research Assistant: AstraZeneca’s Agentic System for R&D


37. Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization


38. Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation


39. Don’t Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese


40. Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing


41. Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments


42. Diagnostic Foundation for Evaluating LLMs’ Research Integrity as Co-Scientists


43. LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure


44. DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data


45. AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models


46. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification


47. CAPRI: Contract-Aware Proof Repair for Isabelle


48. Algebraic Decomposition Theory for Transformer Length Generalization


49. Are You Sure You’re Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity


50. Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference


51. It’s How You Ask: Gender-Associated Linguistic Bias in LLMs


52. Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services


53. Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model


54. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures


55. Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models


56. CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport


57. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint


58. Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering


59. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation


60. EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory


61. Operationalizing Cyber Threat Intelligence with GraphRAG


62. UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations


63. Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents


64. Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference


65. InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers


66. NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents


67. FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation


68. From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options


69. CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives


70. Memorization Diagnostics for Code LLMs Should be Scale-Aware


71. SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization


72. PatientAct: Theory-Grounded Mental Health Client Simulation


73. ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval


74. Error-Aware Reverse Auction Mechanism for Large Language Model Routing


75. Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks


76. Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging


77. Novels generated by language models show compressed formal variation


78. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory


79. LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning


80. Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks


81. Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling


82. Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review


83. A Hierarchical Energy-Based Model for Multimodal Cognition


84. Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models


85. Query Timing Produces Opposite Positional Biases Between LLMs and Humans


86. From Observation to Intervention: Memory in Brains and Large Language Models


87. Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities


88. Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF


89. From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks


90. StreamReason-Bench: Can Large Language Models Reason about Event-Time Stream-Processing Semantics?


91. Mimicry without understanding: the origins of decision bias in large language models


92. StorySpark: Module-wise Evolutionary Search for Story Premise Generation


93. Steering the Language Axis: From Linear Decodability to Causal Control


94. Vision-Language Models are Fragile Multilingual Associators


95. Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching


96. AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement


97. When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models


98. Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance


99. What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting


100. The AI Accountability Ecosystem in the Era of Language Models