LLM 관련 주요 논문 - 2026-08-05

1. ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning


2. Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations


3. Interpretable Adaptive Sampling for LLM Test-Time Scaling


4. TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring


5. The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections


6. Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition


7. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?


8. ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories


9. MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents


10. LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards


11. Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation


12. KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation


13. Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement


14. CARE-Bench: Benchmarking Patient-Facing LLM Triage


15. When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives


16. LiveEvalBench: Toward Open-World Evaluation for Web Generation


17. Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training


18. AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery


19. When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation


20. Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection


21. Formal Verification of Agentic Systems over Operational Data


22. Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents


23. FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation


24. Large language models for partial differential equation workflows


25. From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities


26. Enhancing Tabular Learners with Context-Aware Semantic Embeddings


27. Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve


28. Reversing Arrows in Large Language Models


29. When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs


30. Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks


31. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design


32. ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning


33. ChartAnno: Evaluating MLLMs for Chart Annotation Generation


34. LeanMem: Simple and Efficient Long-Term Memory for LLM Agents


35. LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models


36. Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory


37. AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction


38. Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems


39. MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc


40. Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks


41. DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning


42. AgentPanel: Toward a New Paradigm for Human–AI Collaboration in Exploring Scientific Questions


43. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning


44. Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains


45. The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems


46. Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation


47. Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning


48. Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents



50. Don’t Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR


51. Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls


52. TraceCAD: Trace-Guided Repair for Agentic CAD Generation


53. CastFSR: A Fast–Slow–Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting


54. LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment


55. UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks


56. When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning


57. Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes


58. VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space


59. HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents


60. Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling


61. ISEE: Interactive Semantic Enrichment for Database Fields


62. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility


63. Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?


64. Separating quantum circuits from classical LLMs


65. Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility


66. When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding


67. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement


68. MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning


69. Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning


70. GENESIS: Towards Explainable Causal Discovery


71. Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking


72. UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space


73. VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs


74. Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure


75. Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss


76. Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks


77. MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models


78. Can LLMs Test Terminal User Interfaces?


79. GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models


80. Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation


81. How Closely Do LLM Reviews Align with Human Peer Review?


82. MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble


83. A Security-Oriented Lifecycle Model for Large Language Model Systems


84. DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction


85. AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality


86. Training Documents Reranker with Search Rubrics for Deep Research Agent


87. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels


88. Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs


89. OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet


90. Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces


91. FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact


92. The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk


93. Route-Align-Verify for Functional Correctness in Code Generation


94. Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform


95. The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics


96. GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs


97. Agentic Reinforcement Learning with Self-Distilled Reward Shaping


98. EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners


99. Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach


100. Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation


101. Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning


102. Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning


103. Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models


104. CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation


105. CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning


106. Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping


107. PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory


108. Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation


109. LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs


110. PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning


111. A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning


112. SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels


113. TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation


114. Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits


115. SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling


116. BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?


117. MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory


118. CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning


119. In-Context Collapse in Vision-Language Models and How to Mitigate it?


120. Learning a Vector-Symbolic Model for Socio-Cultural Tasks


121. A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation


122. SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology


123. Don’t Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators


124. Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach


125. Output-Aware Rotation for INT2 KV-Cache Quantization


126. A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models


127. $S^3$: Improving Agent Safety through Multi-Stage Defense


128. TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows


129. DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial


130. When Policies Change Probabilities: Modular Decision-Making for LLM Code Review


131. Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation


132. Single Canonical Prompts Underestimate LLM Safety’s Surface-Form Sensitivity


133. Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures


134. IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation


135. Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation


136. Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models


137. Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety


138. OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning


139. KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization


140. PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement