LLM 관련 주요 논문 - 2026-08-26

1. A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments


2. Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA


3. Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought


4. StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing


5. Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav


6. RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons


7. Meta$^n$: Recursive Self-Improvement through Emergent Depth


8. Confident at the moment of action: belief miscalibration in LLM play under hidden information


9. The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models


10. Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning


11. PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos


12. Joint Optimization of Tool Creation and Use for Large Language Model Agents


13. EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents


14. When “Must” Becomes “Maybe”: Constraint Weakening in LLM Agent Workflows


15. Discovering Adaptive Transmission Programs for Collective Innovation


16. Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models


17. PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents


18. HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning


19. A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation


20. ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping


21. Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs


22. Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems


23. The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents


24. Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning


25. SonarLLM: A Native Sonar–Optical Multimodal Large Language Model for Underwater Perception


26. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight


27. OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning


28. VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models


29. RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards


30. Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing


31. Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks


32. SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction


33. Evaluating Multiple LLM Generations with Validated Task Coverage


34. Constraint-Guided Enterprise Data Mapping with Large Language Models


35. MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG


36. Preference Data Selection for Mitigating the Alignment Tax in Large Language Models


37. Paritok-4B: Intent-Conditioned Context Compression for Coding Agents


38. Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping


39. AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL


40. ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation


41. EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals


42. Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression


43. Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems


44. Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment


45. Relative Time Intervals Representation for Word-level Timestamping with Masked Training


46. Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding


47. Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing


48. Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction


49. Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning


50. When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs


51. Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design


52. More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving


53. More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight


54. Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining


55. MARS: Multi-Specialist LLM Relay System for Competitive Programming


56. Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment


57. Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems


58. BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification


59. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors



61. SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models


62. Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention



64. Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware


65. AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace


66. Do LLMs Understand Limit Order Book Dynamics?


67. Automata from Agent Traces: Failure and Next-Step Prediction


68. Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering


69. MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models


70. Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value


71. Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes


72. Function-Level Execution Feedback for Code Preference Optimization


73. TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery


74. LLM Agents Perform Controlled Experiments Using Simulation Models


75. RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation


76. Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows


77. Automatic Model Card Generation Using an LLM


78. The RAT: A Unified Bayesian Model for RAG Evaluation


79. On-policy Distillation with Verifiable Reward


80. Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems


81. A Literate Programming Environment for Human and Machine Agents


82. COCI: Conference Organisers and Content Identifier


83. When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study


84. FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision


85. SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling


86. Contrastive Branch Policy Optimization


87. ‘Ghaib in Translation’ aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with ‘Missed-in-Urdu’ Scores in LLM Hate Speech Detection


88. Preference Optimization for Non-Verbal Vocalization Synthesis


89. LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes


90. PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment


91. PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control


92. Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection


93. Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents


94. PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding


95. When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs


96. VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference


97. Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings


98. What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions


99. WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents


100. SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding


101. Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation


102. The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem


103. RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation


104. Evaluating Language Models on Cross-Language Code Functional Equivalence


105. NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution


106. The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses


107. RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding


108. Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs


109. Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)


110. Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring


111. Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders


112. EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis


113. When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk


114. TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers


115. EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$


116. The Limits of Automatic Evaluation of Creativity in Large Language Models


117. Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model


118. From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers


119. Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap


120. Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling


121. Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail


122. ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents


123. When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs


124. Identifying Latent Declarative Representations of Code for Assisting Repository Migration


125. REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring