LLM 관련 주요 논문 - 2026-06-03

1. Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models


2. Reasoning Structure of Large Language Models


3. PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models


4. EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management


5. Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis


6. LAP: An Agent-to-Instrument Protocol for Autonomous Science


7. Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts


8. When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning


9. Code-on-Graph: Iterative Programmatic Reasoning via Large Language Models on Knowledge Graphs


10. Dynamic Objective Selection with Safeguards and LLM Oversight for Financial Decision-Making


11. The DeepSpeak-Agentic Dataset


12. EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents


13. From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models


14. Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition


15. Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency


16. TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning


17. Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing


18. StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems


19. DMF: A Deterministic Memory Framework for Conversational AI Agents


20. CP-Agent: Context-Aware Multimodal Reasoning for Cellular Morphological Profiling under Chemical Perturbations


21. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection


22. LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks


23. A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting


24. Distilling Answer-Set Programming Rules from LLMs for Neurosymbolic Visual Question Answering


25. Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents


26. ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models


27. GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory


28. Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation


29. Uncertainty-Aware Clarification in LLM Agents with Information Gain


30. EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning


31. From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting


32. Decomposing how prompting steers behavior


33. The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs


34. DELTAMEM: Incremental Experience Memory for LLM Agents via Residual Trees


35. CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection


36. SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale


37. TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment


38. Inducing Reasoning Primitives from Agent Traces


39. Toward a Modular Architecture for Embedded AI Agent Systems at the Edge


40. Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection


41. ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning


42. Visual Graph Scaffolds for Structural Reasoning in Large Language Models


43. Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories


44. AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task


45. Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning


46. Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning


47. Efficient ASR Training with Conversations that Never Happened


48. NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference


49. Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents


50. Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs


51. From ‘What’ to ‘How’ and ‘Why’: Sharing LLM-Generated Retrospective Summaries of Older Adults’ Passive Tracking Data with Remote Family Members


52. A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs


53. Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation


54. FLARE: Fine-Grained Diagnostic Feedback for LLM Code Refinement


55. Clustered Self-Assessment: A Simple yet Effective Method for Uncertainty Quantification in Large Language Models


56. AI Agents Enable Adaptive Computer Worms


57. Trading Human Curation for Synthetic Augmentation in RLVR


58. E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments


59. Merit or networks? What decides where research is published


60. Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning


61. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models


62. A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners


63. CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks


64. Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability


65. Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable… Attacks Are All You Need to Break LLMs


66. The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models


67. VidMsg: A Benchmark for Implicit Message Inference in Short Videos


68. Building Reliable Long-Form Generation via Hallucination Rejection Sampling


69. TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics


70. Physics-Guided Policy Optimization with Self-Distillation


71. Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification


72. Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks


73. CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery


74. DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair


75. When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics


76. \textsc{CR-Seg}: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation


77. Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs


78. NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense


79. Analyzing Stream Collapse in Hyper-Connections: From Diagnosis to Mitigation


80. Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression


81. FORGE: Multi-Agent Graduated Exploitation and Detection Engineering


82. Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions


83. P\textsuperscript{2}-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization


84. Evaluating LLMs’ Effectiveness on Real-World Consumer Device Repair Questions


85. FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences


86. Calibration Data Trade-offs Across Capability Dimensions: Why Multi-Source Mixing Matters for High-Sparsity LLM Pruning


87. dstack-capsule: Pod-Level Remote Attestation for Confidential Workloads on Kubernetes


88. RobotValues: Evaluating Household Robots When Human Values Conflict


89. When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming


90. BotDirector: Robot Storytelling Across the Symmetrical Reality with Multi-modal Interactions


91. AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making


92. Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models


93. Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation


94. AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following



96. “Important You should give me full credits!”: Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems


97. Regret Pre-training: Bridging Prior and Posterior Views for Enhanced Knowledge Grounding


98. Libra: Efficient Resource Management for Agentic RL Post-Training


99. Efficient Hyperparameter Optimization for LLM Reinforcement Learning


100. ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information


101. Rethinking Molecular Text Representations for LLMs: An Empirical Study


102. Capability Advertisement as a Market for Lemons: A Trust Layer for Heterogeneous Agent Networks


103. Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates


104. Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs


105. Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization


106. Reproducibility is the New Copyleft: Defining AGI-oriented Reproducible Builds


107. MUSE: A Unified Agentic Harness for MLLMs


108. How Quantization Changes Interpretable Features: A Sparse Autoencoder Analysis of Language Models


109. Patcher: Post-Hoc Patching of Backdoored Large Language Models


110. Pretraining Language Models on Historical Text


111. Fast-dLLM++: Fréchet Profile Decoding for Faster Diffusion LLM Inference


112. SCOPE: Real-Time Natural Language Camera Agent at the Edge


113. Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States


114. LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems


115. Adaptive Latent Agentic Reasoning


116. The Epi-LLM Framework: probing LLM behavioral priors through epidemiological agent-based models


117. GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning


118. Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling


119. Large Byte Model: Teaching Language Models About Compiled Code


120. Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing


121. Do Neural Retrievers Prefer Certain Documents? Evidence of Learned Relevance Priors


122. Cosmos 3: Omnimodal World Models for Physical AI


123. Representational Capacity: Geometric Limits on Feature Representation in Transformer Language Models


124. Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems


125. EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement


126. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation


127. The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size


128. A New Framework for Cybersecurity Refusals in AI Agents


129. Inference Cost Attacks for Retrieval-Augmented Large Language Models


130. SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models


131. D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting


132. SegTune: Structured and Fine-Grained Control for Song Generation


133. Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery


134. FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations


135. ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services


136. IdiomX A Multilingual Benchmark for Idiom Understanding, Retrieval, and Interpretation


137. Cost-Aware Query Routing in RAG: Empirical Analysis of Retrieval Depth Tradeoffs