LLM 관련 주요 논문 - 2026-06-10

1. The Role of Feedback Alignment in Self-Distillation


2. ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models


3. ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity


4. What Fits (Into Few Tokens) Doesn’t Overfit: Compression and Generalization in ML Research Agents


5. Superficial Beliefs in LLM Decision-Making


6. Structure from Reasoning, Numbers from Search: On-Premise Open LLMs as Structural Priors for Coupled MIMO Controller Tuning


7. Null-Space Constrained Low-Rank Adaptation for Response-Specified Large Language Model Unlearning


8. Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?


9. Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans


10. Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages


11. Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution


12. Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation


13. Evaluating Research-Level Math Proofs via Strict Step-Level Verification


14. READER: Robust Evidence-based Authorship Decoding via Extracted Representations


15. AutoPDE: Reliable Agentic PDE Solving via Explicitly Represented Solver Strategies


16. Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory


17. One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA


18. ActiveMem: Distributed Active Memory for Long-Horizon LLM Reasoning


19. HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning


20. A complementary study on PlanGPT: Evaluation with defined Performance Metrics and comparison with a planner


21. ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics


22. Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents


23. Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness


24. STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios


25. Instruction Finetuning DeepSeek-R1-8B Model Using LoRA and NEFTune


26. Beyond Static Evaluation: Co-Evolutionary Mechanisms for LLM-Driven Strategy Evolution in Adversarial Games


27. ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience


28. Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning


29. Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts


30. Mobility Anomaly Generation using LLM-Driven Behavior with Kinematic Constraints


31. From Context-Aware to Conflict-Aware: Generalizing Contrastive Decoding for Knowledge Conflict in LLMs


32. Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling


33. Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction


34. RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning


35. Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents


36. From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs


37. EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents


38. Flaws in the LLM Automation Narrative


39. Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation


40. TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning


41. Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA


42. FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model


43. PhantomBench: Benchmarking the Non-existential Threat of Language Models


44. Unifying Local Communications and Local Updates for LLM Pretraining


45. Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models


46. T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains


47. AuRA: Internalizing Audio Understanding into LLMs as LoRA


48. Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning


49. Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature Interactions


50. CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference


51. A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS


52. Optimal Post-Training Quantization Scales and Where to Find Them


53. From Perception to Action: Can UI Interventions Foster Sustainable LLM Chatbot


54. Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs


55. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models


56. K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language Modeling


57. Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks


58. Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use


59. Dep-LLM: Training-Free Depression Diagnosis via Evidence-Guided Structured Multi-factor with Reliable LLM Reasoning


60. Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation


61. Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding


62. Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings


63. Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training


64. Unifying Data, Memory, and Compute Efficiency in LLM training: A Survey


65. Divide and Cooperate: Role-Decomposed Multi-Agent LLM Training with Cross-Agent Learning Signals


66. Decentralized Multi-Agent Systems with Shared Context


67. Dynamic Linear Attention



69. Causal Ensemble Agent: Hierarchical Causal Discovery with LLM-guided Expert Reweighting


70. Towards Diverse Scientific Hypothesis Search with Large Language Models


71. NOVA: Symbolic Regression Discovery of Interpretable Car-Following and Lane-Change Models with Driver Heterogeneity


72. Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction


73. Benchmarking Knowledge Editing using Logical Rules


74. LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization


75. Assessing Automated Prompt Injection Attacks in Agentic Environments


76. Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs


77. Advancing the State-of-the-Art in Empirical Privacy Auditing


78. Decoupling Thought from Speech: Knowledge-Grounded Counterfactual Reasoning for Resilient Multi-Agent Argumentation


79. UPLOTS: A Unified Pretrained Language Model for Constrained Time-series Generation


80. ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs


81. LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake


82. KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data


83. Atomic Intent Reasoning: Bringing LLM Semantics to Industrial Cross-Domain Recommendations


84. Routing-Aware Expert Calibration for Machine Unlearning in Mixture-of-Experts Language Models


85. Baseline-Free Policy Optimization for Neural Combinatorial Optimization


86. Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents


87. The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge


88. LLM-Guided Neural Architecture Search for Robust Co-Design of Physical Neural Networks


89. A Source Domain is All You Need: Source-Only Cross-OS Transfer Learning for APT Anomaly Detection via Semantic Alignment and Optimal Transport


90. Density Ridge Selective Prediction for LLM and VLM Hallucination Detection under Calibration Label Scarcity


91. MMClima: A Framework for Multimodal Climate Science Data and Evaluation


92. $τ$-Rec: A Verifiable Benchmark for Agentic Recommender Systems


93. MetaPlate: Counterfactual-Guided RAG-LLM Tool for Personalized Food Recommendation and Hyperglycemia Prevention


94. Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing


95. What makes a harness a harness: necessary and sufficient conditions for an agent harness


96. Unsupervised Style Representation Learning for AI-Text Detection via Paraphrase Inversion


97. A Controlled Audit of Pretraining Contamination in Public Medical Vision-Language Benchmarks


98. Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces


99. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders


100. Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech


101. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents


102. RKSC: Reasoning-Aware KV Cache Sharing and Confident Early Exit for Multi-Step LLM Inference


103. One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability


104. When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff


105. Trainable Smooth-Rotation Transforms with Learned Channel Scales for LLM Quantization


106. Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling


107. Co-GLANCE: Uncertainty-Aware Active Perception for Heterogeneous Robot Teaming


108. IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference


109. IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts


110. Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History


111. When Attribution Patching Lies: Diagnosis and a Second-Order Correction


112. PreAct-Bench: Benchmarking Predictive Monitoring in LLMs


113. SocraticPO: Policy Optimization via Interactive Guidance


114. SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs


115. TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition


116. Integrating Local and Global Entropy for Uncertainty Quantification in LLMs


117. Rotate2Think: Geometric Priming via Orthogonal Rotation to Improve Language Model Reasoning


118. SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation


119. SPACE: Source-free Proxy Anchor Concept Erasure for MLLMs


120. EstRTL: Functional Estimation Guided RTL Code Generation


121. Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning


122. Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation


123. Blurry Window Attention


124. Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models


125. Using Probabilistic Programs to Train Inductive Reasoning in Large Language Models


126. Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis


127. LLM-Based Code Documentation Generation and Multi-Judge Evaluation


128. CANVAS: Captioning Art with Narrative Visual-Audio AI Systems


129. The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans


130. An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models


131. Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS