LLM 관련 주요 논문 - 2026-06-19

1. Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages


2. What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?


3. Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe


4. SoftSkill: Behavioral Compression for Contextual Adaptation


5. Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving


6. Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference


7. Thermodynamic Measure of Intelligence


8. QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation


9. Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact


10. BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling


11. Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring


12. Process-Verified Reinforcement Learning for Theorem Proving via Lean


13. The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self


14. Multi-Agent Transactive Memory


15. A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models



17. CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models


18. ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?


19. Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning


20. Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks


21. GLARE: A Natural Language Interface for Querying Global Explanations


22. Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents


23. AI4SE and SE4AI Exploration: A Decade Looking Back and Forward


24. Which Pairs to Compare for LLM Post-Training?


25. Analyzing the Narration Gap in LLM-Solver Loops


26. Uncertainty Decomposition for Clarification Seeking in LLM Agents


27. Emergent Alignment


28. LLM Doesn’t Know What It Doesn’t Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data


29. DeXposure-Claw: An Agentic System for DeFi Risk Supervision


30. Hidden Anchors in Multi-Agent LLM Deliberation


31. Diffusion Language Models: An Experimental Analysis


32. Deontic Policies for Runtime Governance of Agentic AI Systems


33. How Transparent is DiffusionGemma?


34. Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software


35. Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems


36. Multi-View Decompilation for LLM-Based Malware Classification


37. LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems


38. AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning


39. ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval


40. Editorial Alignment: A Participatory Approach to Engaging Editorial Expertise in LLM-mediated Knowledge Dissemination


41. The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse


42. SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs


43. ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments


44. Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs


45. MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization


46. From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models


47. Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform


48. IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources


49. The Hidden Evolution of Disguised Visual Context inside the VLM


50. AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models


51. When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents


52. Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution


53. StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation


54. Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning


55. Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services


56. ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models


57. Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA


58. FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming


59. Large Language Models Do Not Always Need Readable Language


60. CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility – Semantic Metrics and Convergence Analysis


61. Uncertainty-Aware Reward Modeling for Stable RLHF


62. Agentic Electronic Design Automation: A Handoff Perspective


63. Towards Engineering Scaling Laws with Pretraining Data Composition


64. SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling


65. Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models


66. Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings


67. NRITYAM: Language Models Meet Art and Heritage of Dance


68. Library-Aware Doubles and Iterative Repair for Large Language Model-Generated Unit Tests in OpenSIL Firmware


69. AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing


70. FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs


71. Efficiently Representing Algorithms With Chain-of-Thought Transformers


72. LOKI: Memory-Free Null-Space Constrained Lifelong Knowledge Editing


73. Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language


74. VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions


75. StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns


76. FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines


77. IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows


78. PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models


79. Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices


80. Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks


81. Secure Coding Drift in LLM-Assisted Post-Quantum Cryptography Development: A Gamified Fix


82. JustDiag!: A Diagnostic Justification Engine for Accountable Root Cause Analysis


83. VERITAS: Verifier-Guided Proof Search for Zero-Shot Formal Theorem Proving


84. Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement


85. DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling


86. Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees


87. Trustworthy Multi-Agent Systems: Mitigating Semantic Drift with the Argent Signaling Protocol


88. Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning


89. Where to Place the Query? Unveiling and Mitigating Positional Bias in In-Context Learning for Diffusion LLMs via Decoding Dynamics


90. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence


91. How LLMs Fail and Generalize in RTL Coding for Hardware Design?


92. Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer


93. Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts


94. Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation