LLM 관련 주요 논문 - 2026-09-08

1. Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe


2. Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence


3. Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models


4. Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models


5. LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams


6. Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness


7. RISE: Recursive Improvement via Self-Extrapolating Policy Distillation


8. Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods


9. GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity


10. Testing Interchangeability in LLM Agent Teams


11. Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference


12. Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents


13. Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory


14. Uncensored Open-weight Models: Redistribution as the Persistence Layer


15. ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs


16. CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review


17. What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection


18. Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens


19. LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28


20. TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents


21. Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment


22. TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing


23. Language models judge war differently when tested for alignment


24. Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing


25. Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing


26. From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments


27. AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems


28. LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models


29. CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution


30. MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models


31. ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults


32. When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models


33. ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing


34. DODR: Deterministic Operator-Driven Reasoning in Latent Space


35. Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges


36. Shadow Queries for Private Retrieval in Vector Databases


37. DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems


38. Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM


39. PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces


40. FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality


41. Model Retirement Creates Reproducibility Risk in Biomedical AI Publications


42. SQL-Zero: Self-Evolving Text-to-SQL


43. ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies


44. Harness-agnostic detection and immunization of reward hacking in self-evolving language models


45. Continual Graph Memory for Adaptive Recommendation under Intent Drift


46. SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents


47. $τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction


48. Extremely Sparse Supervision Incentivizes Reasoning Ability


49. La Agente Óptima: Towards Agentic Self-Driving Laboratories


50. IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion


51. From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs


52. MaxKernel: Agentic Kernel Generation for TPUs


53. HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals


54. Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer


55. Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets


56. A Removal Based Approach to Improve LLM Faithfulness at Test-Time


57. From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance


58. When LLM Decompilers Recompile More and Preserve Less


59. The History Is the Detector: Executing CVE Patch History, End-to-End


60. CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls


61. Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization


62. PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting


63. A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR


64. AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics


65. A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment


66. A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning


67. TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors


68. A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support


69. Beyond Co-purchase Relation: Evolution of Complementary Recommendations at Allegro


70. Qlippy: A Retrieval-Augmented GenAI Assistant for Reproducible Quantum Workflows and Experiment Tracking


71. Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection


72. How a Chatbot’s Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI


73. ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems


74. Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair


75. Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents


76. CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation


77. MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain


78. MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate


79. Reinforcement Learning for improving Large Language Models’ Catalan text simplification capabilities


80. Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution


81. Can Activation Steering Capture Multidimensional Authorship Style?


82. Persistent Teacher Anchoring for Tool-Using Agents


83. Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models


84. Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models


85. Tracing Audio Grounding and Answer Selection in Audio LLMs


86. PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning


87. When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models


88. Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics


89. Hakken: Predicting future discoveries to fill the gaps in today’s knowledge


90. Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective


91. Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning


92. Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation


93. GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion


94. REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation


95. A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models


96. Abstraction Agent


97. Evidence Integration in Large Language Models


98. Scalable Context Orchestration for Serving LLMs Over Voice


99. When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs


100. AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks