LLM 관련 주요 논문 - 2026-09-15

1. Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection


2. Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science


3. AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery


4. Atria Dawn: The Dawn of Agentic Superintelligence


5. EvoOntology: A Self-Evolving Ontology Layer for Data Agents


6. Are LLMs Good Financial User Simulators? A Preliminary Study


7. Data storytelling meets interpretable machine learning: Decoding AI decisions for non-experts without revealing sensitive data and model details


8. New Conditions for Philosophers to Catch the Wave of Citizen Deliberation in the Age of Artificial Intelligence in advance


9. Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction


10. NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities


11. EEG-Xplain: Decoding Neural Black-Boxes of EEG Foundation Models


12. HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments


13. SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution


14. MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving


15. Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting


16. Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures


17. Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain


18. CWM: Controllable White-Box Meta-Prompting for Adaptive Retrieval-Augmented Generation and Reasoning Ability


19. From Ideas to Actions: A Public-Data Decision-Support Toolchain Across the Venture Lifecycle


20. Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election


21. VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries


22. STHMoE: Hypergraph-Enhanced Heterogeneous Dependency Coordination for LLM-Based Urban Traffic Data Forecasting


23. Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes


24. ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models


25. Enabling Creative Exploration for Vibe Design Agents


26. Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML


27. Overflip: Repetition-Induced Label Flips in Guardrail Models


28. CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems


29. Converting Sequenced Fuzzy Cognitive Maps to Causal Virtual Worlds with Large Video Generators


30. MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents


31. Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair


32. Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference


33. El Agente Potente: High-Throughput Agentic Atomistic Simulations


34. ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence


35. Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks


36. Bayesian Intelligence from the Outside


37. DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents


38. MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents


39. Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had


40. Convergent Emergence of In-Context Learning Across Modalities


41. Synthetic Data in Marketing Research: How to Evaluate and When to Trust


42. SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery


43. LoRA Fine-Tuned Models for Control Systems Course Q\&A: A Multidimensional Evaluation of Model Scale and Rank Effects


44. ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information


45. UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics


46. Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture


47. LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems


48. Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs


49. Positioning manuscripts in the scientific landscape with agentic AI


50. How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition


51. Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges


52. Enhancing Event Candidate Acquisition for Event Linking


53. GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents


54. Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself


55. Solar Intelligence


56. FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks


57. Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems


58. Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents


59. Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports


60. OrchSLM: Probing the Dynamics of Small Language Model Orchestration


61. Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures


62. TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models


63. Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?


64. Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents


65. The Router Within: Eliciting Native Skill Routing from a Frozen LLM


66. Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale


67. LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys


68. K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations


69. Before You Poll with LLMs: A Deliberative Diagnostic Framework


70. Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression


71. When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control


72. Look Before You Leap: Factual Decoding with Internal Attribution Signals


73. Don’t Send What You Don’t Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models


74. Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding


75. CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel Optimization


76. Kaininja: Extending Native 3D Generators to the Part Level


77. Predictive Likelihood Ratios for Language Model Watermark Detection


78. ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs


79. VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding


80. A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation


81. Specifying Reward Functions for RL Without Environment Sampling


82. The Misery of Mechanistic Interpretability: A Formal Perspective


83. Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting


84. Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku


85. How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus


86. Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation


87. A Conservative OCR-Enabled Workflow for R214 Sodium Screening of South African Packaged Foods


88. CodeTS: Verifiable Text-to-Time Series Generation via Executable Code


89. IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective


90. Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs


91. Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models


92. Concept-Grounded Reasoning with Prompt-Driven Localization for Interpretable Structured Report Generation


93. Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs


94. The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems


95. Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders


96. Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception


97. TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models


98. EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse


99. AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference


100. MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup


101. Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation


102. Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context


103. Translating the Translator: Decomposing the Cost of English-Forced Inter-Agent Communication


104. Salesforce Koa: An Enterprise Language Model for Agentic Tool Use


105. Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training


106. Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models


107. SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing


108. Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache


109. Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks


110. ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents


111. Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains


112. Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning


113. Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations


114. LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions


115. Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization


116. Enemray: Toward Capable Language Models for Hassaniya


117. Route, Don’t Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection


118. A primer on evaluation methods for large language models in healthcare


119. The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents


120. How broad is that claim? Mapping Generalisation in NLP Research


121. Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination


122. TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps


123. Calibrating Interpretability Instruments Before Trusting Their Verdicts



125. Carryover Drafting: Recycling Rejected States for Speculative Decoding


126. Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing


127. Natural Language Knowledge Graph Query Execution: Leveraging Controlled Semantics in the LLM Context Window



129. Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World


130. AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education – A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory


131. Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection


132. Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation


133. NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass


134. A latent dimension of Condorcet’s jury theorem for multiple AI advisers


135. A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification


136. A Hybrid Dependency-Aware Framework for Task Decomposition and Dynamic Agent Generation in Oracle-to-PostgreSQL Migration


137. LLaTSA: Large Language Model-Aligned General-Purpose Transient Stability Analysis


138. AURA: Unified Multimodal Framework for Conversational Music Editing


139. OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving


140. Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study


141. Modeling, Scaling, and Decoding: Optimizing Controllable Speech Generation with Nonverbal Vocalizations


142. ECAS: An Edge-Controlled Agentic System for Validation-Gated Scientific Application Execution


143. Towards Evolving Context Parameterization for Large Language Models


144. Signatures of Steerability in Activation Space of Language Models


145. LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents


146. RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes


147. AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents


148. Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control


149. Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration


150. Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context


151. CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education


152. When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents


153. CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection


154. Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models


155. PolicyMem: Geometric Policy Memory for LLM Governance


156. TyPatch: Transforming Patches into Typestate Rules for Kernel Bug Detection


157. Oops, Not Now: PEARL, a RAG-Based Support Agent for Gameplay and What Players Want from AI Help


158. Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding


159. LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference


160. An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS


161. Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs


162. Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models


163. Generative Interpretability via Scalable Neuro-Symbolic Models


164. Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning


165. A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion


166. Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment


167. Learning to Solve Hard Problems in RL for LLMs by Never Giving Up


168. ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement


169. Task-Aware Federated Fine-Tuning for MoE-based Large Language Models


170. Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework


171. The Agentic Company OS: Substrate Inversion for Sustained Enterprise Agent Deployment


172. LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents


173. From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration


174. (How) Do MLLMs Report Bistable Images Like Humans?


175. Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding


176. LLMs or Naive Bayes? Old Gems or New Ways


177. From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity


178. Natural-Language to SysMLv2 Translation via Conformance-Driven Iterative Refinement


179. Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB


180. Towards Optimizing SQL Generation via LLM Routing