LLM 관련 주요 논문 - 2026-09-22

1. Emergent Collusion in Long-Horizon LLM Agent Interaction


2. Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models


3. Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection


4. GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes


5. Construting Reverse Thinking: Developing Large Language Models’ Reverse Thingking Ability


6. TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction


7. DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security


8. Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards


9. VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning


10. Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs


11. LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning


12. How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation


13. Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement


14. SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision–Language Models


15. LIMIT: Less Is More for Instruction Tuning in Text-to-SQL


16. APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction


17. Self-Healing Harness for Runtime Oversight of Agent Self-Modification


18. EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation


19. DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents


20. Incremental Consistency Execution for Autonomous Intelligent Systems


21. Representation-guided in-context learning for medical image interpretation with multimodal large language models


22. Structured Decomposition for Reliable LLM-Generated Access Control Policies


23. Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech


24. Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure


25. FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering


26. Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents


27. Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery


28. Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models


29. Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows


30. PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI


31. Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment


32. AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction


33. TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents


34. Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation


35. CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine


36. From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving


37. Tutoring Large Language Models to be Domain-adaptive, Precise and Safe


38. FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics


39. Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees


40. Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World


41. PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models


42. OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling


43. Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model


44. ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures


45. ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control


46. Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems


47. Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation


48. MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators


49. AutoGym: Blueprint-First Generation of Verifiable Agent Gyms


50. IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law


51. Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus


52. The Wisdom of Artificial Deliberative Crowds


53. Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation


54. Goal-driven Variant Categorization


55. Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models


56. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses


57. Rare Event Estimation via Iterative Unalignment


58. OSWorld-Pro: Process-based Evaluation for Computer Use Agents


59. Small-world Networks of Agents Brainstorm AI Risks to Support Ideation


60. SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture


61. Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection


62. When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs


63. PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning


64. LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot


65. Reasoning Topology Matters: A Controlled Study of LLM-Based Cybersecurity Analysis


66. Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference


67. iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs


68. From Semantic Decisions to Feasible Trajectories: Self-Evolving LLM-Guided Optimal Control for Narrow-Space Parking


69. Augmented Hypothesis Testing with Persona-Based LLM Simulations


70. QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation


71. AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos


72. VPRune: Efficient Training-free Pre-LLM Visual Token Pruning


73. Do LiDAR Language Models Really Understand Spatio-temporal Relationships?


74. ActGov: Governing LLM Agent Actions via Policy-Constrained Validation


75. ARM: Attention with Routed-Memory for Learnable Sparse Control


76. Information-Time Proximal Policy Optimization


77. URA-NER: A Unified Retrieval-Augmented Framework with Retrieval Alignment and Uncertainty Reduction for Low-Resource NER


78. DeceptionAnalyser: A Web-Based AI Tool for Performing Structured Deception Analysis with Argumentation Schemes and LLMs


79. Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection


80. Mitigating Entity Type Confusion in Cross-Domain NER via Multidimensional Quantification and Reasoning Enhancement


81. TTSE: A Two-Track Online Self-Evolution Framework


82. MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents


83. Taramandal-GPT: Enhancing Astrodynamics Problem-Solving with Knowledge Retrieval and Structured Thinking


84. Memory vs. Context? Influential Factors of Factual Recall in Language Models


85. MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents


86. TAC-Time: Texts as Channels For Multimodal Time Series Forecasting


87. WidgetVA: A Widget-Centric Framework and Benchmark for Agentic Visual Analytics


88. From Bits to Beliefs: Recoverable Semantic Fingerprints for Black-Box Verification of Large Language Models


89. Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models


90. RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations


91. Djinnlang: Higher-Level Programming by Unambiguous Specification with an LLM in the Compiler


92. HaikuS2S: A Cascaded System For Responding In Verse


93. Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges


94. SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses


95. From UNDRR Reports to Event Records: Schema-Constrained LLM Extraction of Georeferenced Disasters


96. Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking


97. FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model


98. TriFleetRCA: On-Premise LLM Root Cause Analysis for Kubernetes


99. GRACE: Grounded Adversarial Reasoning over Canadian Law


100. When the Agent Becomes the Kernel: A Systematization of Security on the Path to AI-Native Operating Systems


101. PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding


102. PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models


103. Listen Then Reason: Perception-Grounded Test-Time Reinforcement Learning for Large Audio-Language Models


104. RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents


105. WaveletECO: A Closed-Loop Physical ECO Platform and a Specialized Local Language Model


106. Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference Services


107. Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines


108. Semantic Candidate-Job Matching: A Comparative Evaluation of Dense Embedding Models in Hybrid Retrieval


109. Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents


110. LLMs as Linguistic Chameleons: Decoupling Semantics and Structure for Privacy-Preserving Communication


111. QwenVLConnector: A Fast, Unified Medical VLM Chatbot for Fine-Grained Clinical Perception and Text Generation


112. DiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic Simulation


113. MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs


114. From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness


115. Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity


116. Automatic multimodal UX improvement recommendations from LLM agent user simulations


117. Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems


118. Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling


119. Block-Sparse Attention with Semantic-Geometric Decoupled Routing


120. The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI


121. Per-Query Gating of LLM Rerankers for Multi-Hop Retrieval


122. Towards Full Pipeline FP8 Reinforcement Learning for LLMs


123. Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs


124. Testing the Construct Validity of a Functional Valence Axis in LLM Agents


125. The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents


126. Commonsense-Grounded Path Planning from Abstract Instructions


127. Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation


128. SelfOp: An Optimization Algorithm for Self-Improving Security Agents


129. ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding


130. MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning


131. LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator


132. Vision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD Modeling


133. From Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference Implementation


134. Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation


135. From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly


136. Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation


137. Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act


138. SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning


139. Zero-Trust Authorization and Discovery for Enterprise MCP


140. FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation


141. A Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support


142. Toward Personalized Sleep Guidance from Wearable Data Using Language Models


143. Contextual Causality with Large Language Models: A Survey


144. Visual Graph Reasoning via Knowledge Compilation


145. Authority-Preserving Evaluation of Medical Vision-Language Assistants


146. Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models


147. ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI


148. Large language models in medical time series analysis


149. Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection


150. RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks


151. Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning


152. Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations


153. Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation


154. CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models


155. CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation


156. Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews


157. The Corroboration Illusion: When More News Makes LLM Forecasts Less True


158. Do Chess Explanations Reflect Model Decisions? Behavioral and Token-Level Tests of LLM Reasoning Faithfulness


159. Replay-Gated Neural Execution: Decoupling Persistent Behavioral Specifications from Neural Realizations in Frozen Language Models


160. H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution


161. Knowledge Graph-Augmented Ambient AI for Clinical Note Generation


162. Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark


163. Team DArgk at the 2026 ELOQUENT lab for evaluating generative language model quality: Residuals of Humanity: AI Detection Evasion via GRPO Fine-Tuning


164. Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B


165. Toollery: Scaling LLM Agents to Thousands of Skills and Tools


166. PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations


167. Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers


168. The Role of AI in Online Reviews


169. EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines


170. The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation


171. Multiple latent orderings better predict language model preferences


172. Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs


173. Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models