LLM 관련 주요 논문 - 2026-07-21

1. SGA: Plug&Play Geometric Verification for Educational Video Synthesis


2. Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering


3. Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs


4. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting


5. AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models


6. Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective


7. OntoExtend: A Framework for Requirement-driven and Scalable Ontology Extension with LLMs


8. Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking


9. PEARL: Auditable Repair for Scientific Reasoning Graph Extraction


10. Stress Testing Concept Erasure with Large Language Model Agents


11. ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding


12. Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory


13. PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model


14. WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement


15. SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation


16. Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution


17. LaT: LLM-as-Trainer for Multi-Task Vehicle Routing Solvers


18. OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment


19. FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models


20. Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents


21. A Dual-Hypothesis Reasoning Framework for LLM Guardrails


22. ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions


23. Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?


24. Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory


25. Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows


26. Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation


27. Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions


28. Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift


29. Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning


30. Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware


31. An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation


32. LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning


33. Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost


34. Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories


35. Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-Start Prediction


36. When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering


37. Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making


38. Lomekwi: Resource-Bounded Tool Discovery in LLM Agents


39. Environment-free Synthetic Data Generation for API-Calling Agents


40. Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification


41. AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents


42. FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images


43. RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Planning


44. Constraint-Anchored Reasoning Traces


45. RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts


46. DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening


47. TopoTuner: Topological Finetuning of Large Language Models


48. Just A Rather Very Intelligent Spoken Agent


49. Interactive Task Alignment as a POMDP


50. LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models


51. RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents


52. SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation


53. Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers


54. Accurate and Efficient Long-Term Memory for LLM Agents


55. Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning


56. JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models


57. PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization


58. It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches


59. Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL


60. Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment


61. Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models


62. Deterministic Replay for AI Agent Systems


63. PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection


64. Some Large Language Models Exhibit Consistent Risk Attitudes


65. Automated Discovery Has No Universally Superior Harness


66. Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs


67. OR Else: A Differentiable Trust Region for Policy Optimization


68. LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications


69. Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints


70. SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs


71. Human Grounded Evaluation of Large Language Models for Optical Network Automation


72. Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security


73. Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation


74. Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation


75. MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models


76. HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization


77. Harness Engineering for LLM-Driven GPU Kernel Generation


78. A Geometric Perspective on Stabilizing Value Conflict Resolution


79. DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration


80. Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI


81. Persona-as-Configuration: Generative Stakeholder Reporting for Agricultural Floods


82. Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence


83. FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches


84. Autonomous Discovery of Wireless Communications Algorithms


85. MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference


86. Uncovering Latent Reasoning Strategies in Language Models


87. Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture


88. Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification


89. CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning


90. CommitLLM: A Fine-Tuned Pipeline for Git Commit Message Generation


91. Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration


92. SALT: Salience-Aware Lexical Trie for Long-Context Compression


93. CoEvoP&R: Co-Evolving Placement Objectives with Routing Feedback via Large Language Models


94. Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory


95. SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation


96. AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization


97. Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph


98. Distilled Reinforcement Learning for LLM Post-training


99. Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat


100. Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs


101. ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts


102. EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding


103. Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning


104. TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization


105. Trace-Based On-Policy Distillation for Masked Diffusion Language Models


106. Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models


107. Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration


108. JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models


109. How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice


110. Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits


111. TellTale: Blending Multi-Instance LoRA Text Encoders and a Zero-Shot LLM Judge for Ambivalence/Hesitancy Recognition in Videos


112. DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines


113. Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL


114. Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings


115. ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents


116. Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM


117. Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent


118. Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models


119. AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language


120. PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation


121. AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows


122. GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs


123. Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs


124. LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models


125. Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation


126. From Intent to Infrastructure: LLM-Driven Agent Compilers for ISAC Networks


127. Autonomous mechanistic discovery of colorectal cancer vulnerabilities via multi-scale AI swarms


128. From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training


129. Feature Generation Using LLMs: An Evolutionary Algorithm Approach


130. SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling


131. An Agentic Interface for End-to-End Probabilistic Seismic Hazard and Risk Analysis


132. High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration


133. Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation


134. CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents


135. RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants


136. TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment


137. KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?


138. From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language


139. What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?


140. DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth