LLM 관련 주요 논문 - 2026-08-06

1. OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling


2. Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite


3. Item Response Theory for AI Safety


4. From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking


5. When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning


6. Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools


7. Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports


8. Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness


9. CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models


10. Architectural Implications of Agentic AI Workflows


11. Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language


12. The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning


13. MatrAIx: Simulating the World with 8.3 Billion Persona Agents


14. BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding


15. FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents


16. FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables


17. Monte Carlo Tree Search for Table-to-Multimodal Report Generation


18. The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents


19. OPD-V: Visual On-Policy Self-Distillation with Modality Balance


20. Chained Recursive Language Models for Multi-Iteration Reasoning


21. Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models


22. Hardware Design and Security in the Era of Chiplets and LLMs


23. The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations


24. Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning


25. ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation


26. OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents


27. ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit Design Optimizer with Large Language Models-Guided Exploration


28. Protoreasoning in Tiny Transformers


29. SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models


30. When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs


31. A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination


32. Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First


33. Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation


34. RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists


35. Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models


36. What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend


37. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO


38. Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports


39. Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark


40. When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models


41. The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering


42. EASy: Towards Efficient LLM-Based Agentic System


43. Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders


44. PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning


45. Breadcrumbing Search Agents


46. EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks


47. A Model Merging Approach for Continual MLLM Unlearning


48. CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding


49. GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs


50. GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction


51. AFD-Ledger: Deployment Provisioning for Attention–FFN Disaggregation


52. Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning


53. D$^2$F-ReAG: Dynamic Decomposition and Filtering for Multi-Hop Reasoning-Augmented Generation


54. ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation


55. MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages


56. Training-Free Hashing-Based Attention via Binary Principal Components


57. FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation


58. Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework


59. HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models


60. Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO


61. COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation


62. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)


63. MIDAS: Multi-LLM Iterative Data-Adaptive Summarization


64. Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems


65. Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary


66. Patients-like-me: A Variational LM–GNN Framework for Explainable Clinical Prediction


67. Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills


68. Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering


69. Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms


70. An Inline Control Architecture for Language Models in Intelligent Transportation Systems


71. AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection


72. Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs


73. EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis


74. RAG-Stack: Co-Optimizing RAG Serving Performance and Quality


75. Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models


76. TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering


77. AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering