LLM 관련 주요 논문 - 2026-08-27

1. SwarmWorld: Stigmergic technological evolution in societies of language-model agents


2. Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems


3. AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs


4. ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs


5. Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs


6. SciMIF: Understanding Multimodal Instruction Following in Scientific Domains


7. LivingRAG: Augmenting Graph RAG with Experience


8. Candidate supply and answer selection shape the value of LLM judging in multi-agent systems


9. How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation


10. Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems


11. ToST: A Tree-of-Thought Socratic Teaching Framework for Multi-Path Guidance and Parallel Thinking


12. Narcissus: Program Synthesis Using Context-Aware LLM Approximations


13. CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval


14. Training Alignment Auditors via Reinforcement Learning


15. Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness


16. Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs


17. FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review


18. FLARE: Verifying MILP Reformulations with LLM-Based Theorem Proving


19. LLM-Driven, Datasheet-Aware Automated Hardware Compatibility Verification for Early-Stage, Pre-Schematic Embedded System Design


20. Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows


21. Tunable Tool-Call Rates in LLM Agents via Representation Steering


22. FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs


23. Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning


24. PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?


25. Auto-Policy, not Auto-Skill: Compiled Agent Skills for the Physical World


26. LifePlanner: Evaluating LLM Agents for Geo-spatial Planning with Social Media Data


27. Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs


28. Semantic Graph Unification for Industrial Digital Threads: Bridging 11 Heterogeneous Manufacturing Systems Through Ontology-Driven Knowledge Graphs


29. Natural Language Input, Semantic Track Representation, and LLM Inference: Making the Maritime Information Exchange Model Tractable


30. Reliable LLM-Powered Decision Engines for Large-Scale Supply Chain Operations: Architecture, Safety, and Performance Guarantees


31. VLM-based automatic multi-granularity graph representation of building layouts for design informatics


32. TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development


33. Prefix Sliding for efficient test-time scaling


34. VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction


35. One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation


36. TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding


37. Code World Model: Coding Agent as World Brain


38. Unlocking Multimodal Protein Language Models at Inference Time


39. Skill Issue: Are Skills Language-Invariant in LLMs?


40. EVOMAL: Self-Poisoning in Self-Evolving Coding Agents


41. Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models


42. When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies


43. AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation


44. Data Citation for Large Language Models: A Challenge


45. From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis


46. Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty


47. Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context


48. Leveraging Inter-object Affordances for Efficient Planning in Contact-rich Tasks


49. V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning


50. PolyMemDB: A Polyglot Database System for AI Memory Management


51. MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations


52. Controllable Affective Generation via Latent Vector Steering


53. Goodput Maximization for Large Language Model Edge Inference: A Two-Phase Maskable PPO Approach


54. MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities


55. Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction


56. MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration


57. VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality


58. DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models


59. RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps


60. Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks


61. Q&A or Document-Based? The Effects of Interface Type on How Screen Reader Users Access Interconnected Documents


62. Adaptive Triggering for Bias Correction in LLM Reasoning


63. CRAMER: Control via Request-Aware Masking for Editing Recommenders


64. LLMscope: Extracting LLM Assets from Edge AI Chips via Optical Probing


65. InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance


66. Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation


67. What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift


68. Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making


69. Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment


70. Self-Explanation Tutor for Active Study of CS1 Worked Examples


71. AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models


72. Belief Cascades Drive Persuasion in LLM Agent Networks


73. SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation


74. GRAPE: Gradient Refinement and Progress-Aware Exploitation for Query-Efficient High-Dimensional Bayesian Optimization


75. Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures


76. HealthBench-Psych: A Mental Health Subset of OpenAI’s HealthBench


77. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration


78. DataKernelBench: Can LLMs Optimize Database Queries on GPUs?


79. Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels


80. ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies


81. A Primer on Computational Semantics for Artificial Intelligence Systems


82. Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal


83. Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation


84. Targeting the Attention Heads Behind Object Hallucination in LLaVA


85. Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace


86. The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline


87. Demystifying Reinforcement Learning Post-Training of Language Models


88. FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference


89. A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards


90. Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation


91. Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment


92. Analyzing and Correcting Benevolence Bias in Large Language Models


93. From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction


94. Beyond the Chatbot: Co-Learning and Co-Teaching through a Dual-Persona Generative-AI Assistant


95. aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI


96. MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents