LLM 관련 주요 논문 - 2026-07-01

1. PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines


2. Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA


3. TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models


4. Harnessing Textual Refusal Directions for Multimodal Safety


5. Adaptive Cluster-First Route-Second Decomposition for Industrial-Scale Vehicle Routing


6. Large Databases Need Small, Open-Weight Language Models



8. Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision


9. FARS: A Fully Automated Research System Deployed at Scale


10. Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index


11. ACE: Pluggable Adaptive Context Elasticizer across Agents


12. Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2


13. Design and Implementation of Agentic Orchestrations and Orchestration of Agents


14. Surprise as a Signal for Plasticity and Metacognition


15. CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift


16. CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market


17. Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs


18. Xiaomi-GUI-0 Technical Report


19. Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models


20. ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents


21. Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models


22. HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)


23. Benchmarking Large Language Models on Floating-Point Error Classification


24. Spatial Reasoning via Modality Switching Between Language and Symbolic Representation


25. Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling


26. Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents


27. Long-term Traffic Simulation via Structured Autoregressive Modeling


28. Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems


29. Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping


30. AI-Assisted Discovery of Convex Relaxations via Dual Agents


31. ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents


32. Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics


33. Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records


34. The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory


35. DDIAgents: Mechanism-Conditioned Context Flow for Drug-Drug Interaction Prediction


36. MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning


37. OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents


38. AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance


39. RoPoLL: Robust Panel of LLM Judges


40. Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering


41. Investigating Multi-Agent Deliberation in Law


42. When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models


43. BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation


44. Contrastive Reflection for Iterative Prompt Optimization


45. Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision


46. QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents


47. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs


48. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors


49. Amplifying Membership Signal Through Chained Regeneration


50. GR2 Technical Report


51. MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments


52. Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference


53. Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning


54. Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR


55. CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield


56. JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering


57. Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue


58. Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian


59. ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping


60. A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems


61. A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents


62. Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models


63. Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning


64. ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models


65. On the Convergence of Self-Improving Online LLM Alignment


66. FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents


67. Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics


68. DA-Studio: An Agentic System for End-to-End Data Analysis


69. Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?


70. Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?


71. Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models


72. 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance


73. CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs


74. Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents



76. Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?


77. ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries


78. LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music


79. PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding


80. A Modular Vision-Language-Action Robotics Framework for Indoor Environments


81. Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation


82. ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs


83. LLM-Driven Personalities for Decision Making in Emergency Simulations


84. Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG


85. Loc2Repair: A Framework for Evaluating the Impact of File-Level Issue Localization in Repo-Level LLM Repair


86. Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback


87. How Human Feedback Shapes AI-generated Community Notes


88. Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models


89. Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support


90. The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning


91. Test-Time Verification for Text-to-SQL via Outcome Reward Models


92. AI-Generated PowerShell Malware: An Experimental Framework and Dataset


93. When transformers learn “impossible” languages, what do they learn?


94. Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization


95. Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions


96. A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization


97. Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens


98. From Search to Synthesis: Training LLMs as Zero-Shot Workflow Generators


99. BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations


100. LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents


101. Citation Discipline in Spec-Driven Development: A Cross-Model Empirical Study of Output Determinism and Automated Hallucination Detection in LLM-Generated Code


102. Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning


103. ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models


104. Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection


105. Emergent Culture in Minimal LLM Systems


106. ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education


107. The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes