LLM 관련 주요 논문 - 2026-06-11

1. Nonslop: A Gamified Experiment in Human-AI Collaborative Writing


2. A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design


3. MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning


4. AutoMine Solution for AV2 2026 Scenario Mining Challenge


5. Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task


6. SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning


7. Mind the Perspective: Let’s Reason Recursively for Theory of Mind


8. Organize then Retrieve: Hierarchical Memory Navigation for Efficient Agents


9. Lung-R1: A Knowledge Graph-Guided LLM for Pulmonary Diagnostic Reasoning


10. Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning


11. SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior


12. INFRAMIND: Infrastructure-Aware Multi-Agent Orchestration


13. Automated Mediator for Human Negotiation: Pre-Mediation via a Structured LLM Pipeline


14. Position: Hippocampal Explicit Memory Is the Cornerstone for AGI


15. From Explicit Elements to Implicit Intent: A Predefined Library for Auditable Behavioral Inference


16. Reroute, Don’t Remove: Recoverable Visual Token Routing for Vision-Language Models


17. DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?


18. System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5


19. TAHOE: Text-to-SQL with Automated Hint Optimization from Experience


20. APPO: Agentic Procedural Policy Optimization


21. ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing


22. Harness In-Context Operator Learning with Chain of Operators


23. Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition


24. VIA-SD: Verification via Intra-Model Routing for Speculative Decoding


25. Rule Taxonomy and Evolution in AI IDEs: A Mining and Survey Study


26. Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application


27. OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models


28. Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation


29. Augmenting Molecular Language Models with Local $n$-gram Memory


30. MSUE: Multi-Modal Soccer Understanding Expert


31. “That’s AI Slop, You Bot!” Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments


32. On the Limits of LLM-as-Judge for Scientific Novelty Assessment


33. Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding


34. Exploration Structure in LLM Agents for Multi-File Change Localization


35. Categorical Prior Lock-in: Why In-Context Learning Fails for Structured Data


36. Characterizing Software Aging in GPU-Based LLM Serving Systems


37. Beyond representational alignment with brain-guided language models for robust reasoning


38. Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection


39. Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Production


40. Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training


41. Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning


42. Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code


43. WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning


44. Automated Creativity Evaluation of Language Models Across Open-Ended Tasks


45. From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal Reasoning


46. Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild


47. ICA Lens: Interpreting Language Models Without Training Another Dictionary


48. Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning


49. Goal-Autopilot: A Verifiable Anti-Fabrication Firewall for Unattended Long-Horizon Agents


50. Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness


51. Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment


52. Runtime Skill Audit: Targeted Runtime Probing for Agent Skill Security


53. TAROT: Task-Adaptive Refinement of LLM-prior Graphs for Few-shot Tabular Learning


54. Physics-Distilled Neural Network enabled by Large Language Models for Manufacturing Process-Property Predictive Modeling


55. AVIS: Adaptive Test-Time Scaling for Vision-Language Models


56. LLMs+Graphs: Toward Graph-Native, Synergistic AI Systems


57. When Roleplaying, Do Models Believe What They Say?


58. Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality


59. APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection


60. AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable


61. The Power of Test-Time Training for Approximate Sampling


62. JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization


63. MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation


64. Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models


65. Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models


66. Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering


67. When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis


68. The Dynamics of Human and AI-Generated Language: How Semantics Fluctuates across Different Timescales


69. TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs


70. FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse


71. RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways


72. Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation


73. PermDoRA – Understanding Adapter Interference in Language Models: Limits of Parameter-Space Geometry


74. RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark


75. SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving


76. Artificial Intelligence in Ship Finance: Applications, Opportunities, and a Case Study in AI-Augmented Loan Origination


77. Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs


78. Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents


79. An Ethical eValuation Agent (EeVA): Results of a Proof-of-Concept Test on a Prototype Agentic-like Workflow to Assist Ethical Deliberations


80. Preregistration for Experiments with AI Agents


81. The Environmental Cost of LLMs in AIED: Reporting and Practices


82. Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models


83. T2MM: An LLM Supported Architecture For Inquiry-Based Modeling


84. ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward


85. Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention


86. NightFeats @ MMU-RAGent NeurIPS 2025: A Context-Optimized Multi-Agent RAG System for the Text-to-Text Track


87. The Structural Attention Tax: How Retrieval Format Hijacks In-Context Learning Independent of Content


88. PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference


89. From Consumption to Reflection: Designing Human-AI Relations for Stable Reasoning


90. From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data