LLM 관련 주요 논문 - 2026-06-08

1. Online Pandora’s Box for Contextual LLM Cascading


2. Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation


3. Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization


4. Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning


5. Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces


6. Declarative Skills for AI Agents in Knowledge-Grounded Tool-Use Workflows


7. Workflow-to-Skill: Skill Creation via Routing-Workflow-Semantics-Attachments Decomposition


8. Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation


9. AdMem: Advanced Memory for Task-solving Agents


10. OpenSkill: Open-World Self-Evolution for LLM Agents


11. A Geometric Account of Activation Steering through Angle-Norm Decomposition


12. CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions


13. Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory


14. How reliable are LLMs when it comes to playing dice?


15. MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism


16. Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning


17. TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment


18. Watch, Remember, Reason: Human-View Video Understanding with MLLMs


19. The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs


20. Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills


21. A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning


22. Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration


23. SV-Detect: AI-generated Text Detection with Steering Vectors


24. Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition


25. When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations


26. DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios


27. Textual Supervision Enhances Geospatial Representations in Vision-Language Models


28. REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference


29. The Three-Ring Architecture: Governing Agents in the Era of On-Platform Organisations


30. OffQ: Taming Structured Outliers in LLM Quantization by Offsetting


31. GP-Adapter: Gaussian Process CLIP-Adapter for Few-Shot Out-of-Distribution Detection


32. On the Geometry of On-Policy Distillation


33. TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents


34. Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets


35. DataEvolver: Automatic Data Preparation for Large Language Models through Multi-Level Self-Evolving


36. Don’t Pause: Streaming Video-Language Synchrony for Online Video Understanding


37. OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios


38. Auditing Training Data in Domain-adapted LLMs: LoRA-MINT


39. SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models


40. The Fine-Tuning Trap: Evaluating Negative Transfer and the Role of PEFT in Sub-1B Mathematical Reasoning


41. ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning


42. SpectCount: Spectrotemporal Counting via Synthetic Signals Improves Large Audio Language Models


43. EASE-TTT: Evidence-Aligned Selective Test-Time Training for Long-Context Question Answering


44. MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models


45. LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics


46. Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation


47. Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks


48. Progress-SQL: Improving Reinforcement Learning for Text-to-SQL via Progressive Rewards


49. PandaAI: A Practical Agent CQ2 for Neuro-symbolic Data Analysis And Integrated Decision-Making in Quantitative Finance


50. SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling


51. Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations


52. Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection


53. HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec


54. Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles


55. Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation


56. MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models



58. How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures


59. Re-Centering Humans in LLM Personalization


60. NTILC: Neural Tool Invocation via Learned Compression


61. AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps


62. IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems


63. FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models


64. Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration


65. Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems


66. Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?


67. Autonomous heterogeneous catalyst discovery with a self-evolving multi-agent digital twin


68. Trading Engagement for Sustainability: Carbon-Aware Re-ranking for E-commerce Recommendations


69. When Does Multi-Agent Collaboration Help? An Entropy Perspective


70. Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs