LLM 관련 주요 논문 - 2026-05-27

1. MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation


2. Natural Language Query to Configuration for Retrieval Agents


3. Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases


4. Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering


5. Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments


6. The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?


7. VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions


8. ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules


9. Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation


10. Position: AI Safety Requires Effective Controllability


11. Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation


12. Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry


13. Generating Robust Portfolios of Optimization Models using Large Language Models


14. LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation


15. Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)



17. TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews


18. Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation


19. Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs


20. The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context


21. A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks


22. It’s Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers


23. Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation


24. MemFail: Stress-Testing Failure Modes of LLM Memory Systems


25. UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems


26. FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning


27. AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents


28. MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning


29. MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration


30. The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence


31. Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions


32. From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator


33. Automatic Layer Selection for Hallucination Detection


34. Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning


35. OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling


36. Experiments in Agentic AI for Science


37. Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions


38. Can LLMs Introspect? A Reality Check


39. LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding


40. GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesis, Research, and Testing


41. MobileMoE: Scaling On-Device Mixture of Experts


42. Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders


43. EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering


44. It’s Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty


45. Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)


46. Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs


47. Qiskit QuantumKatas: Adapting Microsoft’s Quantum Computing exercises for LLM evaluation


48. Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis


49. Learning When to Think While Listening in Large Audio-Language Models


50. LitSeg: Narrative-Aware Document Segmentation for Literary RAG



52. ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference


53. E3: Issue-Level Backtesting for Automated Research Critique


54. QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents


55. ConVer: Using Contracts and Loop Invariant Synthesis for Scalable Formal Software Verification


56. Tracing Computation Density in LLMs


57. Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination


58. ReasonOps: A Unified Operational Paradigm for Trustworthy Verified LLM Reasoning


59. Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling


60. Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation


61. JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors


62. Beyond Questions: Evaluating What Large Language Models (Actually) Know


63. Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton


64. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models


65. GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought


66. Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations


67. The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection


68. Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study


69. The Kalman Evolve: Closing the Gap in Kalman Filtering via Interpretable Algorithm Discovery


70. ContextGuard: Structured Self-Auditing for Context Learning in Language Models


71. RAGEAR: Retrieval-Augmented Graph-Enhanced Academic Recommender


72. Innovation: An Almost Characterization of Hallucination


73. SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability


74. EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation


75. Ratio-Variance Regularized Policy Optimization


76. MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation


77. Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models


78. L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation


79. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning


80. An In-Vitro Study on Cross-Lingual Generalization in Language Models


81. DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding


82. The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models


83. AI evaluation may bias perceptions: The importance of context in interpreting academic writing


84. Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models


85. More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations


86. Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training


87. Cordyceps: Covert Control Attacks on LLMs via Data Poisoning


88. Linear and Neural Dueling Bandits with Delayed Feedback


89. A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection


90. InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward


91. Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization


92. Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories


93. Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records


94. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization


95. Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines


96. LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness


97. Targeted Remasking: Replacing Token Editing with Token-to-Mask Refinement in Discrete Diffusion Language Models


98. The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP


99. Plans for Evaluating Structured Generative Search Summaries


100. Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection


101. VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes


102. BioFact-MoE: Biologically Factorized Mixture of Experts for Vision-Language Prognostic Modeling in Hepatocellular Carcinoma


103. Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations


104. Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models


105. Curriculum Learning for Safety Alignment


106. Intelligent Detection and Mitigation of Carpet-Bombing DDoS Attacks in SDN Using Retrieval-Augmented Generation and Large Language Models


107. CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations


108. AgentSociety: Incentivizing Agentic Social Intelligence


109. CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly


110. Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training


111. SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?


112. AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations


113. PitchBench: Measuring Pitch Hearing in Audio-Language Models


114. InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization


115. A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration


116. Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets


117. TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models


118. Furina: Fragmented Uncertainty-Driven Refusal Instability Attack


119. Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges


120. MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning


121. VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents


122. Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception


123. Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications


124. GEM: Geometric Entropy Mixing for Optimal LLM Data Curation


125. Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU