LLM 관련 주요 논문 - 2026-06-24

1. OpenThoughts-Agent: Data Recipes for Agentic Models


2. Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models


3. Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System


4. Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment


5. Can Scale Save Us From Plasticity Loss in Large Language Models?


6. Scaling Laws for Task-Specific LLM Distillation


7. LaGO: Latent Action Guidance for Online Reinforcement Learning


8. CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning


9. SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation


10. When CQs Go Wrong: Challenges in CQ Verification with OE-Assist


11. ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling


12. AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability



14. Governed Shared Memory for Multi-Agent LLM Systems


15. Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation


16. A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial


17. On the Smallness of the Large Language Models Scaling Exponents


18. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference


19. Bayesian control for coding agents


20. ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling


21. Cycle-Consistent Neural Explanation of Formal Verification Certificates


22. Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War


23. PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models


24. When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs


25. Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation


26. LemonHarness Technical Report


27. Probing the Misaligned Thinking Process of Language Models


28. T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph


29. VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification


30. Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning


31. Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?


32. Critique of Agent Model


33. RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems


34. IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation


35. Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution


36. EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence


37. Grad Detect: Gradient-Based Hallucination Detection in LLMs


38. DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects


39. UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving


40. Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations


41. FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction


42. AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach


43. Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity


44. Toward Self-Evolution-Ready Workflow Harnesses: A Reversible Migration Path and Convertibility Taxonomy for Expert LLM Pipelines


45. Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams


46. CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation


47. Red-Teaming the Agentic Red-Team


48. video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding


49. The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs


50. On the Stability of Prompt Ranking in Large Language Model Evaluation


51. CALIBER: Calibrating Confidence Before and After Reasoning in Language Models


52. Pigeonholing: Bad prompts hurt models to collapse and make mistakes


53. SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization


54. Social Structure Matters in 3D Human-Human Interaction Generation


55. AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming


56. Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy


57. A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy


58. CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression


59. Towards Version-aware Operations and Transaction Memories for Multi-layer MeMo


60. Fast and Slow Variational Continual Learning


61. Towards Spec Learning: Inference-Time Alignment from Preference Pairs


62. RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring


63. Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization


64. Maestro Order: A Model-Agnostic Orchestration Harness


65. The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models


66. E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis


67. Mind the Heads: Topological Representation Alignment for Multimodal LLMs


68. One Year Later…The Harms Persist, But So Do We!


69. JupOtter: Cell-Level Bug Detection in Jupyter Notebooks


70. From Task-Guided Conversational Graphs to Goal-Oriented Dialogue Runtimes


71. Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification


72. VeriPilot: An LLM-Powered Verilog Debugging Framework



74. Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability


75. Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment


76. SemChunk-C: Semantic Segmentation for C Code


77. Quantifying Prior Dominance in RAG Systems


78. Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code