LLM 관련 주요 논문 - 2026-07-16

1. Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models


2. AIMO Interpretability Challenge


3. Experience Memory Graph: One-Shot Error Correction for Agents


4. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities


5. STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle


6. Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System


7. SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing


8. How Far Can Root Cause Analysis Go on Real-World Telemetry Data?


9. LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning


10. Set-shifting Behavioral Test for Harnessed Agents


11. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable


12. Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management


13. Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools


14. Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution


15. Music-to-Dance Generation via Atomic Movements


16. Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code


17. Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings


18. Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild


19. Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection


20. Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations


21. Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs


22. Social Simulations: from Agent-Based Modeling to Digital Twins


23. Consensus as Privileged Context for Label-Free Self-Distillation


24. Semantic Anchoring for Robotic Action Representations


25. Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities


26. Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents


27. GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning


28. UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors


29. ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level


30. DeepLoop: Depth Scaling for Looped Transformers


31. DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments


32. GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding


33. Data-Efficient Adaptation of LLMs via Attention Head Reweighting


34. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models


35. Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition


36. Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks


37. Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance


38. Reassessing Muon for Matrix Factorization


39. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference


40. RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar


41. What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors


42. SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests


43. WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency


44. Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit


45. Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates


46. Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models


47. Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework


48. Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs


49. Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows


50. The Hitchhiker’s Guide to Monoculture


51. HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models


52. Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes


53. Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems


54. The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI


55. LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents


56. Designing Safety-Constrained LLM Systems for Public Health Information Access


57. Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry


58. FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents