LLM 관련 주요 논문 - 2026-08-17

1. Handover of In-Context Learning State Across Session Boundaries


2. Split the Labor: Separating Evidence Interpretation from Decision Aggregation


3. SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning


4. Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations


5. LLMs Don’t Pay for the Jump


6. Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons


7. AgentRewind: Recoverable Execution for Long-Horizon LLM Agents


8. ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond


9. Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents


10. AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs


11. TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments


12. Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models


13. APTER: Adaptive Post-Training with Expert-Grounded Rubrics


14. Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding


15. BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs


16. Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach


17. A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents


18. Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers


19. Scaling Domain Data Repetition in LLM Pretraining


20. Demystifying Agent Skills: Why They Work-Until They Don’t


21. Agent-Orchestration in Autonomous Chip Design


22. Buy the Rumor, Sell the News: When Is News Priced In?


23. AI Research Preference Models


24. Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact


25. When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict


26. From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL


27. Simulation-Aware In-Context Policy Improvement for LLM-Aided Analog Layout Refinement


28. Second Thought: Reasoning in Parallel as LLM Agents Act and Observe


29. How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights


30. SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data


31. Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis


32. No Universal Signal Predicts Sample-Level LLM Regression under Version Updates


33. Active Perception for Embodied Disambiguation


34. Measuring Cross-Task Behavioral Consistency in Language Model Agents


35. Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors


36. Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents


37. A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing


38. Modular Cognitive Architecture Emerges in Large Language Models


39. Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking


40. Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation


41. Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice


42. DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding


43. A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models


44. A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation


45. Seeing Red, Thinking Bad: Color Bias in Vision Language Models


46. Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions


47. P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems


48. MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation


49. Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency


50. Content Based Video Narration of Gameplay with Vision Language Models


51. MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning


52. Musical Mirrors: The LLM as Sounding Board in Songwriting


53. CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing


54. Agentic Transaction: Towards ACID-Compliant Agent Systems


55. Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions


56. Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions


57. Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models


58. Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline


59. MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation


60. Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT


61. Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation


62. Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning


63. Context Aware AI Assistant and AR Interface for Lunar Extravehicular Activity (EVA) Procedural Guidance


64. From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent


65. Jais 2: A Family of Arabic-Centric Open Large Language Models


66. Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems


67. Think in Latent, Explain in Language: Self-Explainable Latent Reasoning


68. Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support