LLM 관련 주요 논문 - 2026-09-17

1. MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education


2. Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization


3. Function Lives Where Variance Doesn’t: Task-Weighted Charts of a Language Model’s Computation


4. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data


5. CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents


6. Clueing up LLMs with Tool-Augmented Deductive Reasoning


7. Which LLM is Best for Translating Natural Language Goals to PDDL


8. Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning


9. Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection


10. Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making


11. AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution


12. Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models


13. Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery


14. The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models


15. Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland


16. Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents


17. Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition


18. What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models


19. Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum


20. BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs


21. REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement


22. WFM: Wiki Foundation Model for Complex Agentic Reasoning


23. Symbolic Temporal Supervision of LLM Agents Using Contracts


24. AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines


25. When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation


26. Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning


27. Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation


28. When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI


29. Collaborative Memory for Multi-Agent VLM Systems


30. The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?


31. A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning


32. SAGE: Governed Artifact Generation from Enterprise Guidelines


33. GVD: Governed Versioning and Deduplication for Document Repositories


34. GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents


35. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video


36. EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents


37. Objective vs. Search: Decomposing What Makes a Good Tokeniser


38. A Zeroth-Order Paradigm for LLM Preference Alignment


39. MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents


40. TalkMatrix: Generating Character Dialogue that is Both Consistent and Diverse


41. WordPolo: Evaluating Language Models Through Iterative Semantic Feedback


42. BadQubits: An LLM-Based Framework for Static Pre-Execution Detection of Structurally Harmful Quantum Circuits


43. StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction


44. Higher-order pruning of experts in mixture-of-experts language models


45. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking


46. ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions


47. Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection


48. A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages


49. Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening


50. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?


51. Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents


52. VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval


53. GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models


54. Autonomy in Check: Governor-Mediated Adaptive Security at the Edge


55. Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR



57. I code or AI code: A comparative evaluation of AI-rated scores in classroom observations


58. ${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models


59. Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers


60. Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment


61. From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale


62. An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks


63. Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders


64. EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing


65. Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming


66. Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels


67. Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents


68. Information Set Emulation: Causal Certificates for AI Derived EHR Features


69. HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models


70. One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG


71. Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents


72. The Missing “I Don’t Know”: Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention


73. Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models


74. Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models


75. Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds


76. The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models