LLM 관련 주요 논문 - 2026-09-24

1. StudentBench: AI and human tutoring yield equivalent GRE learning gains


2. An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act’s Code of Practice


3. Learning the Cost of Reliable Inference


4. PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety


5. Agentic Governance and Adversarial Verification for Policy-Constrained LLM Healthcare Appeal Generation


6. Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM


7. Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis


8. Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions


9. State-Grounded Conditioning: Wrapping User-Facing LLM Agents Where Direction Depends on Live State


10. Not What You Meant: Can LLMs Follow a Specified Negation Semantics?


11. MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design


12. CART: Closed-Loop Adaptive Red Teaming for Large Language Models


13. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents


14. StateComp: Learning When to Compress History in Long Horizon Agents


15. Memory Control Signals Emerge Before Action in Long Horizon Agents


16. Hunyuan-A13B Technical Report


17. Enhancing Small Language Models for Power Outage Report Generation via Minimum Risk Training


18. Do We Need Complex Topology Control? Distinct-Peer Random Routing Improves Cost-Efficiency in Sparse Multi-Agent Debate


19. Provably Complete Generalized Planning with LLMs


20. Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit


21. Propose, Don’t Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors


22. Math Reasoning in LLMs is Organized by Approach, Not Topic


23. Reinforcement Learning with Decomposed Subtasks


24. Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts


25. Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution


26. Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity


27. Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse


28. Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark


29. Order-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning


30. Agent-Editing World Model: Rethinking World Modeling for LLM Agents


31. Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer


32. AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios


33. Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers


34. Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?


35. Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings


36. Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing


37. Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices


38. RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction


39. Query Implied Generative Engine Optimization


40. Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis


41. EidosDoc: Implicit Structure Encoding for Cost-Effective Semi-Structured Document QA


42. Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures


43. Hard Negatives Reveal What Easy Negatives Hide: Cross-Lingual Harmfulness Representations Degrade with Resource Tier Under Hard Negatives


44. The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA


45. FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation


46. InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation


47. When Context Misleads: In-context Learning with Jurisdiction in Large Language Models


48. Hidden not Deleted: How Networks Suppress Entangled Features


49. DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment


50. FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration


51. NV-Reason-CT: 3D Visual Language Model for CT Analysis


52. Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models


53. What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit


54. Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools


55. Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction


56. Evolving Inspectable O-RAN Slicing xApps with LLMs


57. Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation


58. KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling


59. Combining LLMs and Genetic Search for ARC-AGI-2


60. Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices


61. Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement


62. The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems


63. A Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery


64. Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving


65. EMA: Elastic and Performance Transparent Memory Across GPUs


66. An open benchmark for machine learning-based polymer property prediction


67. Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms


68. Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation


69. COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference


70. Ajar: Measuring Open Privilege in Agent Defenses


71. COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation


72. Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits


73. Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management