LLM 관련 주요 논문 - 2026-07-22

1. Agents in the Wild: Where Research Meets Deployment


2. LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior


3. Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks


4. OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation


5. Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs


6. Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio


7. Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning


8. Measuring Reward-Seeking via Contrastive Belief Updates


9. From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar


10. OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining


11. PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents


12. Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety


13. AI Tour Meeting: Group Travel Planning by LLM Agents


14. SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval


15. AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents


16. Semantic Primes as Explanans for Emotion in Large Language Models


17. SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring


18. Operational Hallucination and Safety Drift in AI Agents


19. Using LLMs for Explainable, Data-Driven Insight Generation from Time Series


20. Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles


21. Fence: Specialized SLM Guardrails for LLM Applications


22. Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models


23. State Compression in Two-Agent LLM Relays: A Closed-World Study of Constraint Preservation


24. MUX: Continuous Reasoning via Multiplexed Tokens


25. When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents


26. FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis


27. Probabilistic Concept-Aware Steering for Trustworthy LLM Inference


28. PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language


29. Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems


30. Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR


31. Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads


32. MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers


33. Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance


34. BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data


35. Calibrated Selective Fact-Checking via Evidence Chain Evaluation


36. SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI


37. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning


38. ISO: An RLVR-Native Optimization Stack


39. Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information


40. They’ll Verify. They Just Won’t Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface


41. Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation


42. PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image


43. Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks


44. Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models


45. Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs


46. The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation


47. Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards


48. Assessment in Team Problem-Solving Exercises in Computing Education



50. SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation


51. From Operations to Elderly Care Outcomes: A Thematic Review of Industrial Engineering and Decision-Support Approaches


52. Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training


53. AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism


54. SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement


55. Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model


56. OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation


57. Data Leakage Prevention in Agentic Applications via Preemptive Hardening


58. Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents



60. AgentTrails: Towards Trust and Reuse for Agentic Tasks


61. Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks


62. Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA


63. Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models


64. Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents


65. CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization


66. LatentMT: Machine Translation with Latent Reasoning


67. Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach


68. For What Reason? Interpreting Models’ Encoding of Causation and Antithesis


69. The Story Shapes the Agent: Narrative Priors in LLM Behavior


70. Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary


71. EduPanel: A Three-Agent LLM Judge for Teaching Videos – Reliability, Complementarity, and Human Trust Calibration


72. Towards an Automated Test of LLM Security Knowledge


73. Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing


74. Structured Output Collapses Answer Diversity Across 44 Language Models


75. Competitive and Complementary Tools


76. Estimating Rare Events in Language Models with Proper Evaluation


77. ChainMark: Model-Free LLM Watermarking with Closed-Form Calibration


78. HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers


79. Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments


80. Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies


81. Binding Drift in Multi-Step Tool-Augmented Agents


82. Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling


83. Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative


84. The Information Shadow: Measuring Structural Limits on What Language Models Can Learn


85. Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime


86. Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression


87. Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models


88. The Economics of Autonomy: Real-Time Risk Indexing for Insurable AI-Driven 6G Systems


89. MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models