LLM 관련 주요 논문 - 2026-08-20

1. Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems


2. Tuning the Stochastic Machine: A Systems Engineer’s Operating Model for Human-AI Engineering


3. What is Missing from AI Post-Training AI: An Empirical Analysis


4. Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models


5. A Theory of Post-hoc Debate Judgement



7. Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models


8. DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning


9. ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery


10. Verifiable abstention makes AI leak diagnosis accountable in water distribution networks


11. Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots


12. A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation


13. Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization


14. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training


15. Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction


16. Preference Reasoning under Indeterminacy in Large Language Models


17. CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence


18. Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference


19. FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems


20. Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson


21. Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval


22. FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents


23. Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions


24. Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam’s 2025 Convex Marking Scheme


25. SESSE: Sketch, Expand, Sort, Summarize, Evaluate – LLM-as-Judge Evaluation via Structured Decomposition


26. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations


27. Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application


28. Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study


29. Redakto - The Incognito Tab for LLMs


30. Looped Language Models Improve Compositional Tool Calling


31. Adversarial Review: Structured Disagreement for Grounded Agentic Code Review


32. Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu


33. Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs


34. Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective


35. FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management


36. Position: Multi-Agent Systems Should Prioritize Concurrency Control


37. Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges


38. Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions


39. SPADE: Self-Play in Adaptive Synthetic Executable Environments


40. Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets


41. Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers


42. ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models


43. From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation


44. rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation


45. Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis


46. MedUAG: Unified Understanding and Generation for Medical Multimodal Models


47. Graphical Design of Interpretable Architectures


48. SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution


49. Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck


50. MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models


51. Identifying Implicit Premises for Logical Reconstruction of Argument Graphs


52. Do Large Language Models Hallucinate Electric Fata Morganas?


53. Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study


54. Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services


55. MemFuse: Multi-Source Memory Fusion from Fragmented Observations


56. Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts


57. Aslema at NADI 2026: Augmentation through Fewshot for SLU


58. OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios


59. From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning


60. MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment


61. CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks


62. Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions


63. DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents


64. Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement


65. Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage


66. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B


67. LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents


68. TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs


69. From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model


70. Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements


71. SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents


72. How AI Prompts Can Teach Us About the Structure of Human Behavior


73. Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation


74. When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators


75. TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving


76. The Deontic Gap: Large Language Models and the Modal Language of Obligation


77. Language Models for Portuguese: A Systematic Mapping Study


78. Temporal Multi-Signal Fusion for Token-Level Hallucination Detection


79. Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings


80. Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation


81. Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals


82. Different Facets of Verbalised Overconfidence: an Interpretability Study


83. StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data


84. DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models


85. Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)


86. Backdoor Learning in Language Models and Vision-Language Models


87. NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages


88. Abliteration Mitigation via Refusal Aliases


89. Self- and Other-Labels Induce Bidirectional Bias in LLM Judges


90. Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities