전체 AI 논문 - 2026-08-20

1. Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication


2. Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems


3. Tuning the Stochastic Machine: A Systems Engineer’s Operating Model for Human-AI Engineering


4. Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk


5. What is Missing from AI Post-Training AI: An Empirical Analysis


6. Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery


7. Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering


8. Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models


9. A Theory of Post-hoc Debate Judgement



11. \textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems


12. Syntactic Simplification of OWL Class Expressions


13. Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models


14. DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning


15. SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents


16. ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery


17. Verifiable abstention makes AI leak diagnosis accountable in water distribution networks



19. Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots


20. A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation


21. Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization


22. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training


23. Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction



25. Preference Reasoning under Indeterminacy in Large Language Models


26. CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence


27. Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference


28. FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis


29. Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement


30. FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems


31. Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson


32. Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval


33. UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval


34. FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents


35. Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions


36. When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification


37. A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations


38. Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam’s 2025 Convex Marking Scheme


39. Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair


40. ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents


41. SESSE: Sketch, Expand, Sort, Summarize, Evaluate – LLM-as-Judge Evaluation via Structured Decomposition


42. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations


43. Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application


44. Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study


45. Redakto - The Incognito Tab for LLMs


46. GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks


47. On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices


48. Looped Language Models Improve Compositional Tool Calling


49. Adversarial Review: Structured Disagreement for Grounded Agentic Code Review


50. RDFdL: Integrating RDF with Differential Dynamic Logic


51. Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu


52. FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud


53. Improving Rural Medication Safety with AI: A Scoping Review


54. Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis


55. Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs


56. Position: AI Leaderboards Are Underserving the Global South: A Case Study from India


57. Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry


58. Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions


59. Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective


60. FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management


61. Position: Multi-Agent Systems Should Prioritize Concurrency Control


62. A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health Monitoring


63. Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models


64. Position: Behavioral Systems Require Behavioral Tests


65. Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges


66. Position: Profiling Game Worlds by Transition Complexity


67. Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions


68. SPADE: Self-Play in Adaptive Synthetic Executable Environments


69. ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning


70. Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning


71. Finetuning Strategies for Querying Sounds by Vocal Imitation


72. Interpretable AI predicts a 2026 summer dry anomaly in central China


73. Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets


74. Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles


75. Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers


76. PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints


77. Discretizing Continuous Time Series for Imputation with Masked Diffusion Training


78. Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation


79. Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift


80. DA-WAM: Decision-Aligned Future Latents for Driving World Models


81. ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models


82. GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting


83. Bernstein-Vazirani Networks: Quantum Machine Learning by Interference


84. Counterfactual Contrastive Analysis


85. One-Stage Object Detectors in Autonomous Driving


86. Harness Continual Learning: Continual Adaptation Beyond Model Parameters


87. From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation


88. GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery


89. DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering


90. rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation


91. AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL


92. Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis


93. MedUAG: Unified Understanding and Generation for Medical Multimodal Models


94. Graphical Design of Interpretable Architectures


95. SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution


96. Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck


97. SMTrap: Cost-Effective DoS Attacks Against Large Reasoning Models via SMT Conflict Guidance


98. Learning-State-Aware Dynamic Generative Data Augmentation on Small-Scale Datasets


99. MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models


100. Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis


101. Identifying Implicit Premises for Logical Reconstruction of Argument Graphs


102. Do Large Language Models Hallucinate Electric Fata Morganas?


103. A strengthening of the MCFL-ness of $O_2$


104. Forgetting, plasticity, and co-observation: a third facet of continual learning


105. Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study


106. SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation


107. Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening


108. Epistemic Subordination: Generative AI and the Infrastructure of Knowledge


109. Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services


110. A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3


111. Budget-First Tariff Recommendation (BFTR): A Complete Algorithmic Framework for Telecom Plan Recommendation without Overcharging


112. The Impact of CutMix on Reliability and Robustness in Semantic Segmentation


113. A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation


114. MemFuse: Multi-Source Memory Fusion from Fragmented Observations


115. Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts


116. Composed Historical Image Retrieval by Modeling Temporal Representations


117. Europe’s Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections


118. Aslema at NADI 2026: Augmentation through Fewshot for SLU


119. Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics


120. Change Point–Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation


121. Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings


122. OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios


123. From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning


124. MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment


125. The Role of Grid Cells in Reducing Spatial Aliasing in Hippocampal Place Representations


126. MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence


127. Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments


128. CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks


129. Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions


130. DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents


131. Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection


132. GCNO: Gramian Chebyshev Neural Operator for Physics-Based Compression of Wireless Channels


133. OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation


134. Science Done on a Machine by a Machine: AI Agents in Computational Chemistry


135. Physics-Unrolled Neural Operator for Wireless Field Modeling


136. Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models


137. Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement


138. ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems


139. Formal Verification of Romanov’s Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas


140. Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage


141. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B


142. Vector Symbolic Policy Gradient


143. LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents


144. TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs


145. Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation


146. One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI


147. Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents


148. Coupled-cluster molecular properties across the main group that extrapolate beyond training size


149. Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring


150. From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model


151. FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning


152. FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation


153. Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements


154. What Makes Software Issue Resolution Tasks Difficult for Agents?


155. SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents


156. How AI Prompts Can Teach Us About the Structure of Human Behavior


157. Visual-Prompt Guided Wildlife Instance-Level Recognition


158. Bidirectional representational alignment between biological and artificial neural networks


159. GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction


160. Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift


161. A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities


162. What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems


163. Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation


164. When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators


165. How Quantum Is the Advantage? A Fair, Calibration- and Noise-Aware Benchmark and Attribution Audit of Quantum Machine Learning for Network Intrusion Detection


166. TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving


167. Entropy-Constrained Adaptive Stochastic Quantization


168. The Deontic Gap: Large Language Models and the Modal Language of Obligation


169. Language Models for Portuguese: A Systematic Mapping Study


170. Global Index on Responsible AI 2026 : Conceptual Framework and Methodology


171. Temporal Multi-Signal Fusion for Token-Level Hallucination Detection


172. Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings


173. Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation


174. Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals


175. Different Facets of Verbalised Overconfidence: an Interpretability Study


176. StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data


177. DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models


178. Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)


179. Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems


180. Backdoor Learning in Language Models and Vision-Language Models


181. NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages


182. Abliteration Mitigation via Refusal Aliases


183. Self- and Other-Labels Induce Bidirectional Bias in LLM Judges


184. Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities


185. Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining


186. SuTRA : Structurally-Unified Tokenization with Root Awareness