전체 AI 논문 - 2026-08-13

1. Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models


2. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies


3. An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS


4. How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models


5. Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation


6. GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings


7. Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges


8. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence


9. CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations


10. Claim-Level Reliability Assessment for Efficient Test-Time Reasoning


11. Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection


12. ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models


13. OEIS Open: How many conjectures can language models turn into theorems?


14. Policy-as-logic for robust reasoning over rules


15. Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents


16. The Sleeping Agent: What Gist-Based Context Compression Loses and Why


17. HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry


18. Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents


19. Proportional Analogies on Probability Distributions via Bayesian Updating


20. Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning


21. HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting


22. FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents


23. AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection


24. XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication


25. CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement


26. Making AI-Generated Feedback Matter: From Provision to Student Enactment


27. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation


28. Foresight Without Seeing: Latent Futures for World Action Models


29. Learning from Online User Feedback for Shopping Agents


30. CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications


31. EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval


32. Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models


33. From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation


34. A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization


35. Benchmarking LLM Judges for Mobile Agent Evaluation


36. Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology


37. When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs


38. From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate


39. Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces


40. Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval


41. Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence


42. Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations


43. Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning


44. Adaptive Hybrid Particle Swarm Optimization with Gradient Descent


45. Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures


46. Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning



48. EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents


49. Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier


50. Towards the Harness of Embodied Agents


51. Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach


52. BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model


53. The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification


54. RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle


55. VQ-bench: A Composable Vector Quantization Framework


56. Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability


57. Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction


58. CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference


59. InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk


60. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs


61. The Edge-based Contiguous p-median Problem with Connections to Logistics Districting


62. Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)


63. Forecasting Side Effects of Activation Steering


64. Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet


65. Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones


66. Harnessing agent memory to build lifelong AI partners for materials scientists


67. A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems


68. LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs


69. From Monolithic to Modular: Segment-level Automatic Prompt Optimization


70. MaSRead: Content-Addressed Reading of Replicated Latent Stores


71. AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research


72. Poor Man’s Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop


73. Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts


74. A Forced-Structure Reduction and Verifiable Bounds for Conway’s 99-Graph


75. Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration


76. Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes


77. DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation


78. AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses


79. Redistribution-based Cost Inference Improves Sparse Safe Offline RL


80. Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations


81. Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence


82. Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages


83. A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery


84. Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents


85. Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams


86. One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL


87. VICBench: A Multi-Language Benchmark for Code Vulnerability Detection


88. HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression


89. How Organizations Use AI: Evidence from ChatGPT


90. Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images


91. Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification


92. SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward


93. Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge


94. Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment


95. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation


96. M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation


97. HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks


98. Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation


99. HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation


100. Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers


101. A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench


102. Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning


103. Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation


104. Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control


105. QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving


106. No One to Blame: A Framework of Constitutive AI Unaccountability


107. Confidence Calibration of Deep Learning Systems


108. Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion


109. Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models


110. Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL


111. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations


112. How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging


113. LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration


114. Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision


115. From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices


116. Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches


117. RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks


118. Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh


119. HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs


120. LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation


121. Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference


122. TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement


123. Accuracy and Order Sensitivity Diverge Under Label-Free Strategies


124. Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians


125. Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models


126. Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework


127. DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation


128. CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation


129. Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra


130. LookBack: Where and How to Score LVLM Responses via Visual Reference Usage


131. User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling


132. How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment


133. Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring


134. Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion


135. TELLME: Test-Enhanced Learning for Language Model Enrichment


136. GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation


137. Instruction Alignment for Binary Code Representation Learning


138. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning


139. JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis


140. G0.5: One Autoregressive Stream for Robot Reasoning and Action


141. Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System


142. Locating and Controlling Implicit Personalization in Large Language Models


143. A 12-CNOT Double Qubit Excitation Gate


144. Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation


145. When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use


146. High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions


147. Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing


148. Consolidator: Learning Persistent Routed Memory Across Context Boundaries


149. REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation


150. The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance


151. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference


152. Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation


153. GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs


154. Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL


155. Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads


156. Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing


157. Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning


158. Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models


159. Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting


160. Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents


161. Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones


162. Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs


163. FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting


164. Dion3: Full-Stack Orthogonal Updates


165. A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases


166. RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation


167. Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs


168. From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection


169. Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents


170. A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era


171. Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment


172. Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks


173. Keep the Future, Drop the Rollout: RIFT for World Action Models


174. Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation


175. Let it Cook: Learning to Wait in Sequential Decision Making


176. Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task


177. Strengthening Full Justified Representation: Efficient Verification and Computation


178. HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging


179. The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark


180. PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition


181. TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation


182. Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards


183. TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs


184. AI Guardrail Survival under Single-Cycle Agentic Self-Summarization


185. Gaze Target Estimation Anywhere with Concepts


186. Dynamics Models for Offline Hyperparameter Selection in Real-World RL


187. Governing Agentic AI in FinTech


188. Self-evolving network verifiers


189. Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation


190. Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings


191. Socioduality: A Relational Process Framework for Human-AI Interaction


192. Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction


193. Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning


194. Backdoor Decontamination Dynamics in LLM Agents


195. CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification


196. SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation


197. Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models


198. Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification


199. Federated Learning for Distributed CNC Tool Wear Prediction


200. Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification


201. Every pooling rule has its world: matching probability combination rules to situations and stakes


202. Agent Safety Should Be a Runtime Contract


203. Methodologies for Improving the Quality of AI Tutoring in K-12 Education


204. Variable Selection in the Context of AI Fairness


205. Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression


206. Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction


207. Reinforcement Learning based DBMS Buffer Pool Auto-Tuning for Optimal Memory Utilization


208. TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation


209. Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets


210. Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs


211. Evaluating LLM Generated Detection Rules in Cybersecurity