전체 AI 논문 - 2026-08-19

1. On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification


2. Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating


3. HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance


4. StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents


5. Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach


6. Towards Zero-Shot Task Transfer with Neurosymbolic World Models


7. Procedural Content Metageneration via Program Search and Continual Abstraction Discovery


8. EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection


9. Adaptive Policy Portfolios for Robust Markov Decision Processes


10. AutoResearch: Insight In, Hallucination Out


11. ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction


12. StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows


13. D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory


14. The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting


15. Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits


16. Evaluating the Diversity of AI-Generated Content with Diversity Profiles


17. Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents


18. Accuracy and Robustness of Model Cascades Under Data Perturbations


19. Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals


20. Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch


21. GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities


22. LLM-Derived Preference Judgments Are Not Self-Consistent


23. Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing


24. Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models


25. Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)


26. MoNe: Modular Neural Memory for Efficient Long Context Inference


27. TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation


28. Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making


29. When to Review: Spaced Repetition for Continual Pre-Training of Language Models


30. Agent Lightning v1.0: Towards Harnessed Agentic RL


31. SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models


32. Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context


33. When AI Designs AI: Innovation or Imitation?


34. SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution


35. Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning


36. Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression


37. Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations


38. LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents


39. Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks


40. LLM-Only PDDL Domain Repair with Open-Weight Models


41. TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration


42. LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap


43. Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents


44. SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning


45. LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models


46. PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs


47. DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation


48. ASI-Bench: At the Dawn of Artificial Superintelligence


49. Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking


50. Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification


51. Fool’s Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models


52. Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models


53. Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection


54. KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn


55. Toward Personal Intelligence Through Cooperative Observation


56. A decodability criterion predicts when hidden-state selection beats majority voting in large language models


57. KernelArc: A Multi-Agent Framework for GPU Kernel Optimization


58. DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization


59. Memory Is Communication: The Frontier Between Remembering and Signaling


60. SkillEffect: Checked Lowering for Memory-Bounded Agent Tools


61. The Problem Is the Problem: Towards Scalable Mathematical Discovery


62. FedPref: Federated Preference Learning for Structured Radiology Report Extraction


63. The Price of Thinking: Reasoning Effort as a Model-Specific API Contract


64. Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution


65. GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents


66. From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation


67. Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors


68. Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System


69. Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents


70. Traceable Trust for action-ready artificial intelligence in bioscience


71. Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media


72. Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity


73. Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection


74. An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models


75. SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE


76. Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation


77. Grading Needs a Rubric, Not Intelligence



79. A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning


80. Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks


81. Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition


82. BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models


83. Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints


84. AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis


85. The Model’s Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges


86. MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure


87. Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses


88. Training with synthetic data for drone detection in thermal imagery


89. Learnware for CSI Feedback: Scene-specific Small Models Can Do Big


90. What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations


91. Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models


92. Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees


93. GADR: Gathering Architecture Decision Records from Meeting Transcriptions


94. Benchmarking Automated Security Patch Backporting: How Far Are We?


95. MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps


96. DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval


97. Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision


98. From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support


99. Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges


100. HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety


101. tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots


102. DMT-Dens: Density-preserving manifold visualization for biological data


103. Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries


104. Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models


105. No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models



107. Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery



109. SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation


110. Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design


111. PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX


112. Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning


113. Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets


114. Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems


115. MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting


116. SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting


117. ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback


118. Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models


119. NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration


120. Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting


121. Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions


122. When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling


123. Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics


124. Nonadaptive Learning in Robust Nonlinear Output Regulation


125. Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction


126. Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL


127. Adaptive surrogate modeling for high-dimensional spatio-temporal output


128. Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data


129. Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement


130. COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models


131. Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer’s Disease Detection


132. PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance


133. Teach and Grow: An Agent-Centered Architecture for General Robot Learning


134. Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories


135. Token Optimization and Context Window Management in Multi-Agent AI Workflows


136. Task Specialization Fine-Tuning for Contextual Reinforcement Learning


137. The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence


138. Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases


139. Expected free energy as an information constraint on the Bethe Lagrangian


140. Q-Learning With World Models


141. Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems


142. Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions


143. From Abductive Explanations to Global Logical Rules for Node Classification in SGCs


144. Structured Driving-State Narratives for Small Language Model-Based GNSS Spoofing Detection


145. Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting


146. Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss


147. Cross-Model Memory Transfer via Target-Side Reader Adaptation


148. The 10th AI City Challenge


149. YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition


150. Without journalists, there is no journalism: the social dimension of generative artificial intelligence in the media


151. PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation


152. Position: Fairness Failure in Generative Models is an Evaluation Problem


153. Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations


154. EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning


155. WIP: LLM Odyssey: A Game-Based Platform for Teaching LLM Engineering Concepts


156. CARA: Cognitive Adaptive Recommendation Agent


157. Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval


158. Average Distance Approximation for Static Large Graphs


159. Education-centered critical policy analysis of AI: Ghana’s AI strategy as a case


160. When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice


161. Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning


162. ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection


163. AI, Brain Death Detection, and Islamic Law


164. QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents


165. CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents


166. What If AI Carried Her Imagination? Black Girls as Creators in an AI Storytelling Weekend Program


167. A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications


168. The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks


169. Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs