전체 AI 논문 - 2026-06-08

1. How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope


2. Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle


3. Online Pandora’s Box for Contextual LLM Cascading


4. Off-Policy Evaluation with Strategic Agents via Local Disclosure


5. DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning


6. TOPSIS-RAD: Ranking According to Desires


7. Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models


8. Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation


9. DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling



11. Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization


12. StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents


13. The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective


14. Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization


15. Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning


16. Accounting for Context: Shaping Moral Credences for Value Alignment


17. Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces


18. Declarative Skills for AI Agents in Knowledge-Grounded Tool-Use Workflows


19. Workflow-to-Skill: Skill Creation via Routing-Workflow-Semantics-Attachments Decomposition


20. Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation


21. AdMem: Advanced Memory for Task-solving Agents


22. OpenSkill: Open-World Self-Evolution for LLM Agents


23. A Geometric Account of Activation Steering through Angle-Norm Decomposition


24. AEGIS: A Backup Reflex for Physical AI



26. Accelerated Fourier SAT (AFSAT): Fully Realising a GPU-based Symmetric Pseudo-Boolean SAT Solver


27. Position: Don’t Just “Fix it in Post”: A Science of AI Must Study Training Dynamics


28. CARVE-Q: Quantum-Proposed, Classically Certified Interactive Driving Repair


29. Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety


30. CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions


31. Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory


32. SafeGene: Reusable Adapters for Transferable Safety Alignment


33. DiBS: Diffusion-Informed Branch Selection


34. Detecting and Mitigating Bias by Treating Fairness as a Symmetry Operation


35. How reliable are LLMs when it comes to playing dice?


36. MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism


37. Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning


38. Twelve quick tips for designing AI-driven HPC workflows


39. Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification


40. Graph Neural Network leveraging Higher-order Class Label Connectivity for Heterophilous Graphs


41. Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders


42. Planning-aligned Token Compression for Long-Context Autonomous Driving


43. PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams


44. TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment


45. Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability


46. Watch, Remember, Reason: Human-View Video Understanding with MLLMs


47. The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs


48. Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills


49. A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning


50. Impact of Synthetic Lesional MR Images in Automated Focal Cortical Dysplasia Detection in Low-Data Scenarios


51. Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests


52. Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge


53. A robust PPG foundation model using multimodal physiological supervision


54. SleepExplain: Explainable Non-Rapid Eye Movement and Rapid Eye Movement Sleep Stage Classification from EEG Signal


55. A Temporal Spatial Minimax Rate for Smoothly-Varying Distributions in Wasserstein Space


56. Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration


57. SV-Detect: AI-generated Text Detection with Steering Vectors


58. CULTURESCORE: Evaluating Cultural Faithfulness in Video Generation Models


59. Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition


60. Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path


61. AI Sovereignty: A Qualitative Model of Strategic Competition as AI Becomes an Instrument of National Power


62. Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation


63. When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations


64. DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios


65. DualGate-Net: A Prior-Gated Dual-Encoder Framework for Histopathology Cell Detection


66. An Abstract Architecture for Explainable Autonomy in Hazardous Environments


67. RETROSPECT: RETROsynthesis via Sequential Prediction, and Chemically Transformed-ranking


68. Textual Supervision Enhances Geospatial Representations in Vision-Language Models


69. UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding


70. From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability


71. REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference


72. The Three-Ring Architecture: Governing Agents in the Era of On-Platform Organisations


73. Native3D: End-to-End 3D Scene Generation via Unified Mesh-Texture Modeling and Semantic Alignment


74. OffQ: Taming Structured Outliers in LLM Quantization by Offsetting


75. DIFFRACT: Neuralized Utility Maximization for Wireless Networks by Differentiable Programming


76. GP-Adapter: Gaussian Process CLIP-Adapter for Few-Shot Out-of-Distribution Detection


77. MetaConfigurator: AI-Assisted RDF Authoring from JSON Data


78. On the Geometry of On-Policy Distillation


79. dots.tts Technical Report


80. SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating


81. TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents


82. STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation


83. Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets


84. Phonetic Error Analysis of Raw Waveform Acoustic Models


85. Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation


86. A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders


87. DataEvolver: Automatic Data Preparation for Large Language Models through Multi-Level Self-Evolving


88. Don’t Pause: Streaming Video-Language Synchrony for Online Video Understanding


89. DaX: Learning General Pathology Representations Across Scales


90. OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios


91. When is 3D Worth It? A Resource-Performance Frontier for CNNs and Transformers in Lung CT


92. Auditing Training Data in Domain-adapted LLMs: LoRA-MINT


93. SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models


94. Didact: A Cross-Domain Capability Discovery System for Defence


95. The Fine-Tuning Trap: Evaluating Negative Transfer and the Role of PEFT in Sub-1B Mathematical Reasoning


96. ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning


97. SpectCount: Spectrotemporal Counting via Synthetic Signals Improves Large Audio Language Models


98. EASE-TTT: Evidence-Aligned Selective Test-Time Training for Long-Context Question Answering


99. Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy


100. FreeAnimate: Training-Free Human Image Animation with Preview-Guided Denoising


101. Neuro-Symbolic Learning for Long-Horizon Task Planning Under Complex Logical Constraints


102. EgoPressDiff: Multimodal Video Diffusion for Egocentric UV-Domain Hand-Pressure Estimation


103. Modeling Nonlinear Feature Interactions with Product-Unit Residual Networks


104. MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models


105. Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces


106. LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics


107. Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation


108. Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks


109. Progress-SQL: Improving Reinforcement Learning for Text-to-SQL via Progressive Rewards


110. PandaAI: A Practical Agent CQ2 for Neuro-symbolic Data Analysis And Integrated Decision-Making in Quantitative Finance


111. SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling


112. Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation


113. Lane Change Trajectory Planning for Personalized Driving Comfort and Mobility Efficiency


114. Exploring Reinforcement Learning for Fluid Transitions Between Clinical Mental Healthcare and Everyday Wellness Support


115. What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media


116. Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations


117. Generalization in Deep Neural Networks: Minimax Rates for Gradient Methods


118. Optimal Rates for Generalization of Gradient Descent Methods with Deep Neural Networks


119. AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation


120. Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection


121. HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec


122. Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations


123. SCOUT: Semantic scene COverage via Uncertainty-guided Traversal


124. MSAIC-Net: A Multi-Scale Attention and Imbalance-Aware Contrastive Network for ECG-Based Myocardial Substrate Abnormality Detection


125. ShallowBench: Benchmarking Generative Drug Design Models on Shallow-Pocket Targets


126. Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles


127. Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation


128. MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models




131. Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers


132. CAF-Gen: A Multi-Agent System for Enriching Argumentation Structures


133. How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures


134. What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?


135. ChronoForest: Closed-Loop Multi-Tree Diffusion Planning for Efficient Bridge Search and Route Composition


136. FIGMA: Towards FIne-Grained Music retrievAl


137. Re-Centering Humans in LLM Personalization


138. Direct 3D-Aware Object Insertion via Decomposed Visual Proxies


139. Generative Models Erode Human Temporal Learning Through Market Selection


140. MalTree: Tracing Malware Evolution from Embeddings at Scale


141. NTILC: Neural Tool Invocation via Learned Compression


142. WAV: Multi-Resolution Block Residual Routing for Deep Decoder-Only Transformers


143. AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps


144. MacArena: Benchmarking Computer Use Agents on an Online macOS Environment


145. IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems


146. Multi-Scale Feature Attention Network for Polymer Classification using THz Dual-Comb Spectroscopy


147. Geometric Second-Order Feature Correlation Learning for Self-Supervised Speech Emotion Recognition


148. FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models


149. Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration


150. Coordinated optimization of departure sequencing and section-track allocation in railway short-term concentrated departure scenarios based on qubo and hybrid quantum algorithms


151. Synthetic Benchmarks Overstate Forward-Forward Scaling: Real-Data Limits of Layer-Local Training


152. Attention-Guided Autoencoder Fusion for Insulator Defect Detection Using UAV Transmission-Line Imaging


153. Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models


154. Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems


155. P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8


156. DxPTA: An Architecture Design Space Exploration with Optical Dataflow-guided Strategy for HW/SW Co-Design of Photonic Transformer Accelerators


157. FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail


158. Which Anatomy Matters Under Limited Labels? A Data-Efficient Anatomy-Aware Benchmark for Cardiac Pathology Prediction


159. A Geometric Gaussian Mixture Representation of Plane Curves


160. Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?


161. Autonomous heterogeneous catalyst discovery with a self-evolving multi-agent digital twin


162. Trading Engagement for Sustainability: Carbon-Aware Re-ranking for E-commerce Recommendations


163. When Does Multi-Agent Collaboration Help? An Entropy Perspective


164. Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs