전체 AI 논문 - 2026-07-01

1. AxDafny: Agentic Verified Code Generation in Dafny


2. PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines


3. Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA


4. TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models


5. Harnessing Textual Refusal Directions for Multimodal Safety


6. An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping


7. Adaptive Cluster-First Route-Second Decomposition for Industrial-Scale Vehicle Routing


8. Creating Intelligence: A Computational Foundation for AGI


9. Large Databases Need Small, Open-Weight Language Models



11. Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision


12. A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols


13. Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist


14. FARS: A Fully Automated Research System Deployed at Scale


15. Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents


16. Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy


17. Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index


18. ACE: Pluggable Adaptive Context Elasticizer across Agents


19. Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2


20. A time-series classification framework for individual-level absenteeism prediction under severe class imbalance


21. Design and Implementation of Agentic Orchestrations and Orchestration of Agents


22. Surprise as a Signal for Plasticity and Metacognition


23. One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution


24. CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift


25. CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market


26. Who Determines the Meaning of an Emotion? Affective Sovereignty as an Epistemic Consequence of Measurement Limits


27. CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes


28. Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration


29. BP-TTA: Balanced and Prototype-Guided Test-Time Adaptation in Dynamic Scenarios


30. Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs


31. Xiaomi-GUI-0 Technical Report


32. Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models


33. World-Model Collapse as a Phase Transition


34. ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents


35. Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches


36. Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models


37. CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM


38. HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)


39. Benchmarking Large Language Models on Floating-Point Error Classification


40. Spatial Reasoning via Modality Switching Between Language and Symbolic Representation


41. Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling


42. Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding


43. Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents


44. Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism


45. Long-term Traffic Simulation via Structured Autoregressive Modeling


46. Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems


47. Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping


48. AI-Assisted Discovery of Convex Relaxations via Dual Agents


49. HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents


50. ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents


51. Cross-Domain Feature Expansion for Tabular Medical Data via Knowledge Graphs Injection


52. Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics


53. Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records


54. The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory


55. Revealing Safety-Critical Scenarios for UTM via Transformer


56. DDIAgents: Mechanism-Conditioned Context Flow for Drug-Drug Interaction Prediction


57. MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning


58. OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents


59. LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents


60. Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization


61. A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management


62. When Regulation Has Memory: Hysteresis and Control Burden in Artificial Agency


63. AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents


64. HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation


65. Neuro-Bayesian-Symbolic Residual Attention Shallow Network: Explainable Deep Learning for Cybersecurity Risk Assessment


66. AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance


67. RoPoLL: Robust Panel of LLM Judges


68. Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering


69. Investigating Multi-Agent Deliberation in Law


70. Beyond expert users: agents should help users construct preferences, not just elicit them


71. When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models


72. BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation


73. How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies


74. Contrastive Reflection for Iterative Prompt Optimization


75. What Drives Interactive Improvement from Feedback?


76. Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision


77. QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents


78. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs


79. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors


80. Freeform Preference Learning for Robotic Manipulation


81. AdaJEPA: An Adaptive Latent World Model


82. FLORA: A deep learning approach to predict forest attributes from heterogeneous LiDAR data


83. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning


84. Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization


85. Amplifying Membership Signal Through Chained Regeneration


86. GR2 Technical Report


87. LUNA: Learning Universal 3D Human Animation Beyond Skinning


88. MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments


89. LeCropFollow: Latent Space Planning for Navigation in Unstructured Crop Fields


90. MVP-Nav: Multi-layer Value Map Planner Navigator


91. Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference


92. Better Understanding, Understanding Better


93. Modal CEGAR-tableaux with RECAR and resolution-based SAT-shortcuts


94. Belief Contraction in Dynamic Epistemic Logic


95. Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models


96. Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling


97. Real-Time Source-Free Object Detection


98. Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning


99. Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR


100. CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield


101. A Technical Typology of AI Systems in Public Administration


102. JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering


103. FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated Learning


104. STEB: Style Text Embedding Benchmark


105. Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue


106. Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian


107. Look But Don’t Touch with Sparse Autoencoders for Unlearning in Diffusion Models


108. RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization


109. ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping


110. When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection


111. Histogram-constrained Image Generation


112. WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models


113. Sparsity-Inducing Divergence Losses for Biometric Verification


114. Improving Certified Robustness via Adversarial Distillation


115. ECHO: Prune to act, trace to learn with selective turn memory in agentic RL


116. A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems


117. Intrinsic decomposition and editing of 3D Gaussian splats


118. A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents


119. Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models


120. Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation


121. Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models


122. Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning


123. Comparative Analysis of Machine Learning based Intrusion Detection in Realistic IoT Networks


124. Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment


125. ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models


126. DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers


127. Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer


128. Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning


129. FLARE-AI: Flaw Reporting for AI


130. CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors


131. Improving multichannel speech enhancement through accurate room-acoustic simulations


132. On the Convergence of Self-Improving Online LLM Alignment


133. FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents


134. Robustness of Robotic Manipulation: Foundations and Frontiers


135. Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets


136. Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics


137. UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation


138. DA-Studio: An Agentic System for End-to-End Data Analysis


139. Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors


140. Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?


141. Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models


142. Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images


143. Stage-Transition Dense Reward Modeling for Reinforcement Learning


144. Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?


145. From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation


146. PGUDA: Pressure-Guided Unsupervised Domain Adaptation with Cross-Modal Knowledge Distillation for sEMG-Based Gesture Recognition


147. Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models


148. 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance


149. From Idea to Prototype in an Afternoon: Scaffolded, AI-Assisted Rapid VA Prototyping


150. CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs


151. Minimizing Quantized Semantic Age of Information (QSAoI) in Foundation Model-Based Semantic Communications


152. CLIMB: Centroid-Based Hierarchical Memory for Online Continual Self-Supervised Learning


153. Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents


154. TDGT: A Tabular Data Generation Toolkit supporting adaptive GPU-accelerated Bayesian mixture models, diffusion-based models, and latent-space generative modeling


155. SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation



157. Information-Aided DVL Calibration


158. Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?


159. Distilling Temporal Coherence into 2D Networks for Transrectal Ultrasound Prostate Video Segmentation


160. Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer’s Disease Detection


161. Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation


162. AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL


163. MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents


164. ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries


165. LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music


166. One Retrieval to Cover Them All: Co-occurrence-Aware Knowledge Base Reorganization for Session-Level RAG


167. PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks


168. PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding


169. A Modular Vision-Language-Action Robotics Framework for Indoor Environments


170. UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling


171. SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos


172. What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning


173. Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation


174. When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking


175. Beyond But-for Test: Counterfactual Explanation in Abstract Argumentation via Actual Causality (Extended Version)


176. Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks


177. ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs


178. Learning Video Dynamics with Predictive Differentiable Rendering


179. Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition


180. LLM-Driven Personalities for Decision Making in Emergency Simulations


181. OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models


182. Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG


183. Loc2Repair: A Framework for Evaluating the Impact of File-Level Issue Localization in Repo-Level LLM Repair


184. Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-Ultrasound Prostate Cancer Detection


185. Motion Planning in Compressed Representation Spaces


186. Physics-informed Conditional Normalizing Flows for Angles-only Cislunar Orbit Determination


187. Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback


188. Budget-Adaptive Routing: Skipping the Weak When the Strong Answers Anyway


189. How Human Feedback Shapes AI-generated Community Notes


190. Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models


191. Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support


192. The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning


193. Test-Time Verification for Text-to-SQL via Outcome Reward Models


194. A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection


195. AI-Generated PowerShell Malware: An Experimental Framework and Dataset


196. When transformers learn “impossible” languages, what do they learn?


197. Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization


198. Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions


199. Security–Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense


200. Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin


201. A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization


202. Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens


203. Hierarchical Global Attention (HGA)


204. Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts


205. From Search to Synthesis: Training LLMs as Zero-Shot Workflow Generators


206. Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification


207. An AI-Based Solution for Secure Service Provisioning in IoT


208. BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations


209. LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents


210. Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses


211. DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning


212. Citation Discipline in Spec-Driven Development: A Cross-Model Empirical Study of Output Determinism and Automated Hallucination Detection in LLM-Generated Code


213. A Coherence Law for Trainability in Noisy Equivariant Quantum Neural Networks


214. Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout


215. Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning


216. ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models


217. Locker-based Truck-Drone Routing with Integrated Considerations of Pickups, Deliveries, and No-Fly Zones


218. Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection


219. Local Pheromone Network: Sparse Local Learning with Multi-Scale Synaptic Trails, Consolidation, and Replay


220. Emergent Culture in Minimal LLM Systems


221. Estimating the Effect of Timing on Coupon Effectiveness


222. ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education


223. Improving Survey Participation in Low-Literacy Populations Through Value-Sensitive Conversational AI


224. Agentic AI Enhances Physician Trust in Clinical Decision Making


225. AI for Quality Assurance in the Operating Room


226. Mapping the Artificial Intelligence Divide in Africa: Infrastructure, Accessibility and Capacity


227. Toward AI-Resilient Assessment in Computer Science Courses in an AI-Native World


228. The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes


229. AI Transparency: Governance Compliance or Stakeholder Requirements?


230. Can Physician Expertise Improve Machine Learning Identification of Delirium?


231. Qualified Educational Capacity Planning under Heterogeneous Student Support Needs: A Synthetic Benchmark and Decision-Support Framework


232. Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation


233. ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection


234. Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design