전체 AI 논문 - 2026-06-17

1. EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation


2. Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers


3. The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data


4. DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction


5. Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure


6. WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning


7. Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So


8. Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models


9. Knowledge Reutilization in Meta-Reinforcement Learning


10. First Proof Second Batch


11. Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding


12. IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus


13. A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation


14. Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications


15. PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience


16. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents



18. LLM Consumer Behavior Theory: Foundations of a Novel Research Field


19. STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training


20. MoCo-AIS: A Contrastive Learning Framework for Similarity Computation of Vessel Trajectories


21. Small Initialization Matters for Large Language Models


22. How Inference Compute Shapes Frontier LLM Evaluation


23. PreAct: Computer-Using Agents that Get Faster on Repeated Tasks


24. DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue


25. Learn to Quantify Social Interaction with Constraints for Pedestrian Walking


26. MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning


27. Structural Preservation and the Logical Expressiveness of Graph Neural Networks


28. StepGuard: Guarding Web Navigation via Single-Step Calibration


29. FlowRAG: Synergizing Explicit Reasoning via Frequency-Aware Multi-Granularity Graph Flow


30. A homotopy-type-theoretic generalization of neurosymbolic inference


31. WallZero: Mastering the Game of WallGo with Strategic Analysis


32. DecoSearch: Complexity-Aware Routing and Plan-Level Repair for Text-to-SQL


33. Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs


34. LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings


35. EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent


36. FllumaOne: A Code-Native Multimodal CAD Dataset with Executable Programs and Kernel-Validated Feature Histories


37. Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games


38. From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs


39. Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns


40. FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness


41. Brick-DICL: Dynamic In-Context Learning for Automated Brick Schema Classification


42. Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning


43. Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow


44. DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack


45. SEAGym: An Evaluation Environment for Self-Evolving LLM Agents


46. LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline


47. Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation


48. Dissecting model behavior through agent trajectories


49. MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors


50. A Machine-Learned Comorbidity Index


51. Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems


52. Treatment Response Optimized Clinical Decision Support AI System via Digital Twin Simulation


53. Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes


54. SpeechDx: A Multi-Task Benchmark for Clinical Speech AI


55. MemTrace: Probing What Final Accuracy Misses in Long-Term Memory


56. Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty


57. Nothing from Something: Can a Language Model Discover 0?


58. Skill-Constrained Model Predictive Control for Resilient Manufacturing Supply Chains


59. SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions




62. Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement


63. ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues


64. Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents


65. Looped World Models


66. RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills


67. A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models


68. Kolmogorov Regression for Robust Diffusion Policies


69. IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction


70. All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code



72. ReAge3D: Re-Aging 3D Faces with View Consistency


73. Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD)


74. Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour


75. Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines


76. Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping


77. Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models


78. Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning


79. Querying an astronomical database using large language models: the ALeRCE text-to-SQL system


80. S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices


81. EAGG: Embodiment-Aligned Grasp Generation via Geometry-Aware Graph Conditioning


82. Volterra Generative Models


83. When LLMs Analyze Scars: From Images to Clinically-Meaningful Features


84. Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond


85. When AI Says “I have been in similar situations”: Synthetic Lived Experience in Peer-Like Caregiver Support


86. When English Isn’t the Best Teacher: Source Language Effects in Cross-Lingual In-Context Learning


87. Catastrophic Forgetting is Low-Rank: A Function-Space Theory for Continual Adaptation


88. LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling


89. C2FL: Clustered Continual Federated Learning under Spatial and Temporal Drift


90. A T-API-Compliant ReAct Agentic Loop for Optical Networks: Generic vs. Domain-Specific Tool Abstractions


91. Multiple cyclicity and Wavelet Decomposition with Channel Correlation for Long-term Time Series Forecasting


92. Recover Semantics First, Generate Better: Improved Latent Modeling for 3D MRI Reconstruction and Cross-Contrast Synthesis


93. SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation


94. A Neuro-Symbolic Approach to Strategy Synthesis for Strategic Logics


95. Robustness of Similarity-based Positional Encoding Under Rotations: Theoretical Analysis and Experimental Validation


96. SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs


97. Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model


98. KANLib – An Modular, Extensible and Fast Kolmogorov-Arnold Network Implementation


99. PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space


100. Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization


101. Non-negative Elastic Net Decoding for Information Retrieval


102. Dimensionality Controls When Modularity Helps in Continual Learning


103. AI Adoption Across a Multinational Workforce: Sociotechnical Conditions for GenAI Acceptance in Human Resources


104. AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor


105. A Quantitative Analysis of Multimodal Biomarkers in Alzheimer’s Disease


106. High-Fidelity 3D Geometric Reconstruction of Pelvic Organs from MRI: A Hybrid Deep Learning and Iterative Optimization Approach


107. Perceptual compensation for tonal context in self-supervised speech models


108. Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode Connectivity


109. When Multiple Scripts Matter: Evaluating ASR in Clinical Settings


110. Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows


111. A Framework for Evaluating Agentic Skills at Scale


112. Conservation Laws for Modern Neural Architectures


113. No-Free-Fairness: Fundamental Limits and Trade-offs in Learning Systems


114. Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering


115. LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams


116. MIVE: A Minimalist Integer Vector Engine for Softmax LayerNorm and RMSNorm Acceleration


117. A Neuromorphic Trigger for Efficient Audio Event Detection


118. Talking to Your Data: Exploring Embodied Conversation as an Interface for Personal Health Reflection


119. Symplectic Transversality and Endpoint Green Estimates for Finite-Horizon Pontryagin Systems


120. ED3R: Energy-Aware Distributed Disaster Detection Enabled by Cooperative Robotic Agents


121. Structured Adversarial Camouflage via Voronoi Diagrams


122. Vision-language models for chest radiography do not always need the image


123. Confusion-Aware Transfer Teacher Curriculum Learning Framework: Disentangling Scoring and Pacing Effects


124. SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology


125. SuCo: Sufficiency-guided Continuous Adaptive Reasoning


126. See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL


127. ASTEROID: A Spatiotemporal Information Transformer for Forecasting Multi-Step Time Series of Molecular Dynamics


128. Handling Feature Heterogeneity with Learnable Graph Patches


129. FacProcessTwin: An LLM-Based System for Process Twin Development


130. Temporal Preference Optimization for Unsupervised Retrieval


131. TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins


132. A Risk Decomposition Framework for Pre-Hoc Fine-Tuning Prediction


133. SketchXplain: Intuitive Visual Explanations of Image Classifiers with Sketches


134. Bounding Box Label Propagation for Re-Annotation of Document Layout Analysis Datasets


135. Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition


136. SkillMoV: Mixture-of-View Routing with Prototype-Conditioned Gating for Unified Multi-View Proficiency Estimation


137. Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations


138. Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics


139. LLM Features Can Hurt GNNs: Concatenation Interference on Homophilous Graph Benchmarks


140. Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery


141. An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts


142. Reversal Q-Learning


143. Offline Preference-Based Trajectory Evaluation


144. Reinforcing Dual-Path Reasoning in Spatial Vision Language Models


145. OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation


146. Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery


147. FoundCause: Causal Discovery with Latent Confounders from Observational Data


148. Unlocking LLM Code Correction with Iterative Feedback Loops


149. Geometry-Aware Post-Hoc Uncertainty Quantification in Operator Learning


150. MagicSim: A Unified Infrastructure for Executable Embodied Interaction


151. Online LLM Selection via Constrained Bandits with Time-Varying Demand


152. Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing


153. AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows


154. AUTOGATE: Automated Clock Gating via Toggling-Aware LLM-based RTL Rewriting


155. MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation


156. Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure


157. Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos


158. Feynman Kac Reweighted Schrödinger Bridge Matching for Surface-Based Tau PET Harmonization


159. L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification


160. Enhancing Pathological VLMs with Cross-scale Reasoning


161. Discrete Autoregressive Transformer for Generative Mechanism Synthesis


162. Graph Neural Networks for Semi-Supervised Image Classification with Multi-Feature Aggregation


163. Bridging Spatial And Frequency Views For Disaster Assessment: Benefits And Limitations


164. The Discrete-Log Clock: How a Transformer Learns Modular Multiplication


165. SoK: AI-Augmented Binary Reversing


166. NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama


167. Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models


168. TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations


169. Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation


170. MeiBRD: Meta-Learning Intraoperative Biomechanical Residual Deformation


171. Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication


172. DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models


173. Translating the Untranslatable: An Operationalizable Ontology for Untranslatability


174. Do Large Language Models Always Tell The Same Stories?


175. Counterfactual Optimization of Baseball Pitch Sequences and Estimation of Its Impact on Season-Level Statistics


176. Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation


177. Transformer-Based Warm-Starting for Feasible and Optimal Terminal Approach to Tumbling Objects with Space Manipulators


178. From Democracies to Autocracies: How AI Systems Enable Authoritarianism by Design


179. ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software


180. Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering


181. MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task


182. Physics-Informed Attention Mechanism and Generalization Capability of Deep Learning-Based Grain Growth Evolution Prediction


183. Rift: A Conflict Signature for Deception in Language Models


184. Trust-Aware Multi-Agent Traceability: Confidence-Calibrated Knowledge Graphs for Consistent Software Artifact Management


185. PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation


186. Cluster-Aware Dual-Level Test Specification Generation for Large-Scale Automotive Software Requirements


187. Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference


188. PromptMN: Pseudo Prompting Language


189. Agentic Discovery of Non-Canonical Antimicrobial Peptides with AMPGAN v3


190. Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control


191. LineageMark: Multi-user White-box Watermarking for Contribution Tracing in Model Derivation Chains


192. TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations


193. Graph neural networks at war: integrating cybersecurity and drone intelligence in the Israeli-Iranian conflict


194. MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs


195. Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis


196. An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios


197. Timestamp-Aware Spatio-Temporal Graph Contrastive Learning for Network Intrusion Detection


198. Models Take Notes at Prefill: KV Cache Can Be Editable and Composable


199. Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators


200. Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models


201. Software Delegation Contracts: Measuring Reviewability in AI Coding-Agent Work


202. ANEForge: Python for direct computation on the Apple Neural Engine


203. ZIVARI-TLBO: A Zero-Cost Inter-Group Evaluated-Elite Relay Mechanism for Teaching-Learning-Based Optimization


204. ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking


205. The Price of Anarchy in Disaggregated Inference


206. HRDX: A Large-Scale Vector HD-Map Dataset


207. Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework


208. CMIP-Forge: An Agentic System that Retrieves, Computes, and Self-Reviews Climate Science


209. Surveying GenAI-based Automation in Printed Circuit Board Design and Test


210. Extracting Semantics: LLM-Guided Automatic Population of Robot Ontology from URDF


211. KFTD: Koopman-Fourier Time-Differentiable Network for Continuous Ocean Spatiotemporal Forecasting


212. PIVOT: Bridging Black-Scholes Implied-Volatility and Price Objectives via Differentiable Jäckel Operator


213. Towards Distributed Inference of LLMs on a P2P Network


214. Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs