전체 AI 논문 - 2026-06-12

1. Automated reproducibility assessments in the social and behavioral sciences using large language models


2. Agents-K1: Towards Agent-native Knowledge Orchestration


3. EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery


4. Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization


5. Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks


6. AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility


7. Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning


8. Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch


9. EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis


10. Reward Modeling for Multi-Agent Orchestration


11. Multiagent Protocols with Aggregated Confidence Signals


12. A Three-Layer Framework for AI in Scientific Discovery


13. Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation


14. Uncertainty-Aware Hybrid Retrieval for Long-Document RAG


15. CloudCons: A Comprehensive End-to-End Benchmark for Cloud Resource Consolidation


16. Why Sampling Is Not Choosing: Intentionality, Agency, and Moral Responsibility in Large Language Models


17. Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems


18. Optimizing Appliance Scheduling for Solar Energy Management Using Metaheuristic Algorithms


19. Neuro-Symbolic Agents for Regulated Process Automation: Challenges and Research Agenda


20. MiniMax Sparse Attention


21. A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget


22. IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing


23. Can I Buy Your KV Cache?


24. ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning


25. Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video


26. ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space


27. From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification


28. MOSAIC: Modality-Specific Adaptation for Incremental Continual Learning in Parkinson’s Disease Gait Assessment


29. Multi-Field Hybrid Retrieval-Augmented Generation for Maritime Accident Root Cause Analysis


30. EPIG: Emotion-Based Prompting for Personalised Image Generation


31. Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm


32. LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis


33. Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints


34. A Minimal Model of Bounded Trade-Off Screening in Multi-Attribute Choice


35. ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning


36. Under What Conditions Can a Machine Become Genuinely Creative?


37. Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach


38. Mental-R1: Aligning LLM Reasoning for Mental Health Assessment


39. TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?


40. Rethinking RAG in Long Videos: What to Retrieve and How to Use It?


41. AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction


42. Augmentation techniques for video surveillance in the visible and thermal spectral range


43. Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior


44. SciR: A Controllable Benchmark for Scientific Reasoning in LLMs


45. Otters++: A Time-to-first-spike Based Energy Efficient Optical Spiking Transformer


46. The Illusion of Multi-Agent Advantage


47. APCyc: Property-Informed Design of Cyclic Peptides via Automated Cyclization


48. Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation


49. A Mathematical Forum Platform for Collaborative Problem Solving and Dataset Generation for AI Reasoning


50. Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models


51. OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models


52. Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory


53. PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization


54. MARS: Margin-Adversarial Risk-controlled Stopping for Parallel LLM Test-time Scaling


55. Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce


56. MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback


57. Zero-source LLM Hallucination Detection with Human-like Criteria Probing


58. The Hidden Power of Scaling Factor in LoRA Optimization


59. HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness


60. DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks


61. WISE: A Long-Horizon Agent in Minecraft with Why-Which Reasoning


62. (Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable


63. Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement


64. Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and an Early-Warning Signature for Emerging Topics


65. GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models


66. Teach-and-Repeat: Accurately Extracting Operational Knowledge from Mobile Screen Demonstrations to Empower GUI Agents


67. MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs


68. The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements


69. A Tutorial on World Models and Physical AI


70. Constructing Evaluation Datasets for Procedural Reasoning: Balancing Naturalness, Grounding, and Multi-Hop Coverage


71. Prefill Awareness in Large Language Models


72. Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices


73. Benchmarking AI Agents for Addressing Scientific Challenges Across Scales


74. Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior


75. The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism


76. Definitional alignment before capability alignment: a Design-Science framework for adjudicating claims about AGI


77. Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System


78. From AGI to ASI


79. Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents


80. TrajGenAgent: A Hierarchical LLM Agent for Human Mobility Trajectory Generation


81. “Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms


82. PersonaDrive: Human-Style Retrieval-Augmented VLA Agents for Closed-Loop Driving Simulation


83. Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation


84. Strategic Decision Support for AI Agents


85. Arbor: Tree Search as a Cognition Layer for Autonomous Agents


86. ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs


87. Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning


88. Mana: Dexterous Manipulation of Articulated Tools


89. SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning


90. SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation


91. Valid Inference with Synthetic Data via Task Exchangeability


92. One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders


93. Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models


94. EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution


95. LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories


96. ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages


97. Existence Precedes Value: Joint Modeling of Observational Existence and Evolving States in Time Series Forecasting


98. Contrast-Informed Augmentation and Domain-Adversarial Training for Adult-to-Neonatal MR Reconstruction Generalization


99. Adaptive Turn-Taking for Real-time Multi-Party Voice Agents


100. AgentRivet: an automated system for producing Rivet routines from journal publications


101. Measurement-Calibrated Multi-Camera Fusion for Vision-Based Indoor Localization


102. Heterogeneous LiDAR Early Fusion and Learned Re-Ranking Strategy for Robust Long-Term Place Recognition in Unstructured Environments


103. CRAFTIIF: Cross-Resolution Analytic Four-Type Interpretable Isolation Forest for Multivariate Time Series Anomaly Detection


104. SupraBench: A Benchmark for Supramolecular Chemistry


105. MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling


106. Understanding the Rejection of Fixes Generated by Agentic Pull Requests – Insights from the AIDev Dataset


107. Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations


108. Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests


109. OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data


110. PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update


111. Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities


112. Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents


113. SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation


114. An LLM System for Autonomous Variational Quantum Circuit Design


115. Real-Time Execution with Autoregressive Policies


116. IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds


117. Dual-Domain Equivariant Generative Adversarial Network for Multimodal CT-PET Synthesis


118. Rarity-Gated Context Conditioning for Offline Imitation Learning-Based Maritime Anomaly Detection


119. Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories


120. HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers


121. Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality


122. Once-for-All: Scalable Simultaneous Forecasting via Equilibrium State Estimation


123. Different Layers, Different Manifolds: Module-Wise Weight-Space Geometry in Transformer Optimization


124. Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes


125. Towards Personalized Federated Learning for Dysarthric Speech Recognition


126. Towards More General Control of Diffusion Models Using Jeffrey Guidance


127. ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm


128. Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier


129. ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling


130. Proprioceptive-visual correspondence enables self-other distinction in humanoid robots


131. Transformer-Guided Graph Attention for Direct Cardiac Mesh Reconstruction: A Structural Digital Twin Framework


132. Modern analog computing for solving differential and matrix equations


133. MemRefine: LLM-Guided Compression for Long-Term Agent Memory


134. NTS-CoT: Mitigating Hallucinations in LLM-based News Timeline Summarization with Chain-of-Thought Reasoning


135. Iterative Visual Thinking: Teaching Vision-Language Models Spatial Self-Correction through Visual Feedback


136. Cascade Classification of Dermoscopic Images of Skin Neoplasms with Controllable Sensitivity and External Clinical Validation


137. MiniPIC: Flexible Position-Independent Caching in <100LOC


138. Select and Improve: Understanding the Mechanics of Post-Training for Reasoning


139. NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation


140. MP3: Multi-Period Pattern Pre-training forSpatio-Temporal Forecasting


141. G-Long: Graph-Enhanced Memory Management for Efficient Long-Term Dialogue Agents


142. Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents


143. Emotional regulation improves deep learning-based image classification


144. The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems


145. “Is This Not Enough?”: Asymmetries in Institutional Accountability and Collective Sensemaking in the Case of Canada’s Algorithmic Visa Triage System


146. TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization


147. EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation


148. Fault Lines: Navigating Ethics and Responsible AI Where National Policy Meets Local Practice in Public Sector Transformation


149. TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment


150. Democracy in the Era of Artificial Intelligence


151. CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts


152. scLLM-DSC: LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering for Single-Cell RNA Sequencing


153. A Machine Learning Framework for Real-Time Personalized Ergonomic Pose Analysis


154. Diffusion Transformer World-Action Model for AV Scene Prediction


155. Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models


156. An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics


157. Order Is Not Control


158. LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold


159. MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems


160. Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning


161. PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning Agent


162. Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement


163. Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming


164. JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications


165. TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for Masked Diffusion Language Models


166. OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction


167. The Internet of Agentic AI: Communication, Coordination, and Collective Intelligence at Scale


168. Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning


169. DIMOS: Disentangling Instance-level Moving Object Segmentation


170. Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata


171. Localizing Anchoring Pathways in Language Models


172. Stubborn: A Streamlined and Unified Reinforcement Learning Framework for Robust Motion Tracking and Fall Recovery for Humanoids


173. SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning


174. Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning


175. Agentic MPC for Semantic Control System Resynthesis


176. LLMs Can Better Capture Human Judgments–With the Right Prompts


177. PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections


178. AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages


179. SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems


180. LLM-Powered Personalized Glycemic Assessment in Type 2 Diabetes with Wearable Sensor Data


181. Two-Layer Linear Auto-Regressive Models Estimate Latent States


182. EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence


183. M*: A Modular, Extensible, Serving System for Multimodal Models


184. A Zero-shot Generalized Graph Anomaly Detection Framework via Node Reconstruction


185. Free-Placement Optimization of Ground Station Locations for Low-Earth Orbit Satellites


186. CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents


187. BASENet: Band-Adapted Speech Enhancement Network with Cross-Band Attention


188. Token Complexity Theory for AI-Augmented Computing


189. Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents


190. Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns


191. HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection


192. From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation


193. Emerging Flexible Designs for Geospatial Multimodal Foundation Models


194. Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs


195. Graph Reduction in Multirelational Networks: A Spreading-Oriented Reduction Benchmark


196. EDEN: A Large-Scale Corpus of Clinical Notes for Italian


197. Foresight: Iterative Reasoning About Clues that Matter for Navigation


198. Boosting Direct Preference Optimization with Penalization


199. A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints


200. Improving Crash Frequency Prediction from Simulated Traffic Conflicts Using Machine Learning Based Microsimulation


201. Speculative Rollback Correction for Quality-Diverse Web Agent Imitation


202. Representing Time Series as Structured Programs for LLM Reasoning


203. ReCal: Reward Calibration for RL-based LLM Routing


204. Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics


205. SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems


206. Occupational Prompting Reveals Cultural Bias in Large Language Models


207. Reframing AI Loss of Control: What It Is, How to Have It, How to Lose It


208. Generativism: Toward a Learning Theory for the Age of Generative Artificial Intelligence


209. Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots


210. Algorithmic Constitutionalism


211. Will AI Agents Free Us From Meaningless Work? A Human-Centered Analysis


212. Muse Spark Safety & Preparedness Report


213. Mapping AI Programs in the U.S: A Status Report from Early 2026 and an Analysis of AI Majors and Minors


214. An Explainable AI Assistant for Introductory Programming Education: Improving Feedback Reliability with Instructor-AI Collaboration


215. AI-Automation Tooling in Computer Engineering Education: Mixed-Methods TAM/UTAUT Evidence for a General Acceptance Attitude


216. The Challenges of Balancing AI Compliance and Technological Innovations in Critical Sectors: A Systematic Literature Review


217. Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering


218. Eigenism: Ethics for a Human-AI Future


219. GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns


220. Divination by Prompt: LLM-Mediated Xuanxue on Chinese Social Media



222. AI SciBrief as a Gateway to Research: A Framework for Onboarding Students into New Research Areas