전체 AI 논문 - 2026-07-09

1. Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety


2. SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents


3. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops


4. RL Post-Training Builds Compositional Reasoning Strategies


5. Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows


6. Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning


7. SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis


8. The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents


9. InductWave: Inductive Multi-Hop Logical Query Answering on Knowledge Graphs


10. Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents


11. Agentic Data Environments


12. MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning


13. Physics-Audited Agentic Discovery in Scientific Machine Learning


14. From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents


15. Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations


16. Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks


17. Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety


18. Measuring Intelligence Beyond Human Scale


19. Learning social norms enhances compatibility in dynamic human-AI coordination


20. Large Behavior Model: A Promptable Digital Twin of the Retail Customer


21. Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix


22. The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI


23. Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics


24. Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1


25. QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron


26. LLM-powered reasoning in agent-based modeling


27. When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning


28. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation


29. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning


30. Co-LMLM: Continuous-Query Limited Memory Language Models


31. Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass


32. Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF


33. Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning


34. DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation


35. ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation


36. QCNN with Rough Path Signature Kernels


37. Future Confidence Distillation in Large Language Models


38. Towards Agentic AI Governance: A Preliminary Assessment


39. CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis


40. Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning


41. Creativity from Friction: Human-AI Interaction for Exploratory Structural Design


42. Stability of Flow Models for Graph Signals


43. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning


44. HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models


45. TimEE: End-to-end Time Series Classification via In-Context Learning


46. Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26


47. Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents


48. Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data


49. SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation


50. RLVP: Penalize the Path, Reward the Outcome


51. Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation


52. When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs


53. On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces


54. Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report


55. Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors


56. HumAIN: Human-Aware Implicit Social Robot Navigation


57. Latency-Aware Bid Acceptance under Operational Feasibility: A Public Benchmark with Hindsight Ceilings


58. Quantum simulation of real-world nonlinear dynamics via Koopman method


59. Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation


60. HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting


61. FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation


62. POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process



64. CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training


65. Bayesian Optimization of Genetic Algorithm Hyperparameters in a Multi-Fidelity Framework for Efficient Lattice Material Design


66. FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series


67. ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies


68. DiPhon: Diffusion on Graphons for Scalable Graph Generation


69. Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation


70. Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 – A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency


71. Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators


72. Predicting LLM Safety Before Release by Simulating Deployment


73. Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning


74. Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning


75. GeoProp: Grounding Robot State in Vision for Generalist Manipulation


76. AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer’s Disease Diagnosis


77. Navigating Hierarchy: Hyperbolic Learning on Brain Graphs for Disorder Diagnosis


78. Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits


79. Complexity-Budgeted, Interaction-Aware Interpretable Model for Tabular Data


80. Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production


81. Making Implicit Preservation Intent Explicit in Conversational Image Editing


82. Riemannian Geometry for Pre-trained Language Model Embeddings


83. On the Principles of Deep Feedforward ReLU Networks


84. Intrinsic Green’s Learning: Supervised Learning on Manifolds via Inverse PDE


85. AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning


86. Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies


87. Latent graph encoding of multimodal neuroimaging features with generative AI architectures


88. Multimodal Spatiotemporal-Frequency Fusion with Peak Enhancement for Cellular Traffic Forecasting


89. Physics-guided spatiotemporal neural models for fuel density prediction


90. WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time


91. Hybrid Least Squares/Gradient Descent Methods for MIONets


92. End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent


93. Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies


94. Self-Supervised Pretraining Improves Cross-Site and Cross-Scale Robustness of Point Cloud Leaf-Wood Segmentation


95. Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System


96. Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery


97. MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations


98. LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models


99. Computing with Stochastic Oracles in AI-Augmented Computation


100. ReMoDEx: A Local-to-Global Relevance-Based Model Decision Explainability Framework for large-Scale Image Datasets


101. GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model


102. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong


103. Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs


104. Ad Headline Generation using Self-Critical Masked Language Model


105. When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems


106. A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora


107. What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study


108. Enhancing deep learning models for time series classification via knowledge distillation


109. From Agentic to Autogenic Network Management for AI-Native 6G and Beyond: A Standards Perspective


110. AirPASS: Over-the-Air Federated Learning via Pinching Antenna Systems


111. SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models


112. A Continual Learning Framework for Adaptive Control of Modular Soft Robots


113. Reliable and Developer-Aligned Evaluation of Agents for Software Engineering


114. Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review


115. SPEAR: A Simulator for Photorealistic Embodied AI Research


116. tsbootstrap: Distribution-Free Uncertainty Quantification and Conformal Prediction for Time Series


117. Digital Fragmentation and Generative AI Use Across 103 Million Application Events


118. Diffusion enabled Optimal Transport distances for graph matching


119. Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering


120. The Rank-One Corner: How Much Value Equivalence Does a Task Need from a World Model?


121. At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics


122. Specification Grounding Drives Test Effectiveness for LLM Code


123. ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities


124. Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation


125. Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking


126. Open-Ended Scenario Reasoning for Specialist Model Adaptation


127. LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting


128. SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts


129. Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation


130. Inertia-1: An Open Exploration of Wearable Motion Foundation Models


131. WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning


132. STAGformer: A Spatio-temporal Agent Graph Transformer for Micro Mobility Demand Forecasting


133. PRoVeFL: Private Robust and Verifiable Aggregation in Federated Learning


134. Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts


135. Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization


136. D2PO: Optimizing Diffusion Samplers via Dynamic Preference


137. Security and Privacy in Agentic AI: Grand Challenges and Future Directions


138. NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts


139. Do Counterfactually Fair Image Classifiers Satisfy Group Fairness? – A Theoretical and Empirical Study


140. TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation


141. MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices


142. Non-contact, Real-time, Heart-rate Measurement using Image Processing with Commodity Cameras and AI Agents


143. When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents


144. LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection


145. Can Reinforcement Learning Efficiently Discover Price Manipulation?