전체 AI 논문 - 2026-07-15

1. Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution


2. Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model


3. Dynamic Resource Allocation for Ensemble Determinization MCTS


4. Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation


5. Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs


6. FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation


7. Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes


8. MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations


9. A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study


10. Solution of the Hempel’s statistical ambiguity problem and Causal AI


11. Human-AI Agent Interaction as a Neuroplastic Training Environment


12. Visual Access Boundaries in Vision-Language Model Reasoning


13. Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents


14. Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?


15. Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative


16. Tracing Agentic Failure from the Flow of Success


17. LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos


18. MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku


19. Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration


20. A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism


21. Atomic Units of X: The Compression Layer of Intelligence


22. Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing


23. Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring


24. Evidence-Grounded AI for Musculoskeletal Care


25. The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank


26. TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments


27. From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery


28. Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models


29. Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting?


30. EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading


31. Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting


32. Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions


33. Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents


34. PM-Bench: Evaluating Prospective Memory in LLM Agents


35. How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks


36. On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage


37. Rethinking the Evaluation of Harness Evolution for Agents


38. Good Benchmarks


39. A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models


40. Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems


41. The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning


42. Connected by Construction: Learning Tractable Near-Tour Marginals for Traveling Salesman Problems


43. Representing and Generating Levels Over Time through Playtrace Reconstructive Partitioning


44. Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking


45. Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations


46. Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability


47. LP Mining with LP2Graph: A Use Case for Railway Rescheduling


48. Optimization Is Not All You Need


49. Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses


50. GRID: Grammar-Railed Decoding for Enterprise SQL Generation


51. Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study


52. In-Context Reinforcement Learning under Non-Stationarity: A Survey


53. Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets


54. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale


55. PalmClaw: A Native On-Device Agent Framework for Mobile Phones


56. Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models


57. ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark


58. Real-time fall detection based on vision for low-power edge platforms


59. UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies


60. Unveiling Complex Collective Behaviors from Simple Rewards


61. ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation


62. Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature Sensitivity


63. Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques


64. PixelLoop: Shortcut Topological Navigation with Pixel-Level Loops


65. Autonomous Tracking and Terminal Guidance of Moving Targets for Fixed-Wing UAVs


66. The One-Word Census: Answer-Choice Conformity Across 44 Language Models


67. Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels


68. When Close Enough Is Not Enough: Autoregressive Drift in Quantum Circuit Synthesis


69. HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition


70. Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination


71. Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems


72. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation


73. Weakly Supervised Spatio-Temporal Candidate Discovery of Dairy Farm Sites from Seasonal Satellite Imagery


74. Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty


75. Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities


76. Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing


77. Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts


78. From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation


79. Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings


80. Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference


81. Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs


82. Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?


83. Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs


84. Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction


85. Deep Learning-based Surrogate Modelling of the LOD Method for Multiscale Problems


86. Traceback Translators Against Forgetting in Continual Fake Speech Detection


87. Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel


88. OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning


89. Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric


90. An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge


91. The Computational Basis of Confidence in Large Language Models


92. ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning


93. Demonstration of the common dual-channel feature decoupling characteristic of front-door mediation causal inference methods in whole-slice image classification


94. IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment


95. Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction


96. Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)


97. LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes


98. A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education


99. A Longitudinal Analysis of Public Discourse on AI Ethics in Education Using Twitter Data


100. The Sound of Absence: Audio-Language Embedding Models Struggle with Negation


101. Code-MUE: Measuring Code LLMs’ Uncertainty through Execution-based Semantic Interaction Graphs


102. Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents


103. Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals


104. Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor


105. RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls


106. The Benjamini–Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests


107. Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing


108. TRAIL: A Platform for Configurable Human–AI Teaming Experiments


109. From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data


110. Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models


111. GaitSpan: Growing Humanoid Locomotion from Walking to Running


112. Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap


113. Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning


114. PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs


115. Sparse Autoencoders for Interpretable Out-of-Distribution Detection


116. Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation


117. Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation


118. AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration


119. Representation and Reference Selection in Training-Free Synthetic Image Attribution


120. An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering


121. HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning


122. Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs


123. Mitigating The Effect of Class Imbalance in Data with Hierarchical and Dependable Structure


124. Sparse Inter-Layer Dependencies of Transformer FFN Neurons


125. Removable Defects: The Economics and Limits of Deliberate Deficiency


126. Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring


127. Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design


128. Signal-Guided Optimization for Machine Unlearning


129. Learning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance


130. Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems


131. Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse


132. Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability


133. Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction


134. Graph-Constrained Policy Learning for Extreme Clinical Code Prediction


135. Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG


136. BAT-RM: A Boundary-Aware Transformer with Region-Aware Multi-Directional Mamba for Clinically Deployed Cervical Cancer Radiotherapy Auto-Contouring


137. Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite


138. BattVAE-GP: Generative Modeling of Long-Horizon Battery Degradation with Uncertainty Quantification


139. How Query Visibility Changes KV-Cache Compression Rankings: A Matched-Budget Audit


140. CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA


141. Mathematics of Data Science


142. Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making


143. Do You Remember? Toward Memory-Centric Multimodal AI


144. AAAI-26 Dual Submissions: Novel Challenges


145. QDEvo: A Multi-Objective Quality-Diversity Framework for Automated Heuristic Design


146. Burst Spiking Neural Networks


147. Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming


148. SeqGPT: A Constrained Transformer Agent for the Inverse Designof Multi-Panel Composite Structures


149. OmniPMNet: Bridging discrete and gridded PM10 forecasts via omni-query neural processes


150. Graph-Based Detection of Disinformation Narrative Diffusion between Russian and Ukrainian Telegram Channels


151. I’m Sorry, but I Can’t Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs


152. G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis


153. CANDI: Contextual Alignment for Niche Domains Question Answering


154. So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis


155. Scaling Point-in-Time Language Models


156. FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis


157. Answering Without Referring: How AI Search Rewrites the Web’s Economic Bargain