전체 AI 논문 - 2026-07-16

1. Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models


2. Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education


3. AI-accelerated End-to-End Framework for Rapid Professional Upskilling


4. Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0


5. A Self-Evolving Agent for Longitudinal Personal Health Management


6. AIMO Interpretability Challenge


7. Experience Memory Graph: One-Shot Error Correction for Agents


8. CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems


9. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities


10. When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects


11. Explaining Reinforcement Learning Agents via Inductive Logic Programming


12. UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following


13. STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle


14. Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System


15. SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing


16. AI advice suppresses people’s willingness to say “I don’t know”, even when the advice is wrong and accuracy is incentivized


17. Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling


18. How Far Can Root Cause Analysis Go on Real-World Telemetry Data?


19. LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning


20. Set-shifting Behavioral Test for Harnessed Agents


21. EZSMT Version 3, Matured


22. Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases


23. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable


24. Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management


25. AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation


26. Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science


27. CayleyR: Solving the TopSpin puzzle via cycle intersection


28. Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models


29. Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents


30. Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools


31. Self-Improvements in Modern Agentic Systems: A Survey


32. Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap’s Typed Intensional FOL


33. Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution


34. SPINE: Bridging the Cyber-Physical Gap with Agentic AI


35. OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets


36. Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study


37. Early Adoption of Agentic Coding Tools by GitHub Projects


38. Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study


39. Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth


40. Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation


41. The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce


42. Music-to-Dance Generation via Atomic Movements


43. Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code


44. Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings


45. Verifying formulas for interventional distributions


46. Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild


47. AI-Augmented Human Resource Management? Insights from German companies


48. NodeImport: Imbalanced Node Classification with Node Importance Assessment


49. Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning


50. Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection


51. Prospective clinical indication, post-hoc report leakage, and fusion design in multi-image chest radiograph classification: a patient-clustered evaluation


52. CAS I: A Geometric Coding Theorem


53. Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations


54. MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model


55. Anatomically Faithful but Temporally Blind: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography


56. How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement


57. Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs


58. Social Simulations: from Agent-Based Modeling to Digital Twins


59. Barnamala: Parameter-Efficient Handwritten Devanagari Recognition at Benchmark Saturation


60. Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models


61. Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction


62. Consensus as Privileged Context for Label-Free Self-Distillation


63. OvisOCR2 Technical Report


64. From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception


65. The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models


66. Semantic Anchoring for Robotic Action Representations


67. Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities


68. Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents


69. From Prediction to Collaboration: Interactive Symbolic Music Analysis


70. Agile perceptive multi-skill locomotion for quadrupedal robots in the wild


71. IMMNet: Hybrid Fusion of Model-based and Data-driven Approaches for Maneuvering Target Tracking


72. Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification


73. GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning


74. Spectral-Informed Neural Networks Outperform Spectral Methods in High-dimensional PDEs


75. UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors


76. Grounded world models in biological organisms and future embodied AI


77. Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning


78. ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level


79. DeepLoop: Depth Scaling for Looped Transformers


80. Explainable Artificial Intelligence for Anomaly Detection in Banking Transactions: An Internal Audit Perspective


81. DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments


82. GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding


83. Adversarial Prompting Framework for AI Safety Assessment


84. Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection


85. Learning Physics-Guided Residual Dynamics for Deformable Object Simulation


86. Discrete Diffusion Models: A Unified Framework from Tokenization to Generation


87. Data-Efficient Adaptation of LLMs via Attention Head Reweighting


88. ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding


89. Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents


90. Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification


91. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models


92. Price of Fairness in Bandits: A Tight Minimax Characterization


93. The Café in Amsterdam: When the Incumbent Becomes the Oracle


94. Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System


95. Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition


96. The Refusal Residue: When Probes Catch Alignment Faking and When They Don’t


97. Efficient Text-to-Audio Generation via Pruning


98. Privacy Preserving Recommender Systems Balancing Personalization with Privacy


99. Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains


100. Tabular Foundation Models for Discrete Choice Estimation


101. Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks


102. Faithful Autoformalization of Natural Language Assertions


103. Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners


104. Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance


105. Reassessing Muon for Matrix Factorization


106. EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting


107. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System


108. Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift


109. Classifying daily activities needs posture, reconstructing them needs motion


110. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference


111. RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar


112. SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy


113. What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors


114. Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach


115. Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation


116. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation


117. AI in Cyberpsychology: A systematic literature review of Cybersecurity enhancement by using AI for analyzing psychology of Victims, Attackers, and Defenders


118. CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion


119. SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests


120. A Hybrid Mamba for Audio-Visual Navigation


121. STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting


122. Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing


123. TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling


124. WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency


125. Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit


126. Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates


127. Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models


128. Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework


129. Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing


130. SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification


131. Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs


132. Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows


133. The Hitchhiker’s Guide to Monoculture


134. The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators


135. When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data


136. HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models


137. Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes


138. Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review


139. Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems


140. Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges


141. The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI


142. Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning


143. LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents


144. Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance


145. Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants


146. Designing Safety-Constrained LLM Systems for Public Health Information Access


147. Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry


148. FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents