전체 AI 논문 - 2026-07-31

1. AISPA: User-Centric System Prompt Auditing for Large Language Model Applications


2. OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models


3. DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation


4. Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs


5. MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems


6. Selective Credibility-Limited Belief Update


7. InfoOps Bench: A live information operations safety benchmark


8. SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination


9. A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks


10. Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation


11. A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports


12. LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in Lean


13. SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute


14. Metaphor Tracer: A Theory-Informed Analysis of Hidden States


15. A foundation model of numerical intelligence with cross-disciplinary generalization


16. When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence


17. WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning


18. QuantWAMs: Calibrating at the Right Granularity for World Action Models


19. GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation


20. When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences


21. HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection


22. How Benchmarks Mis-Score Computer-Use Agents


23. Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners


24. Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents


25. PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?


26. One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence


27. Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3


28. MemHarness: Memory Is Reconstructed, Not Replayed


29. LLM-Guided Evolutionary Search for Constraint Model Reformulation to Improve Solver Efficiency


30. Operationally Guided Placement-Aware Learning for Industrial Online 3D Bin Packing


31. AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge


32. CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising


33. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents


34. Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation


35. AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach


36. ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs


37. BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints


38. Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training


39. MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck


40. SciDataSailor: Deep Scientific Data Exploring


41. An Instrument to Evaluate Governance Proposals: AI Policy Analysis at Scale


42. PerturbMap: Cross-Context Transfer of Single-Cell Perturbation Responses


43. Diversifying Personalized Research Ideation against AI-Induced Homogenization


44. Distilling Answer Set Programming Theories from Large Language Models


45. Group-Reflective Self-Distillation for Agentic Reinforcement Learning


46. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale


47. SemPIC: Learning Semantic Position-Independent KV Caches


48. IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD


49. SKILL-KD: Contrastive Skill Distillation for LLM Agents


50. DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness


51. MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images


52. MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware


53. SKIMIX: Multi-Agent Harness-Time Scaling with Skill Mixture for Dynamic Harness Engineering


54. Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation


55. MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation


56. From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents


57. Shapes from Examples: Foundations of Shape Learning in Recursive SHACL


58. The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty


59. Meta-Task: Turning Terminal Task Synthesis into a Terminal Task for Scalable Agent Training


60. One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs


61. IFHierBench: Hierarchical Instruction Following for Large Language Models


62. A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models


63. MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos


64. Dynamic Spectral Filtering for Temporal Graph Learning: Learning Evolving Propagation Operators


65. Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning


66. An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding


67. Search as Computation Allocation


68. Orca: Neural Operators for Causal Reasoning in Continuous Time


69. Back to All-Entity Ranking: Sampler-Dependent Evaluation in Continuous-Time Dynamic Graphs


70. Simplifying Neural Networks During Training


71. Virtual Process Dossier: A Process-Aware Data Catalogue


72. Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration


73. MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory


74. Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding


75. STEREODISCO: Discovering Stereotypicality in LLMs


76. MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes



78. SpecCal: Ambiguity-Aware Candidate Calibration for Infrared Spectrum-Based Molecular Structure Reconstruction


79. VeriSkill: A Self-Evolution Framework for Program Verification Skills


80. Baikal: Structured Search for Deep Research over Data Lakes


81. New Synchronous Computation Dynamics for Hopfield Networks


82. MECA: A Mechanism-Centered Agent for Constructing Well-Specified and Valuable Mathematical Conjectures


83. Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration


84. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them


85. Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling


86. Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch


87. Evaluating and Pricing Advertisements in AI-Generated Responses


88. HALO: Heterogeneous Admission through Localized Obligations for Safe Agentic Execution


89. HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data


90. ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning


91. SCOPE: Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization


92. Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures


93. CORE: In-Context Reconstruction for Unified Tabular Anomaly Detection


94. World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models


95. Wiring diagram extraction and gluing: a case study in classifying figure skating jumps using 3D dataset


96. From Minds to Models: The Intersection of Psychology and LLM Behaviours


97. What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering


98. DeepResearch Agent System


99. Evaluating Agentic Bioinformatics through Function, Evidence, and Validation


100. Using Large Language Models for Idea Generation in Innovation


101. AI Literacy: An Exercise in Power-Knowledge


102. Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks


103. A dataset of rated conceptual arguments


104. MedLLM: An Open Medical Language Model at the Sub-Billion Scale


105. INCLAIR: Inception-Based Longitudinal Clinical Anomaly Detection with Informed Reasoning


106. Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents


107. VAmoS Bench: Voice Agent Simulation Bench


108. Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems


109. Dimensionality and Measurement Precision in HLE’s Multiple-Choice Subset


110. Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs


111. Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models


112. SkillMentor: LLM Agent Self-Evolution via Learning Blind-Spot Diagnosis


113. PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments


114. PIE-APT: A Unified Framework for Temporal Planning and Contradiction Hunting via Incremental Direct-Derivation Abduction


115. Divergence Decoding: Training-Free Capability Fusion


116. Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition


117. RadHarmony: Radiological Data Handling in the Era of Agentic AI


118. KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation


119. Multi-Head Attention Residuals


120. Learning to Trace Seiberg Dualities


121. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval


122. PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball


123. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis


124. PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks


125. Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B


126. Algorithms for Structured Elections under Thiele Voting Rules


127. APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems


128. ORCA-bench: How Ready Are Language Model Agents for Oncall?


129. What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration


130. Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation


131. TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval


132. Machines that know they are aging: a framework for hardware-aware autonomous intelligence


133. QQWorld: Quantile-Quantile Matching for World Model Regularization


134. On-Policy and Off-Policy Learning for Large Action Spaces


135. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow


136. Teffic-Audio: Tell Fact from Fiction


137. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding


138. From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis


139. MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians


140. CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance


141. Agentic Method for Deterministic Validation of Legacy Code Migration


142. Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation


143. EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE


144. Agentic Metaverse Services: A New As-a-Service Paradigm


145. EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents


146. Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation


147. Vibe-FDTR: An agent-oriented framework for reproducible frequency-domain thermoreflectance data analysis


148. The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials


149. Persistent Gaussian Perturbations Prevent Oversmoothing in Recurrent Graph Neural Networks


150. Integrating AI into Requirements Quality Learning in Software Engineering Education: A TPACK-Guided Empirical Study


151. Search Strategies for Optimal Classification and Regression Trees


152. Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models


153. OPLD: On-Policy Latent Distillation for Multimodal Reasoning


154. Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game


155. Asymmetric Communication: Large Language Models and Language Games


156. Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models


157. Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging


158. Information Bottleneck Learning for Faithful Time Series Forecasting Explanations


159. On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems


160. Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction


161. Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs


162. Stimulus-Evoked Network Dynamics in Human Cortical Organoids: From a Graph-Computational Framework to Repeated-Stimulation Depression


163. VISA: A Structured Description Protocol for Agent-Based Simulation Models Towards Machine Reproducibility


164. Scaling, Lock-In, and Proxy Compliance: A Political Economy of Responsible AI


165. Flux-OPD: On-Policy Distillation with Evolving Contexts


166. RepBench: Compiling Benchmarks into Capability Representations for Large Language Models


167. Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures


168. Driving up Inference Energy on SNNs: Per-Sample and Universal Sponge Attacks


169. TAPO: Transition-Aware Policy Optimization for LLM Agents


170. Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models


171. $Σ$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems


172. SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions


173. LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference


174. Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs


175. Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting


176. Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation


177. ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance


178. Class-Aware Reinforcement Learning for Counterfactual Explanation Generation


179. RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents


180. ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents


181. EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits


182. FinanceHarness: Autonomous Financial Deep Research Framework


183. AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification


184. SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification


185. Deep Learning for Accelerated Long-Horizon Forecasting of Multicomponent Multiphase Microstructure Evolution in High-Entropy Alloys


186. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation


187. Can AI Follow In Einstein’s Footsteps?


188. Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis


189. LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts


190. Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation


191. RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy


192. VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition


193. Train Small, Deploy Large: Zero-Shot GNN Transfer Through Geometric Renormalization


194. Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation


195. Hierarchical Latent Reasoning for LLM-based Recommendation


196. A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery


197. Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities


198. ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation


199. Towards joint scaling laws with optimal batch size schedules


200. RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation


201. LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents


202. Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness


203. JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles


204. Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective


205. Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement


206. From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models


207. Arm2Air: Cross-Embodiment Skeleton Transfer for 3D Relay Formation


208. DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation


209. Revisiting the Adversarial Robustness of Graph-Based Traffic Forecasting


210. A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response


211. Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories


212. Is Solving Better Than Evaluating GenAI Solutions?


213. Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings


214. Cross-Embodiment Transfer via Behavior-Aligned Representations


215. Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games


216. ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping


217. Hierarchical Reranking for Scalable Financial RAG System


218. Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning


219. Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models


220. Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights


221. SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups


222. Rethinking Artificial Intelligence in Medical Imaging: Assumptions, Reality, and Reframing


223. Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models


224. SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements


225. Benchmarking LLM Competence on Logical Inference over Probability Operators


226. ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders


227. AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes


228. FunL2O: LLM-Guided Feature Function Design for Learning to Optimize


229. Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models


230. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System


231. RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation


232. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation


233. Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance


234. Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents


235. FAVA: Formal Authorization for Verified Agents with Evidence-Backed Permission Graphs


236. LLM-Guided Initialization for Accelerated Hybrid Quantum-Classical Medical Image Classification


237. Recursive transformers for semiconductor thermo-mechanical reliability


238. Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories


239. Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage


240. Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups


241. AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026


242. Estimating Treatment Effects for Depression in Longitudinal Therapy Switching Settings


243. Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation


244. Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review


245. MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking