전체 AI 논문 - 2026-07-24

1. Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana


2. Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning


3. OpenForgeRL: Train Harness-native Agents in Any Environment


4. MIRROR: Learning from the Other View for Multi-Modal Reasoning


5. The Boundaries of Automation: A Theory of Persistent Human Participation


6. Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation


7. Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems


8. Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry


9. Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks


10. AREX: Towards a Recursively Self-Improving Agent for Deep Research


11. Detecting LLM-Generated Tokens in Human–LLM Coauthored Text


12. Agent-Guided Relational Concept Discovery: Toward Interpretable Surgical Margin Assessment


13. Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for Dynamic Graph Systems


14. PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning


15. Logical Regression for Planning with Axioms


16. Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog


17. MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning


18. Multimodal Pretraining for Generalizable EEG Representation Learning


19. Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls


20. SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning


21. Regulating autonomous and agentic AI


22. Expert Behavior Prior Reinforcement Learning


23. An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models


24. BasketEvent: Understanding Who Did What and When in Basketball Videos


25. Logic Programming Semantics for Causal Processes


26. ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders


27. How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic Programming


28. A New Well-Supported Semantics for Description Logic Programs


29. Bound-Founded Semantics for Answer Set Programming with Difference Constraints: Preliminary Report


30. Identifying Good Rules for Efficient SAT Encodings of Single-Constant Multiplication Using Machine Learning


31. Differentiable Logic Programming to Mitigate Reasoning Shortcuts in Neurosymbolic Systems


32. Explaining Weather Bulletins via ILP


33. Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications


34. SafeStep: AI-powered Travel Assistance for Elderly People with Frailty or Dementia


35. V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure


36. AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning


37. Can Generative Recommendation Reach Cold Items? A Temporal Perspective on Semantic-ID Generation


38. Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs


39. HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices


40. EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization


41. Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers


42. Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory


43. Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills


44. GuardianAgentBench: Where Agents Fail and How to Guard Them


45. Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence


46. Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents


47. From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data


48. Clustered Edge Intelligence: Beyond Just Convergence of Edge Computing and AI


49. SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration


50. Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority


51. OPOD: On-Policy Omni Distillation


52. Traceable Scholarship: Page Anchors and Ariadne’s Thread for Humanistic Inquiry in the Age of Generative AI


53. Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning


54. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions


55. Code Monitor Red Teaming for Public-Test-Passing Code


56. Auditing Evidence Use in Medical LLM Diagnosis


57. Auditing Provenance Sensitivity in LLM Agent Action Selection


58. Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks


59. Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs


60. Profiling Lightweight Large Language Models


61. Can an AI System Be Creative? A Critical Perspective from Art and Engineering


62. Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling


63. The Human-AI Substitution Principle: When will you be replaced by AI in your organization?


64. ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management


65. NVIDIA-labs OO Agents: Native Python Object-Oriented Agents


66. WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms


67. KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback


68. AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optimize Graph Neural Networks


69. CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning


70. StrideDiffusion: Accelerating Diffusion Models for Time-series Generation


71. AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use


72. DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers


73. PromptPack: Scaling LLM Annotation Agents for Online Recommendation


74. Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis


75. ConfidenceBench: Evaluating Confidence Calibration in Large Language Models


76. Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro


77. Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models


78. Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving


79. CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits


80. Reliability-Aware LLM Alignment from Inconsistent Human Feedback


81. SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning


82. Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain


83. MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference


84. Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval


85. LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization


86. Inducing Comparability of Factorised Probability Distributions


87. MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation


88. FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts


89. ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models


90. AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs


91. From Errors to Rules: Iterative Prompt Optimization for Text Classification


92. Workload-Aware Caching for Multi-Agent Systems


93. Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation


94. Attention-based Experience Replay Framework for Continual Learning of Agnostic Time Series Forecasting Models


95. DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making


96. CRAWO: Custom Resources for Adaptive Workload Orchestration


97. EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL


98. Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants


99. Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering


100. OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining


101. Expectation Alignment of Language Models for Real-World User Expectations


102. The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path


103. Tractable Hierarchical Control of Autoregressive Language Models


104. PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails


105. Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating


106. Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source Inference


107. Beyond Liars’ Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs


108. Semi-Supervised Text-Attributed Graph Distillation


109. Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment


110. SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification


111. VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification


112. Incomplete Prompt Jailbreaks in Large Language Models


113. Robust Critics: Defending LLMs Against Multi-Turn Attacks


114. Benchmarking the Personalization Capabilities of Large Language Models


115. PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs


116. DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions


117. InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents


118. DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding


119. JAXBench: Benchmarking Autonomous TPU Kernel Optimization


120. Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs


121. ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models


122. Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts


123. AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics


124. 3D-Aware VLMs with Implicit and Explicit Geometries


125. GraphVid: Interactive Graph-Controllable Video Generation


126. Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$


127. Synthetic data generation framework for quality control automation in gravure printing


128. Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity


129. Visual Contrastive Self-Distillation


130. From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs


131. ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing


132. GS-Agent: Creating 4D Physical Worlds With Generative Simulation


133. Improved lower bounds for the Shannon capacity of odd cycles


134. Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it


135. Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections


136. Error Certificates for KV-Cache Eviction via Randomized Design


137. Thinkink: 2D Spatial Ink-native Interaction with LLMs


138. RUMBA: Russian User Memory Benchmark


139. Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping


140. Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models


141. Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas


142. When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation


143. VoLN: Vision-Only Long-Horizon Navigation—Paradigm, Benchmark, and Method


144. Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy


145. DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation


146. Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks


147. M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data


148. Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin


149. From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics


150. Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation


151. GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG


152. PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing


153. Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study


154. AI Assistants Overassist


155. Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning


156. Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning


157. A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset


158. pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development


159. slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek


160. Explainable Belief Harmonization under Dynamic Epistemic Partitions


161. Explainability Framework for Policy-Aware Autonomous Agents


162. Hybrid MKNF with Classical Negation in the Rule Component


163. Towards a Certifying Grounder


164. Declarative Problem Solving in UAM Strategic Deconfliction


165. Case study: solving P-99 with LPTP and an LLM


166. Chess_db: A framework for working with large chess game datasets


167. Animation, Verification and Visualisation of Prolog Transition Systems with ProB


168. Encoding Event-B Proof Rules in Prolog: An Interactive Sequent Prover for ProB


169. Case study: proving sqrt(2) irrational with LPTP and an LLM


170. Representative Sets in Propositional Abduction


171. CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA


172. One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs’ Clarification Policies


173. Demographically-Informed Heat-Mortality Risk Curves via Risk Graph Neural Networks


174. Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core


175. Relative Value Learning


176. GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes


177. TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning


178. Training Large Language Models for Self-Explanation Faithfulness


179. Sparse Concept Channels in Frozen 3D CT Vision Encoders


180. HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving


181. Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter–Tissue Interaction


182. Scientific exploration, collaboration and labor division in the large language model era


183. Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation


184. Anti-Goal Reasoning: Rethinking the Theory of Goal Reasoning in Non-Axiomatic Logic


185. TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging


186. Probabilistic Residual Learning for Online Recommendations


187. Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery


188. Beyond Heavy Log Curation: Perplexity-Based APT Detection via Unsupervised, Context-Augmented Language Models


189. The Geometry of Personality: Activation Steering with Jungian Cognitive Functions


190. Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test


191. Robostral Navigate


192. HARP: The Human–AI Research Platform


193. Emergent Compositional Skills in Mixture-of-Experts VLAs


194. Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles


195. IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests


196. Self-Supervised Bio-Inspired Robotic Trajectory Planning with Obstacle Avoidance


197. GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning


198. Operational Identity: A Finite Audit of Declared and Implemented Rules of Sameness



200. U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation


201. A Framework for Reputation Aware Uninorm-driven Consensus Algorithms for Blockchain Networks


202. DS@GT ARC at ImageCLEFmed GANs 2026: Geometric Filtering for Privacy-Preserving CT Slice Generation


203. Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis


204. From Agent Failures to Text Policies: What Works and What Breaks


205. Adaptive Multi-Horizon Reinforcement Learning


206. SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking


207. Scaling Interpretable Transformers with Parity Bottleneck Layers


208. Frontier Financial Judgement: Can agents tell what might move a stock?


209. Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents


210. RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring


211. When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers


212. Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans


213. Bayesian uncertainty estimation improves clinical decision making in medical AI agents


214. Geometric Configurations of Perturbed Jailbreak Prompts


215. Joint Utilization of Geospatial and census proxies for Autoencoder-Assisted Downscaling (JUGAAD) of socioeconomic indicators in India


216. StabilityBench: Benchmarking Instability in LLMs


217. Monkey King Bang: A Unified Scientific Multimodal Foundation Model


218. SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction


219. Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design


220. SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales


221. When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion


222. Improving Access to Essential Medicines via Decision-Aware Machine Learning


223. HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws


224. From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime


225. Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling


226. Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches


227. ReliableTableQA:How Much Supervision Does Reliability Annotation Need?


228. A Graph Neural Network approach to zero-shot Digital Twins


229. Grounding Investor Views: Neural Predicates in the Black-Litterman Model


230. Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development


231. CLOE: Christoffel Loss Autoencoder for Anomaly Detection


232. Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement


233. Scaling Closed-Loop Feature Channel Configuration with LLMs


234. The Active Ingredient in Muon’s Grokking


235. PhantomFill: When the Form Demands an Answer, Language Models Invent One


236. Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation


237. Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study


238. Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?


239. THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA


240. CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents


241. Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention


242. Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc


243. RE-AD: Real-Time Requirement Adherence for Data Labeling


244. Response drift across frontier large language models


245. A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction


246. The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs


247. Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception


248. Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility


249. Preference Tuning as Spectral Update Reorganization


250. Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models


251. Making Open-Source Text LLM Watermarks Durable Against Merging


252. Break Through the Compression Bottleneck: From Theory to Practice


253. Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing


254. LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining


255. More Is Not More: What Matters for Diversity in LLM Opinions?


256. Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought


257. Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs’Hallucinations


258. From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring


259. Through-the-Earth Magnetic Induction Communication and Networking: A Comprehensive Survey


260. Deblurring in the Wild: A Real-World Image Deblurring Dataset from Smartphone High-Speed Videos