전체 AI 논문 - 2026-08-18

1. What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models


2. Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning


3. Quipu: A Governed Bitemporal Knowledge Graph Store


4. Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment


5. When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding


6. GRIP: Grounded Reasoning via Information-Restricted Premises


7. LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing


8. FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy


9. Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents


10. Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies


11. PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning


12. Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement


13. A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory


14. Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate


15. CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction


16. Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents


17. Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I)


18. CUBICS: Situation-aware performance estimation for safety-relevant ML components


19. DeepInsight II: One Trace from Benchmark to Robot


20. Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling


21. Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation


22. JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills


23. HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents


24. Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics


25. The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach


26. Drive, Pack, Fly: The Travelling Thief Problem with Drone


27. ParaTempo: Efficient Parallel Reasoning via Temporal Confidence


28. Reasoning-supported Robustness Validation of Automotive E/E Components


29. A Policy Algebra for Trust-Preserving Agentic AI Execution


30. Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152


31. AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems


32. What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics


33. DriveCache: Action-Aware Caching for Driving World Model Inference


34. AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment


35. Process-Constituted Intelligence: A Shared Criterion for Humans and Machines


36. BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics


37. Competing at Every Price Point with Agentic Evolution over a Menu of LLMs


38. Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior


39. Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication


40. Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain


41. TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents


42. FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection


43. When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification


44. Assessing LLMs’ mathematical abilities requires understanding the various mechanisms of mathematical creativity


45. Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling


46. Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics


47. Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance


48. Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency


49. MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment


50. ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction


51. Solvable Sokoban Without a Solver via Diffusion


52. Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces


53. Augmenting Text to Increase Translation Difficulty


54. UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations


55. Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning


56. Breaking and Defending LLM-Powered Social Media Bot Detection Systems


57. Bounded Agents: Delegation Security for Multi-Agent AI Systems


58. Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation


59. CoupVisor: Strategy Optimization by Round and Challenge Decision Support


60. RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration


61. Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs


62. The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale


63. RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning


64. Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State


65. KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving


66. Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration


67. Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents


68. Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign


69. Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps


70. PLeDO: Pain Level Detection for Osteoarthritis from EMR Data


71. HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation


72. Adaptive Mixing of Policies from Searching and Policies from Learning


73. Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment


74. THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts


75. A Responsible Artificial Intelligence Framework for Groundwater Modeling


76. Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict


77. Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts


78. Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds


79. VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation


80. TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation


81. When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction


82. Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback


83. From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM


84. Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling



86. From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation


87. Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation


88. EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints


89. A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network Evolution


90. Dynamic Multi-Byte Prediction With Hierarchical Language Models


91. Mental Model Management: An Operator-Based Framework for LLM Memory


92. Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization


93. OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks


94. Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs


95. A survey of AI-generated voices and their detection


96. Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs


97. Large Language Model Assisted Operational Monitoring for Battery Energy Storage System Integrated Power Distribution Networks


98. Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish


99. TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions


100. Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL


101. Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot


102. FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA


103. UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms


104. Incoherent by Design? On the Moral Self-Consistency of LLMs


105. A concentration result for multilayer feedforward neural networks


106. The Benchmark Trap: Structures of Power and Injustice in AI Evaluations


107. Physics-informed VAE-EVT for Tail Aware Radio Map Prediction


108. MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning


109. Physiological World Models for Human State Transitions


110. Understanding Cognition-Induced Risks in Agentic AI Systems


111. Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis


112. ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning


113. $D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction


114. VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?


115. Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication


116. Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark


117. Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts


118. LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures


119. SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion


120. Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World


121. ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models


122. Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy


123. ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits


124. Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark


125. Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads


126. Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents


127. Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation


128. Validation-Frontier Representation Selection under Constrained Observation


129. StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling


130. Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems


131. Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents


132. Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning


133. LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents


134. GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG


135. TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning


136. Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research


137. SCOPE: Score-Isolated Agentic Optimization for Video World Models


138. LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning


139. Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form


140. S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices


141. Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5


142. Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design


143. T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework


144. RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder


145. Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL


146. Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid


147. When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation


148. Small Models Scout Bottleneck Order for Large-Model Data Control


149. LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks


150. Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence


151. Personalized Auto-Research: Towards a True AI Co-Scientist


152. JarvisBench: Always-on Intelligence Between Humans and Agents


153. Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning


154. What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering


155. MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment


156. Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking


157. Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning


158. Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous


159. CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs


160. Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations


161. From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving


162. Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs


163. Advanced modelling and data analytics in aviation


164. Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation


165. Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems


166. Synchronized Logit Steering: Real-world Steganography


167. A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks


168. When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry


169. Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition


170. Beyond Correctness: Toward Automated Novelty Verification with Lean 4


171. Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems


172. Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring


173. When Uncertainty Isn’t Enough: An Empirical Study of Self-Correction in Code Generation


174. Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance


175. Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks


176. Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study


177. Learning Agent Execution for KV-Cache Management in Agentic Serving


178. A Human-Centred Approach to Benchmarking LLMs for Parenting Advice


179. Large Language Models and their Awareness of Mechanics and Spatial Geometry


180. Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP



182. Position: Medical AI Neglects Real Treatment Outcomes


183. Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement


184. The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines


185. An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case


186. Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry


187. OGX: An Open-Source, Vendor-Neutral Generative AI Application Server


188. SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization


189. Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study


190. Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System


191. Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration


192. Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws


193. From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change


194. Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture


195. Position: AI Lock-In Is in Progress, and We Must Be Prepared


196. Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review


197. When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL


198. The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning


199. Large Language Models Show Metacognitive Sensitivity in Medical Reasoning


200. FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment


201. Don’t Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory


202. Improving the matrix multiplication exponent with modern optimization and AlphaEvolve


203. AutoSR: Automatic Symbolic Regression by Searching Research States


204. Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text


205. Proteus: Incremental Memory Activation for Long-Context Sequence Modeling


206. HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL


207. Model Hypnosis: Strong control of AI via additive subliminal effects


208. CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?


209. When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents


210. Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models


211. ClawGym II: Exploring Black-Box RL on Agent Harness


212. UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation


213. Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot


214. Neurosymbolic Embodied Agents


215. Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching


216. Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis


217. TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation


218. Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments


219. TDD-Agent: Test-Driven Reasoning for Code Generation


220. GoalEvolve: From Handcrafted Algorithm Priors to Goal-Driven Evolution of Physical Design Algorithms


221. Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI


222. MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter


223. Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors


224. Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors


225. UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures


226. Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank


227. Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL


228. Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning


229. X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization


230. Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents


231. Toward Better Assessment of LLMs’ Performance in Clinical Error Detection


232. When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness


233. HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes


234. Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning


235. Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity


236. VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience


237. Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization


238. When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation


239. Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans


240. MLLM-Guided Semantic Correction for Text-to-Video Generation


241. NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation


242. Graph Machine Learning: An Opportunity for Power Systems


243. RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning


244. A Two-Stage Learning PINN Approach for Solving the Inverse Problem of the 1D Porous Medium Equation


245. A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling


246. A Human-LLM Teaming Framework for Privacy Risk Analysis: An Illustration with CBDC-Based Welfare Schemes


247. Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos


248. Visualizing Uncertainty-to-Action Composition for Human Oversight


249. PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data


250. Towards Risk-free AI Agent Deployment


251. Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs


252. Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics


253. Synthetic Data Augmentation for Satellite-Based Analysis of Battle-Damaged Agricultural Fields in Ukraine


254. Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation


255. Coverage-Maximizing Multinomial Subset Routing under Operational Constraints


256. OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations


257. MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories


258. HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals


259. SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection


260. Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning


261. Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026


262. Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning


263. Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps


264. Audio-Visual Segmentation via Depth-Guided Collaborative Modeling


265. Decoupled Temporal Encoding for Generative Recommendation


266. Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic


267. Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection


268. CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills


269. Software Engineering for AI-driven Building Operation


270. A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation


271. STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering


272. HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction


273. Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology


274. Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline


275. LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents


276. Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics


277. MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems


278. Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations


279. Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm


280. QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents


281. Domain-Specific Text Embedding Models for Entity Resolution


282. Digital Twin Degradation: Detecting Cyber Physical Attacks via Temporal Inconsistencies


283. A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis


284. Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System


285. TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening


286. RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction


287. AsyTO: Asymmetric Temporal Operator for Parameter-Efficient Multivariate Time Series Forecasting


288. Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner


289. Learn What’s Left, Not What’s Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization


290. OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation


291. CAPO: Constraint-Aware Prompt Optimization for LLM Agents


292. Walk Before You Run: The Importance of Data Exploration for Data Analysis Agents


293. Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer’s Disease Detection


294. NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption


295. RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection


296. Dynamic Evidence Collection Ecosystem for Assessment Integrity and Authentic Competence


297. From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents


298. A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency


299. CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications


300. LLMs Get Smarter from Targeted Synthetic Multilingual Data


301. Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation


302. Information Geometry of Message Passing


303. Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery


304. Pre-training Visual Dexterity in Simulation


305. Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology


306. Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive


307. Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning


308. Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes


309. Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning


310. Characterising cardiac tissue properties with graph neural networks


311. CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling


312. A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations


313. Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation


314. Decomposing Staleness in Recommender Systems: A Dual-Filter Framework for Supersession and Decay


315. ALKEMIE Agent: an autonomous platform for computational materials design


316. Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection


317. TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity



319. FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction


320. Beyond Single Object: Learning 3D Relations with Large Language Models


321. RRFC: Recursive Refinement via Feedback Conditioning for Iterative Image-to-Image Generation


322. Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer


323. Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media


324. Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation


325. PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails


326. When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations


327. Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation


328. When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation


329. Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification


330. Sparse Prototype Code Underlies Classification and Prediction Across Modalities


331. Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis


332. EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input


333. FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy


334. GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix


335. Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair


336. ARENA: Automated Red-Teaming for Large Audio Language Models


337. Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study


338. Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors


339. MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration


340. Spectral Saliency for Machine Unlearning


341. EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation


342. Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability


343. Optimal Lower Bounds for Networked Information Aggregation


344. Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off



346. NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models


347. ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems


348. An Evaluation Framework for National AI Regulation


349. Invariant Pretraining for Robust Code Representations


350. FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge


351. Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees


352. Afterlife Delegation Protocol: Speculative Design of Self-Sovereign Agents that Outlive Their Principals


353. Agent Inheritance Protocol: Speculating on Feralized Agents After Principals Die


354. AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization


355. SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning


356. ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution


357. When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text


358. Logical Embeddings for Argument Analysis


359. Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning


360. MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation


361. No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage


362. PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies


363. VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments


364. VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction


365. UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection


366. CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction


367. Valhalla: A Layered Knowledge-State and Service-Governance Framework for Long-Term Scientific Knowledge Work


368. The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests


369. FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection


370. LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset


371. A Unified Backbone–Expert Framework with Relation-Token and Residual–Classifier Interfaces for Automatic Modulation Recognition


372. Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models


373. Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems


374. Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models


375. From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems


376. Fast Test-Time Refinement for Robust Learned Image Compression


377. CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation


378. Beyond Direct Access: Resource Hijacking in LLM Agents


379. WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing


380. Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning


381. GATTA: Graph Active Learning with Test-Time Augmentation


382. Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints


383. MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM


384. DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest


385. Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference


386. SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system


387. MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems


388. FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making


389. RamseyGadgets: A Graph Construction Dataset for LLMs


390. PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity


391. GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation


392. Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?


393. Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning


394. Generative data assimilation highlights fronts as key regulators of ocean energy cascade


395. Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment


396. PaSTel: Anchoring Histology in Spatial Transcriptomics via Multi-Scale Hierarchical Bio-Prior Contrastive Pretraining


397. SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable


398. Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task


399. The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency


400. Workspace Topology as an Attack Vector in Agentic Coding Assistants


401. Evaluating Agentic Code Repair Capabilities in Distributed Systems


402. Writing Style Similarity Reflects Academic Genealogy


403. Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce


404. Handover Analysis for Vehicular Communication with Explainability on the Fly


405. Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters


406. ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning


407. Artificial Intelligence as a Tool for Combating Child Labour: A Real-Time Edge Vision Pipeline for Child Detection and Age Estimation


408. NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving


409. Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening


410. Synthesizing Post-Acetazolamide Cerebral Blood Flow Maps from Baseline MRI in Moyamoya Using 3D Generative AI


411. PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models


412. Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews


413. Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks


414. A Novel Fourier Feature Network for Solving Partial Differential Equations


415. Tail-Aware Top-$k$ On-Policy Distillation


416. Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering


417. Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra


418. DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis


419. Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data


420. Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics


421. Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning


422. Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning


423. Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level


424. Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation


425. Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans


426. Information-Theoretic Causal Modelling of Semiconductor Process Dynamics


427. Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation


428. Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection


429. ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry


430. BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement


431. Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning


432. Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow


433. pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier


434. P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval


435. FedImp: Enhancing Federated Learning Convergence with Impurity-Based Weighting


436. Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms


437. Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling


438. Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings


439. SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation


440. iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge Integration


441. BDIP-Net: Dual-Interaction Graph Learning for Property Prediction of Bilayer Materials


442. Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays


443. Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network


444. DeMTS: Denoising Trajectories as Multivariate Time Series for Hallucination Detection in Diffusion Language Models


445. Characterizing Rhetorical Misalignment in Decision-Making with Language Models


446. Inference-Time Mitigation of Adversarial Political Bias in Large Language Models


447. LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review


448. Local AI pre-screening for human triple-blind peer review in health sciences



450. Explaining Reinforcement Learning Decisions in Self-adaptive Systems


451. Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion


452. DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs


453. Understanding AI Anxiety in the Workplace: A Multimethod Investigation Using Fear Acquisition Theory and the Technology Acceptance Model


454. Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents


455. Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models


456. Extend the Safety Horizon for Intelligent Transportation Systems through Semantic-Aware Cooperative Perception


457. Intelligent Base Station Deployment in Urban Wireless Networks: A Geographic Data-Informed Digital Twin Approach


458. Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures


459. Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework


460. HarmProfile: Characterizing Harmful Distributions in Frontier LLMs


461. From Reactive to Autonomous: Evolution of AI Operations in Cloud Network Infrastructure


462. WARA: Toward Automated Wireless Optimization Research with Closed-Loop LLM Agents


463. Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)


464. Orbital AI Computing: Carbon Tradeoffs Across Satellite Scale


465. A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation