전체 AI 논문 - 2026-09-09

1. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents


2. A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes


3. Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails


4. ExecCritic: Learn to Test, Test to Improve for Coding Agents


5. A Generalization of Amari’s Bayesian Duality


6. MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents


7. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?


8. The Surprising Effectiveness of Approximate Value Iteration in Self-Play


9. Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training


10. Time-Varying Data as Sheaves: an Invitation to Narratives


11. Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning


12. Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths


13. Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack


14. PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving


15. SkillAdam: Stable and Efficient Skill Evolution for Agents


16. API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces


17. Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course


18. It’s All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction


19. When Can One Obtain Certificates of Optimality Using Positivstellensaetze?


20. Application of curiosity driven exploration methods for hardware interference identification


21. GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data


22. CLAMP: Constrained Decoding for Vision-Language Embodied Planning


23. Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation


24. A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation


25. AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems


26. BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents


27. Personalizing LLM Agent Memory Using Biometrics


28. SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs


29. EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering


30. Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models


31. FastE: Readout-Triggered Token Compression for LLM Embedding Inference


32. LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation


33. Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation


34. MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging


35. Three Types of Negation of Triple and its Elements and an Extension of Triple


36. Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation


37. Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems


38. CircuTutor: Transforming Static Circuit Problems into Intelligent and Dynamic Tutoring


39. Agentic ML Exploration (A-MLE) for Ads Ranking


40. zScore-N: A Neural Network for On-Chain Wallet Reputation Scoring


41. Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts


42. SE-GoS: Self-Evolving Graph-of-Skills for Skill Library at Scale


43. TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs


44. Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI


45. A Better Spur Should Start From Each Objective


46. Qiushi Engine on AstaBench E2E-Bench-Hard


47. Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference


48. Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment


49. Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models


50. Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models


51. Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits


52. OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?


53. Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering


54. WorldAgen: Unified State-Action Prediction with Test-Time World Model Training


55. SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents


56. SchemeArena: Factorized Stress Testing of Scheming in LLM Agents


57. Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training


58. Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture


59. CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information


60. RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts


61. Inference-Time Nash Alignment


62. Automated Design of Inventory Policy with Large Language Models: An Exploratory Study


63. ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?


64. Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning


65. A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate


66. From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents


67. Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans


68. Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study


69. When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability


70. From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining


71. Support Topology and Gradient Mixing in Sinkhorn Layers


72. CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows


73. Beliefs and Behavior in Language Models


74. FrogNano: Training a 4B Coding Agent via Online Task Synthesis


75. PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations


76. Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression


77. Explainable Temporal Attention-based Defect Detection For Fillet Joints in Real-Time Gas Metal Arc Welding Based on Multi-modal Data


78. Do Large Language Models Know What They Don’t Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty


79. Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging


80. What Does an LLM-Agent Leaderboard Rank Actually Compare?


81. xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems


82. When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems


83. The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs


84. A radiographic world model for clinical reasoning and evidence generation


85. The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing


86. APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents


87. Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision


88. Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best


89. AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era


90. FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?


91. A Tool-Augmented, GPT-4 Chatbot for Real-Time Repository Data Analysis


92. From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction


93. Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End


94. Quantile-Led Feature Extraction for Multi-Horizon Predictive Maintenance in Industrial Manufacturing Systems


95. Modus Tollens and Counterfactuals and Counterfactual Reasoning Based on Three Types of Negation


96. The Internal Anatomy of Strategic Choice in Large Language Models


97. CIT-CAD: Constraint Intent Tree-based CAD Code Generation and Verification


98. RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems


99. Human-like moral judgments conceal divergent motive attributions in large language models


100. AAS-RAIL: Improving Information Extraction for Asset Administration Shells through Retrieval-Augmented In-Context Learning


101. DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning – Extended Version


102. Weakly supervised neural network: segmentation of complex structures in X-ray microCT


103. World Models Under Asynchronous Sensor Observations


104. SkillAlign: Aligning Skill Interfaces for LLM-based Agents


105. Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning


106. Distance-Aware Attention and Wall-Distance Expert Routing for Transformer-Based 3D Flow Prediction


107. Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective


108. Agentic Algorithm Engineering: Improving Shared-Memory Exact Minimum Cuts


109. EmoMed: An Emotionally-Aware Agent for Multimodal Medical Support with Real-Time Information Retrieval


110. PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians


111. An Auditable Symbolic-RAG-Generative AI Architecture for Goal-Oriented Conversation Orchestration


112. Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner’s Expertise


113. EEG-Driven Decoding Framework for Passenger Hazard Perception in Highly Automated Vehicles


114. Beyond Sparse Rewards: A New Benchmark and Structure-Aware Graph Alignment for Micro-Drama Understanding


115. Risk Is Not Review Value: Wrong-Answer Exposure Under Bounded Review Budgets


116. A Hierarchical Consistency Framework for Auditing Retrieval-Augmented Generation Systems


117. VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery


118. Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval


119. RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving


120. SSP-DMGTimeNet: Physics-Constrained Learning for Spatiotemporal Trajectory Prediction of Vehicle Platoons


121. iBrain: A Unified Foundation Model Reading the Brain from Surface to Spikes


122. When and Why LLM Causal Priors Help: Closed-Loop Prior Selection for Amortized Causal Inference


123. A visual large language foundational model for medical image recognition using clinician-oriented social media


124. Learning transferable human physiology from two million hours of sleep with SleepFM-2


125. NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures


126. Formation of structural attractors in neuromorphic systems


127. Unsound Search with Policy and Value Networks in Legends of Code and Magic


128. Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool Actions


129. Reason Through the Latent! Making Latent Visual Reasoning Necessary


130. Simulating the Marginal Green Contribution of AI Modules in a Smart-Agriculture Platform: Evidence from Two Monte Carlo Experiments


131. Monte Carlo-Based Ex-Ante Assessment of the Green Benefits of an AI-Driven Smart Agriculture Platform in Hainan


132. We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consciousness


133. A Computational Implementation of a Goal-Directed Theory of Affect


134. SerenAI: State-transition system inspired by text-based world AI models


135. A Translational Note on AI Safety Evaluation


136. MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games


137. A Unified Policy Architecture (UPA): The Governance Kernel for Enterprise AI Operating Systems


138. Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification


139. From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts


140. Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction


141. AutoKD: Autonomous Knowledge Discovery


142. Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts


143. SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores


144. Customer Relationship Intelligence: Integrating CRM and MDM for Enhanced Customer Engagement


145. MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition


146. IIns-VAE+: A Robust Transfer Learning Framework for Environmental Identification in Wireless Sensing


147. Substrate-Portable Execution for Production LLM Workflows



149. SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use


150. LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies


151. Generating Instance Generators in PDDL Planning


152. Explaining AI Agents Through Execution Traces


153. DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents


154. Generator-Independent Runtime Assurance under Partial Observation


155. Agentic Pressure: The Endogenous Entropy of Reliable Autonomy



157. Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review


158. The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble


159. Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization


160. AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories


161. Learning Counterfactual World Models for Embodied Reasoning under Partial Observability


162. Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents


163. Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools


164. Exposing Weaknesses in Emotion Recognition in Conversations


165. Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration


166. Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment


167. More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review


168. Distilling Vision-Language Models for On-Device Fire Understanding


169. DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents


170. Inference-Time Graph Engineering for Multi-Agent LLM Workflows


171. From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale


172. The Normalization of Deviance in AI Development


173. Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses


174. Recovering Temporal and Geographic Signals from Language Model Embeddings


175. CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning


176. What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets


177. The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry


178. Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools


179. Planning and Scheduling Business Processes under Control-Flow Uncertainty


180. EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent


181. Deep belief networks are exact


182. EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph


183. Beyond “AI Helps Humans”: Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era


184. The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies


185. When and What to Teach: Budget-Aware Online Adaptation for Web Agents


186. Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment


187. SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction


188. SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews


189. PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories


190. RAPID: Reliability-Aware Pair Importance Distillation


191. ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models


192. Compiling VGDL into Causal Models


193. Damage-Aware Bandit Pruning for Vision and Language Transformers


194. AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents


195. When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents


196. CriticGen: Generation-Aware Evaluation as Actionable Feedback


197. Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models


198. TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model


199. NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting


200. Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs


201. DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination


202. Measuring LLM Sycophancy under Sustained Multi-Turn Pressure


203. GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting


204. ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR


205. Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics


206. Training-Free Task Vectors for LLM Behavioral Control


207. The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits


208. It Is Not My Code Anymore


209. Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks


210. Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling


211. Omni Interaction Agent Technical Report


212. GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks


213. SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL Evaluation


214. Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation


215. OntoKG-EQ: A provenance-grounded, competency-question-governed knowledge graph for auditable analyst querying


216. Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems


217. Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling


218. Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports


219. Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers


220. Adaptive Anisotropic Attention for Axis-Structured Signals


221. Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks


222. Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics


223. Hyperparameter Scaling Laws Across MoE Sparsity


224. CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling


225. Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning


226. BIFTA: Brain-Inspired Few-Shot Tactile Adaptation for Unknown Sensors


227. X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR


228. MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models


229. TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection


230. Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR


231. SUN: Reaching for Novelty in Reinforcement Learning


232. CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation


233. From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video


234. Suan: Rectifying Direct Preference Safety Alignment in Large Language Models


235. SynthRCT: Scalable Conditional Deformation Synthesis for Synthetic Repeat CT Generation


236. Leveraging contextual events on structure-aware next activity prediction


237. Neptune: An AI model for Global Ocean Subseasonal Prediction


238. The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?


239. Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-Constrained Settings


240. Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?


241. Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values


242. SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching


243. AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation


244. Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks


245. Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method


246. Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models


247. Equivariance Breaks the Learning Rate


248. IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring


249. Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking


250. RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation


251. A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing


252. Tracing Stereotypes from Representation to Output in Multilingual LLMs


253. AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents


254. Exploring Bottom-Up Clustering for Creating Semantic IDs


255. A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware


256. FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy


257. What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory


258. Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding


259. ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing


260. CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations


261. CUNO: Curriculum and Preference Optimization for Stable Graph Unlearning under Mass Deletion


262. CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning


263. 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints


264. WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos


265. AI for AI: Optimizing Additional Infrastructure Build-out to Power Artificial Intelligence Data Centers


266. Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation


267. IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA


268. KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization


269. Sparse Data Augmentation for Optimization with Provable Guarantees


270. DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning


271. LLMs for Social Network Modeling: From Network Generation to Dynamic Processes


272. SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities


273. HyCO: A Hybrid Neural Solver for Combinatorial Optimization


274. Rethinking Sign Language Translation: The Impact of Signer Dependence on Model Evaluation


275. TDDN: Text-aligned Diffused DINO Network for Puzzle Understanding


276. AVCG: A Generalized Variational Framework for Counterfactual Generation under Hypothesis Distributions


277. The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]


278. Foundation Models for Generalizable Semantic and Goal-Oriented Communication


279. SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs


280. A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM


281. Kalman Delta Networks: Uncertainty-aware Associative Memory


282. VoT: Vision-of-Thought for Unified Multimodal Representation Alignment


283. You Can’t Prefer Emotions You Don’t Sample: Intensity Undershoot in DPO-Tuned LLMs


284. Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining


285. Quantifying the Engagement Trap: Impact of Short-form Video Recommender Systems on Users with ADHD


286. CodeTD: Topology of Attention Detects Hallucinations in Code LLMs


287. Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment


288. Do AI Coding Assistants Check Before They Install? A Pre-Registered Demand-Side Audit of Trust Signals in the Research Software Supply Chain


289. TFTrack: A Template-Free Framework for Efficient 3D Point Cloud Tracking


290. An emancipatory vision for designing (generative) AI for learner flourishing


291. Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web


292. Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions


293. From Citations to Contributions: LLM-Assisted Credit Scoring of Research Articles


294. Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection


295. Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs


296. Noēsis: Deterministic-First Retrieval with Two-Tier Context Hydration for Factuality-Critical Queries on Small Local Models


297. How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement


298. Thermodynamic Cyclic Processes with Markov Samplers in Bayesian Inference


299. Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery


300. Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?


301. Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem


302. Decentralized Safe Multi-Agent Reinforcement Learning via Predictive Shielding


303. ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making


304. Mapping the Emerging Social Science of Large Language Models


305. Beyond the Matrix Sign: Quadratic Spectral Descent


306. Large-Scale User Behavior Analysis in Multimodal AI-Assisted Manual Task Execution


307. Topology Obstructs Pure Foundation Neural Quantum States


308. Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection


309. Efficient Exploration Is Enough


310. We’re Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation


311. Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining


312. Generation of Vectorized Maps Beyond Vehicle View


313. Human mutation field reveals an equilibrium-like structure with irreversible circulation


314. Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy


315. When Superpixels Fail on Documents: A Study of Segmentation for LIME Explanations


316. Latent-to-Latent Flow for Volumetric Stochastic Segmentation


317. Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection


318. TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning


319. TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables


320. Revisiting Thinning Methods for Kernel Learning Problems


321. RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting


322. Federated Binary Gating with Server-Side Vision-Language Inference for Surveillance Anomaly Classification


323. Riemannian Optimization for Multi-Player Quantum Games on Product Unitary Manifolds


324. Distributed Lag Neural Additive Models


325. Beyond Single-Negative Preference: Multi-Negative DPO for LLM-Centric Historical Entity Linking


326. BlueprintAgent: Constraint-Triggered Targeted Revisits for Simulation-Ready Generation from Scanned Structural Blueprints


327. D3ARC: Time-Critical Distributed Disaster Detection for Asynchronous Cooperative Multi-Robot Systems


328. Monadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint)


329. Staying on the Attack Path: Structured State for Long-Horizon Automated Penetration Testing


330. PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout


331. PLATOS: A Power and Latency-Aware Task-Oriented Scheduling Strategy for Healthcare IoT in Fog Computing


332. LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances


333. Quality Metrics for LLM-Generated Asset Administration Shells: A Perturbation-Based Evaluation Approach


334. Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval


335. Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval


336. MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling


337. Mathematical Programming in Machine Learning and Artificial Intelligence: A Unified Taxonomy of Models and Applications


338. Towards a Resilience-Theoretic Foundation for Adversarial Robustness in Industrial Control System Anomaly Detection


339. Parallelism Strategy Chaining for Fast Training Convergence


340. REFINE: Trajectory Representation Learning via Closed-Loop Transcription – Extended Version


341. Recompilation Is Not Enough: Test-Guided Decompiled-C Repair


342. Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families


343. FedRAW: Preserving Rare-Label Influence in Asynchronous Federated Learning


344. Tensor network representations of discrete maximum entropy distributions via mean polytopes


345. Deep Learning for Biopsy-Free Subtyping of Basal Cell Carcinoma from Dermatoscopic Images


346. In-Place Instruction Following in Diffusion Language Models


347. Mind the Approximation: Fisher-Weighted SVD Compression for ViTs


348. FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity


349. Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer


350. AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing


351. From LLM-Generated Specifications to Learned Quadruped Locomotion


352. Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models


353. AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis


354. Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection


355. Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision


356. SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem


357. MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation


358. Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning


359. Frequency Estimation Based on SNR-adaptive Frequency Estimator Under Wide SNR Range


360. LoGAN: Multilingual Font Localization with Generative Agents


361. CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records


362. ARNAI: Artifact Removal Network based on Autoencoding and Inpainting for Robust Spinal Image Segmentation and Measurement


363. AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation


364. Input-to-State Stability Framework for Fully Distributed Primal-Dual Dynamics for Quadratic GNEPs Without Multiplier Consensus


365. Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning


366. AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories


367. Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models


368. DPSF-Net: A Dual-Prior Spatial-Frequency Network for Real-World Remote Sensing Image Dehazing


369. MSSP: Multi-Scale Spatially-Constrained Partition for Unsupervised Semantic Segmentation of 3D Point Clouds


370. Mind the Phase: Effective Rank and Representation Health in Legged Locomotion


371. Steering Interference Reflects the Model’s Defaults, Not the Behavior Directions


372. PCSDiff: Diffusion-Based Bias Correction and Super Resolution Toward Practical Operational Medium-Term Precipitation Forecast


373. The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists


374. Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras


375. Constrained Online Learning with Noisy Constraint Values


376. From Synthetic Priors to Model Behavior: Structural Coverage in Tabular Foundation Models


377. Human-agent discovery of reconfigurable in-plane ferroelectric superdomain control


378. Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning


379. AutoLexSteer: Automatic Contrast Construction for Lexical Activation Steering


380. Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting


381. Characterizing Contention-Induced Reliability Collapse in KV-Cache Timing Side Channels for Multi-Tenant LLM Serving


382. You Are What You Read: Misalignment via In-Context Persona Induction


383. XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?


384. WAPP: Safe Learning of Positive Security WAF Policies from Live Traffic


385. Skynet: Workflow-Level Anomaly Detection for Agentic AI via Semantic and Structural Modeling


386. Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents


387. Comparative Study of Anatomical and Learned Features in AI Models for Structural Brain MRI


388. Hardware Trojan Threats to Multi-Chiplet Photonic Neural Network Accelerators


389. AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories


390. DrugReason: Dynamic Multi-View Reasoning over Knowledge Graph and Language Evidence for Drug Repurposing


391. AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths



393. Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction


394. Companion-style QA Assistance in Ego-Vision


395. A Trustworthy Watermarking Framework for LLM-Generated Food Safety Content


396. Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models


397. Attention-Enhanced Deep Features with Heterogeneous Ensemble Learning for Glaucoma Detection


398. Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces


399. Tracking the Moving Frontier: Long-Short Term Advantage Estimator


400. ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics


401. FSAN: Flow State Attention Network for Aerodynamic Prediction


402. Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model


403. SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration


404. Inducing Emergent Misalignment from Reward Hacks with Iterative DPO


405. When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization


406. TD-STGT: A Spatio-Temporal Graph Transformer for Mobile Traffic Demand Forecasting


407. Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus


408. MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference


409. SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives


410. Deep Barycentric Regression for Optimal Transport Map Estimation and its Statistical Optimality


411. Discovering Translation-Worthy Languages with E-Values


412. Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier


413. Certifying cooperation: a novel approach to cooperative multi-agent task generation


414. Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection


415. A TTP by TTP Approach: Precise Malware Detection via Malicious TTP Recognition


416. Recovering topological information of light by topological learning


417. SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure


418. ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language


419. OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution



421. Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing


422. One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints


423. Collision Snapshot Guided Time-Reversed Safety-Critical Scenario Generation


424. On BatchNorm Forward Modes in Value-Based Reinforcement Learning


425. Parameterized and Streaming Algorithms for Euclidean Fair $k$-Center Clustering


426. Recovering Weak Signals with Normalizing Flows


427. Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction


428. AGSA-Net: Abundance-Guided Self-Attention Network for Spectral Unmixing-Aware Hyperspectral Remote Sensing Image Classification


429. Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression


430. Steering Geometry: Validating Human Value Geometry in LLM Steering Space


431. SIDE: Sensor Impersonation Detection at the Edge via Sequence Prediction


432. It is Not Yet Another Tool: Creating and Deploying an Agentic AI Companion in a Security Operations Center


433. VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification


434. SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction


435. Scratchy: Visual-Scratchpad Multimodal Reasoning for Cryptographic Proof Generation in EasyCrypt


436. Diamond Agent: Agentic Control of Federated HPC Resources as a Service


437. Decision-Aware Suffix Prediction and Reasoning of Business Processes


438. Multiple Myeloma Lesion Segmentation on Whole-Body Diffusion-Weighted Imaging via Efficient Anatomical Anticipation and Multimodal Confirmation


439. All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs


440. From Splats to Silicon: Rethinking Computational Efficiency of 3DGS


441. ExpertLens: Visualizing Embedding Spaces for Post-Hoc Explainability in MoE Enhanced Retrievers


442. SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs


443. What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark


444. PAGR: Proof-Carrying Algebraic-Geometric Retrieval: A Quiver-, Provenance-, and Sheaf-Theoretic Framework for Grounded LLM Retrieval


445. Programmable Cellular Automata


446. VERPO: Verified Evidence Regularized Policy Optimization


447. Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education


448. Explainable Deep Learning for Price-Trade Dynamics: From Black-Box Forecasts to Effective Parametric Models


449. PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us


450. FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon


451. ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs


452. The Role of Gradient Modification in Heavy-Tailed Nonconvex Stochastic Min-Max Optimization


453. Calendar-SPCA: Interpretable Representation Learning for Multi-Periodic Electricity Consumption Profiles


454. SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness


455. Image-Scale Robustness and Visual Recognition Performance: A Cross-Architecture Analysis


456. A solution to the Erdős Problem #1040


457. Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning


458. What Does Animal Re-Identification Learn? Linear Biological Concepts and Their Origins in Visual Representations


459. Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts


460. Memory in Deep Time-Series Models


461. Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning


462. DART: Distributional Adversarial Recurrent Training for Algorithm Learning


463. Beyond Cross-Lingual Transfer: Benchmarking Propagation Boundaries in Multilingual LLM Unlearning


464. Intent Drift at SME Scale: Deployment Practice, Not Model Capability, Determines Agentic Compliance


465. Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems


466. AVSplat: Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning


467. Versioned Transitive Dependency-Closure Binding and Operation-Time Effect Governance for Agent Skills: ClosureBound


468. STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models


469. Structurally Close, Temporally Distant: Measuring Security Exposure in Long-Horizon LLM Agents


470. UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms


471. From Review to Authorization: Key-Isolated Threshold Signing for LLM Agents


472. AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection


473. What if LLMs Ate Their Words: Causal History Effects in Multi-Turn Interaction


474. Grounded and Faithful P&ID Reasoning: Constraining Vision-Language Models with Recovered Evidence Graphs


475. FACT: A Forensic Agent with Compiled Tool-Use Trajectories for AI-Generated Image Detection


476. Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs


477. SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement


478. Do Quantum AIs Dream in Paths? Path-Integral Slow Thinking through Grover Interference


479. AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents


480. Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks


481. Closed-Loop Evaluation of Bird’s-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies


482. Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models


483. Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora


484. Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding


485. GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims


486. RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems


487. Concord: A Video Relational Algebra for Cross-Modal Query Optimization


488. SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation


489. Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance


490. Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling


491. XAI-SDN: An Explainable Entropy-Guided Machine Learning Framework for Real-Time DDoS Detection in Software Defined Networks


492. Analysis of Respiratory Sinus Arrhythmia with Neural Networks


493. Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance


494. PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement


495. Full-Page Optical Music Recognition of Handwritten Monophonic Scores


496. TamilEOT: A Dataset and Model for Semantic End-of-Turn Detection in Tamil Telephone Speech


497. Adaptive Cost-Sensitive Machine Learning for Autonomous Robot Navigation Failure Prediction: When Not All Errors Are Equal


498. WolfSociety: Understanding Collective Risk from Harmful-Agent Scaling in Financial Agent Societies


499. ViT3Flow: A Test-Time Training Transformer MeanFlow for Postoperative Radiograph Synthesis in Scoliosis


500. An Agent Model Abstraction for Human-AI Teaming Cognitive Coupling