전체 AI 논문 - 2026-09-01

1. OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques


2. When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning


3. BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing


4. Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations


5. Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data


6. Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization


7. Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence


8. Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores


9. Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents


10. MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents


11. Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration


12. CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models


13. Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation


14. CAER: Causal Action Effect Reweighting for World Model Training


15. Predicting Residential Rents in Dakar Using Machine Learning


16. VFR-Audit: Verdict-Level Reliability for Fairness Audits in Hospital Length-of-Stay Prediction


17. HSRM: Hidden-State Reward Models for Test-Time Verification


18. SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents


19. Which Rules Matter Now? Policy-Centroid Routing Before an Intelligent System Acts


20. Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models


21. Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization for Sentiment Analysis


22. ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents


23. MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning


24. HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving


25. PyKEEN-NSX: A Modular Framework for Static, Dynamic and Schema-Aware Negative Sampling in PyKEEN


26. Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning


27. Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle


28. TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI


29. AdaPath: Query-Adaptive Path-Finding via Path-Bank for Multi-Hop Implicit Biomedical KGQA


30. GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns


31. Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis


32. DiffPDE: Masked Diffusion Language Models as PDE Solver


33. Learning-Assisted Congestion-Aware Route Scheduling for Semiconductor Fab Material Control Systems


34. ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions


35. CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework


36. CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target


37. EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents


38. From Metaheuristics to Exact Methods: A CP-SAT Approach for Multi-Objective Healthcare Workforce Scheduling


39. DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark


40. Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models


41. Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation


42. Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence


43. Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents



45. Answer Probing-Guided Search for Diverse Solution Exploration of LLMs


46. Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents


47. SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning


48. Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs


49. LLM-Based Knowledge Graph Completion Combining Discrete Structural Coding with Similar Entity Information


50. CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding


51. Rethinking the Test-Time Prompt Tuning Objective from the Perspective of Calibration


52. LaMoC: Loss-Aware Modular Compression for LLMs


53. SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature


54. FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation


55. A.X K2 Technical Report


56. VERA: Authority-Preserving Edge Revocation for Federated AI-Agent Workflows


57. Game-Agnostic Value Functions through Automatic JSON Feature Extraction


58. Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide


59. Spec2Twin-Chain: Orchestrating Bi-Level Optimization with LLMs for Blockchain Digital Twin Construction


60. Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks


61. Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation


62. Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula


63. Interpreting and Steering for Safe and Correct Code Generation


64. Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models


65. AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning


66. An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecasting, and On-Chain Fraud Detection


67. EDGE: Engine for Deterministic Graph Evaluation through Conversation Simulation from Graph Structured DSL Configuration


68. Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records


69. SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking


70. Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval


71. AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies


72. On the Instance Hardness as a Decision Criterion in TinyML Systems


73. Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization


74. FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production


75. PAGE-RAG: Provenance-Aware Graph Evidence Promotion for Fixed-Budget Multi-hop Retrieval-Augmented Generation


76. Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment


77. Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems


78. LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge


79. Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security


80. Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines


81. Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation


82. Can escalation channels redirect reward hacking toward defect disclosure?


83. Toward Latent Language Model Skills Steering and Optimization: An Empirical Study


84. EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants


85. Evaluating Tiny Recursive Models Across Training for Code Generation


86. FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents


87. Reviving our data foundations is the most disruptive step to data maturity


88. TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning


89. LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO


90. Plant-Inspired AI: Plants as Inspiration for Novel Problem Formulations, and Two Case Studies


91. APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes


92. Cross-Relational Preference Learning for Better LLM Instruction Following


93. Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images


94. BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations


95. Formal Concept Analysis with Three Types of Negation


96. Predicting Future Organ Dysfunction in ICU Patients Using Temporal Convolutional Networks on MIMIC-IV Data


97. Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling


98. MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs


99. Understanding Deep Learning via Entropy Space Theory


100. EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation


101. RACER: Reinforced Agent Collaboration for Explainable Reasoning on Knowledge Graphs


102. Dynamic Important Example Mining for Reinforcement Finetuning


103. GuardianAgent: Policy-Conditioned Risk-Adaptive Anonymization with Verified Adversarial Escalation


104. Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge


105. Localizing Emergent Failures in Agentic AI: Recovering Minimal Repair Families via Counterfactual Replay


106. Computational Depth Measurement in Thermographic Video: Overcoming Spatial Overfitting via Spatio-Temporal Decoupling


107. Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation


108. Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling


109. Benevolent Bias in Multi-Turn Human-Agent Dialogue


110. How Identity and Opinion Shape Political Sycophancy in LLMs


111. An Explainable Coherence Score for Detecting Temporal Inconsistencies in Political News


112. JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning


113. More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning


114. APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows


115. Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges


116. Emergent Misalignment Is Not Magical


117. Clustering as Approximation by Constrained Projectors: Theory and Guarantees


118. SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models


119. EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation


120. HANIA: Planner-Guided Multimodal Graph Evidence Selection for Grounded Question Answering


121. Nested Convex-Body Chasing for Online Optimization with Evolving Feasible Sets


122. Revolutionizing Turn-by-Turn Navigation with Cloud-Edge Deep Learning


123. Agent2UCB: Agentic System for Generative Engine Optimization


124. Let Prompts Bridge Defense Knowledge: Transferable Graph Purification via Vulnerability-Aware GPL


125. EmoLASP: Emotion Recognition with Language Models and Answer Set Programming


126. Learning to Follow In-Context Watermark Instructions via Self-Distillation


127. Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs


128. Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis


129. Frequency Selective Neural Networks as a Foundation Architecture for Time Series Learning


130. Multi-Step Forecasting of Grape Berry Temperature based on LSTM Model with Feed-Forward Attention


131. Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free


132. Agentic AI uncovers conserved cross-tissue protein co-abundance programs inaccessible to single-dataset analysis


133. Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents


134. The Role of Network Topology and Opponent Information in Shaping Cooperation in Multi-Agent Reinforcement Learning Systems


135. From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction




138. Automated Researchers Can Reliably Mitigate Alignment Failures


139. Oculi: A Conversational Agentic Platform for Automated Credit Risk Analysis


140. MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft


141. Evaluating the Hidden Costs of Personalization in Large Language Models


142. Discovering Machine Correlates of Consciousness


143. Explainable Artificial Intelligence (XAI) in Computational Pathology: Definitions, Taxonomy, and Recommendations


144. Capability-Stratified Degradation in Ternary Language Models


145. Enhancing SAE-based Steering via Neighbor Integrated Feature Selection


146. Efficient Geothermal Well-Control Optimization via Diffusion-Surrogate Reinforcement Learning


147. PermitGPT: A Unified Generative-AI Pipeline for Construction Hazard Forecasting, Permit Prediction, and Community Impact


148. Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference


149. Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification


150. ORDDAR: Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive Recovery


151. FRAC-MAS: A Safe and Explainable Multi-Agent System for Fracture Diagnosis


152. A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration


153. How Language Models Choose Sides: Internal Representations of Instruction Hierarchy


154. Self-Specialized Teachers for Domain Post-Training


155. BiasMix-Finance: Post-Generation KYC Guardrails for LLM Portfolio Advice


156. From Extraction to Governed Memory: Multi-Agent Knowledge Graph Construction with Domain-Expert Review


157. Reward-Oracle MCTS for Formal Theorem Proving: Sample-Efficient Search and the Need for Kernel-Level Proof Auditing


158. Self-Evolving Skills via Surrogate-Guided Solve-and-Reproduce


159. AI Scientist Mission Control (AIMC): Visual Analytics for Human Oversight of Autonomous Scientific Discovery


160. AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment


161. CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science


162. CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence


163. Machine Learning-Enhanced Tabu Search for Tactical Wireless Network Design


164. Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values


165. InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal


166. TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback


167. RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences


168. MedTVL: Harnessing Vision and Language for Medical Time Series Classification


169. C3-UniMM: Causal Cycle-Consistent Unified Multimodal Modeling via Super Alignment and Shared Decoding Space


170. Integrating Triaxial IMU Sensors and Ensemble Learning for Effective Parkinson Disease Severity Classification


171. Leveraging Generative AI to Design Accessible Interactive Visualizations for Undergraduate Mathematics: A Six-Phase Workflow


172. SHAPE of Chain-of-Thought in Math Reasoning


173. CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis


174. The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys


175. Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation in Applied Sciences


176. The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification


177. From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics



179. A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making


180. Expert-validated STEM QA


181. DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation


182. SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies


183. Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification


184. LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering


185. Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents


186. Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring


187. One note in three: a verified census of three deployed AI scribes, and the instrument that counted it


188. LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It


189. Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning


190. Evaluating and Improving LLM Self-Modeling


191. MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI


192. CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations


193. CogEvol: Towards Efficient and Reliable Learning Environment Generation


194. A Universal Context-Reuse Layer for Cross-Model KV Sharing


195. LOCI: A Locator-Critic with Refinement Loop


196. Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse


197. MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative Music AI


198. LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation


199. Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper


200. Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning Methods


201. Safety Screening for Voltage Control in Active Distribution Grids via Distributionally Robust Conformal Screening


202. Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models


203. Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice


204. Exponential random graph models with soft clique constraints


205. TAMI: Temporally Aligned, Missingness-Aware, and Interpretable Multimodal Fusion for Mental Health Assessment in Older Adults with Mild Cognitive Impairment


206. Pretrained, Curriculum-Tuned, and Ensembled: A Tracer-Aware Interactive Segmentation Pipeline for AutoPET V


207. Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis


208. Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling


209. A Composition-Aware Pretraining Framework for Geospatial Foundation Models


210. Aggregate Disambiguation Systems


211. Conjoint Audio-to-Spikes Encoding and Processing for Efficient Neuromorphic Speech Recognition


212. On the Prospects of Dynamic LLM Conversations in Software Development


213. Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval


214. Calibrating Small Language Models for Claim Check-Worthiness Detection


215. RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation


216. BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks


217. RailSyn: Diagnosis-Guided Image Generation for Traceable Data Completion in Railway Foreign Object Detection


218. SingProbe Technical Report


219. An Agentic Retrobiosynthesis Framework with Learned Frontier Selection


220. Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning


221. Learning Materials Properties from Scarce Labels and Unlabeled Crystals


222. LCoT-GV: Graph Attention Networks for Verifying Long Reasoning Chains in Large Language Models


223. CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy


224. Fine-Grained Multi Image Object Hallucination Benchmark


225. BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs


226. GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning


227. Cost-efficient Active Learning for Referring Image Segmentation and Grounding


228. Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text


229. Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training


230. Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster


231. DiffSAC: Diffusion-guided Sampling for Consensus-based Robust Estimation


232. Collapsibility of Performance Metrics in Clinical Predictive AI


233. Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs


234. Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval


235. TSExplorer: An interactive data annotation and exploration tool for time-series data


236. Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems


237. Lot Machine: Multimodal Lot Extraction from Auction Catalogs


238. Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability


239. Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework


240. ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation


241. Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer


242. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions


243. Towards Cognitive Process-Aware Proactive Writing Support


244. ObjectSplat: Improving Mesh Fidelity and Interactivity for 3D Scenes via Object-Level Mesh Splatting


245. Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols


246. SePArate: Segmenting Patterns from Defects in Wafer Manufacturing Using Weak Supervision


247. ImageCAS-X: a dataset and benchmark for coronary artery segmentation and centerline extraction in coronary CT angiography


248. SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation


249. TopGQ: Fast GNN Post-Training Quantization Leveraging Topology Information


250. Using Grounded Theory for Agent Behavior Analysis at Scale


251. PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning


252. DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving


253. PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies


254. Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation


255. Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends


256. Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering



258. One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread


259. Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs


260. ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation


261. CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration


262. BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning


263. Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization


264. Centering before Pruning: Lightweight Geometry Correction for Diversity-Based Visual Token Pruning in LVLMs


265. Using Prosody to Predict Syntactic Structure


266. Stratified Consistency Distillation for Natural Language Formalization


267. Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs


268. Motus2: A Self-Evolving General World Model for Dexterous Manipulation


269. The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce


270. Label Semantic Expansion via Label Guided Neural Topic Modeling


271. SIR: Self-improving Red-teaming for Compute Use Agents


272. Science sandboxes measure the scientific capability of AI agents


273. CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation


274. E-SENS: Exclusion-Sensitive Penalization for Negative-Constraint Retrieval


275. VIBE: Video Instruction-aligned Background music gEneration


276. TPR-Attention for Combinatorial Generalization


277. Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving


278. Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators


279. AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP


280. Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear Optimization


281. Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer


282. How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account


283. Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark


284. Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction


285. Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection


286. “Act Like a 5th Grader” is Not Enough: Bounding Knowledge in LLM-Based User Simulators


287. Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models


288. TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models


289. Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models


290. Generating Clinical Vignettes that Preserve Cognitive Formulations


291. Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models


292. Training-Free Action Correction for VLA Model Failures via Language Feedback


293. The Policy Deficit in AI x Social-Emotional Learning Research


294. Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents


295. Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization


296. Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?


297. IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil


298. When Less is More: Understanding When Token Filtering Helps and Fails in AI-generated Text Detection


299. INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event Prediction


300. REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling


301. R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment


302. SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video


303. Higher-Dimensional Rotary Position Embedding


304. A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction


305. MedSegBenchmarker: A Raw-Count-First Framework for Controlled 2D Medical Image Segmentation Benchmarks


306. Cost-Effective Repository Exploration for Agentic Issue Localization


307. Conducting Stylistic Analysis of Paintings through an Art-History Agent


308. LLMODE: Aligning ODEs with LLMs via Gated Token Injection for Irregular Spatio-Temporal Forecasting


309. MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation


310. AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing


311. CineForge: Self-Improving Agents for Long-Horizon Video Generation


312. Memory-First Fact-Checking: A Knowledge-Graph-Grounded Multi-Agent System for Misinformation Detection


313. Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps


314. Wide Learning: Learning to Reach Evidence


315. SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs’ Robustness to Contradictory Evidence


316. HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning


317. PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation


318. Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift


319. AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies


320. Integrating adaptive human behavior into epidemic models with large language models


321. The Emergent Symbolic Structure of Artificial Neural Networks


322. Argument-Aware Semantic Alignment of Normative Texts: A Toulmin-Based Neuro-Symbolic Approach


323. On the Plasticity Collapse in Continual Machine Unlearning


324. Denoising as Projection: Constrained Optimization with Gradient-Guided Diffusion


325. Generalizable Multi-Agent Planning from Signal Temporal Logic Specifications via Diffusion


326. Applications of Risk Science to AI Fairness Evaluation: Principles, Challenges, and Best Practices


327. MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects


328. Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance


329. Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered


330. Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation


331. Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities


332. AI Can Be Easily Persuaded in Clinical Decision Making


333. SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning


334. Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations


335. Polis: 3D Self-Supervision at City Scale


336. Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD


337. Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multipliers


338. Safe to Resume? Breaking Execution Continuity of Agent Execution via Rollback


339. Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration


340. StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue


341. SGE: Semantically-Guided Exploration for Unstructured Environments via Image-Space Waypoint Sampling


342. Hyper3-CLIP: Hierarchy-Conditioned Hyperbolic Vision-Language Training


343. Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase


344. Evaluating LLM-based AI agents integrated with materials synthesis tools: the case of atomic layer deposition


345. Improving Randomized Metric Distortion to 2.3282


346. Detecting and Repairing Hallucinations in Retrieval-Augmented Generation


347. Learning Simple Test-Time Environments for LLM Web Agents


348. When Do Larger Batches Help Scale LLM Reinforcement Learning?


349. AOI-Net: Structural Face AOI-Guided Eye-Gaze Track Representation Learning for Autism Spectrum Disorder Detection


350. RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation


351. SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization


352. Measurement Validity in LLM Cultural Alignment


353. Adaptive Multi-Branching for Shallow Decision Tree Induction


354. QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation


355. AGRICAM: A Track-Mounted Crop Pollination Monitoring Robot


356. Background-Free Objectness Learning for Class-Agnostic Detection


357. AgentLogs: A Dataset for Opening the Black Box of GitHub’s Cloud Agent


358. PokaiTrainer: Scaling Belief-State Search to Competitive Pokémon VGC


359. Rate-Coding Bundle Memory: A Unified Model of Memory and Control for Symbolic Computation in the Brain


360. Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space


361. TAAL: Mitigating Early Beam Pruning in Generative Recommendation via Temporal Autoregressive Alignment


362. Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA Challenge


363. Training-Free Hidden-State Refinement for Flow-Matching Image Generators


364. STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation


365. Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs


366. HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding


367. CGFM-Nav: Cognitive Graph-Field Memory for Semantic-Guided Lifelong Multimodal Embodied Navigation


368. Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding


369. Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks


370. Not All or None: Dynamic Construction of Target-aware Memory Graph for Conversational Stance Detection


371. DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering


372. A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability


373. Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models


374. RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos


375. The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era


376. Free Speech and Artificial Intelligence


377. Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction


378. The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal scheduling


379. From the Loss Landscape to Diverse Feature Learning in Neural Networks


380. The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice


381. ActiveAugment: Online Active Learning for Augmentation Selection in Deep Learning


382. Structured State Reconciliation for Human-AI Task Handover


383. Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation


384. Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management


385. No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus


386. The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning


387. Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs


388. A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives


389. Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis


390. Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks


391. Delegating Before Learning: Where Generative AI Sits in Students’ Professional Communication


392. Representation Learning with Quantum Signal Processing


393. Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings


394. FigMirror: Ground It, Code It, Plot It


395. A Large-scale Evaluation of Text-guided Models for Facial Editing


396. The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents


397. ASTRA - Agentic System for Ticket Resolution and Analysis


398. Peer Oversight in Collective Decision Making


399. RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction


400. Defending Wearable VLMs Against Private Attribute Inference


401. Measuring Similarity between Artistic and AI Generated Images using Siamese Neural Networks


402. GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon


403. Test-Time Scaling for Scientific Equation Discovery


404. Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?



406. Redesigning and Auditing Deep Research Writing for Faithful Reports


407. Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture


408. PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework


409. Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA


410. PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation


411. Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework


412. Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models


413. Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects


414. Asymmetric Within-Document Predictive Learning for Scientific Document Representation


415. MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson’s Disease Assessments


416. Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure


417. PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora


418. From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education


419. STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study


420. Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System


421. Parametric Multimodal User Memory: Storing What Captions Cannot Carry


422. NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts


423. PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response


424. Agent-Based Model Framework for the North Carolina Modeling Infectious Diseases Program (NC MInD ABM) Overview, Design Concepts, and Details Protocol