전체 AI 논문 - 2026-05-29

1. Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software


2. SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations


3. Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection


4. Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents


5. Demystifying Data Organization for Enhanced LLM Training


6. MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection


7. ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure


8. mcp-proto-okn: Natural-language access to open scientific knowledge graphs through the Model Context Protocol


9. When Should Models Change Their Minds? Contextual Belief Management in Large Language Models


10. Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit


11. Double-Edged Sword or Sharp Tool? Designing and Evaluating Triadic LLM-Teacher Collaboration for K-12 Writing at Scale


12. Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance


13. BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders


14. Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents


15. Temporal Stability and Few-Shot Prompting in Math Task Assessment


16. Anchorless Diversification for Parallel LLM Ideation


17. AgentSchool: An LLM-Powered Multi-Agent Simulation for Education


18. Enhancing Multi-Agent Communication through Attention Steering with Context Relevance


19. VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing


20. PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers


21. Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison


22. Conformal Certification of Reasoning Trace Prefixes


23. Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers


24. Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection


25. Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning


26. Teaching Values to Machines: Simulating Human-Like Behavior in LLMs


27. RAISE: RAG Design as an Architecture Search Problem


28. From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs


29. KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning



31. Accelerating Constrained Decoding with Token Space Compression


32. Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent


33. Meta-Programming for Linear-time Temporal Answer Set Programming


34. Formalizing Mathematics at Scale


35. MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization


36. Make LLM Learn to Synthesize from Streaming Experiences through Feedback


37. Its All About Speed: AIs Impact on Workflow in Music Production


38. Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment


39. On the Geometry of Games and their Solvers


40. Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories


41. Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation


42. OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields


43. OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation


44. Quantifying and Optimizing Simplicity via Polynomial Representations


45. Harnessing non-adversarial robustness in large language models


46. PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing


47. AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security



49. MEMENTO: Leveraging Web as a Learning Signal for Low-Data Domains


50. SkillsInjector: Dynamic Skill Context Construction for LLM Agents


51. Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk


52. Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations


53. From XXLTraffic to EvoXXLTraffic: Scaling Traffic Forecasting to Sensor-Evolving Networks


54. LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs


55. Benchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Models


56. Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence


57. Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering


58. Uncertainty-Aware Transfer Learning for Cross-Building Energy Forecasting: Toward Robust and Scalable District-Level Energy Management


59. NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs


60. BitTP: The Lightweight Trajectory Prediction Model with BitLLM for Edge-Devices


61. Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling


62. FHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Series Inpainting and Forecasting


63. Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability


64. NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs


65. Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems


66. GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents


67. TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation


68. PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?


69. Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured Health Text Generation


70. LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning


71. VikingMem: A Memory Base Management System for Stateful LLM-based Applications


72. Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures


73. Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models


74. HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering


75. Mind-Omni: A Unified Multi-Task Framework for Brain-Vision-Language Modeling via Discrete Diffusion


76. FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification


77. GPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activity Chain Generation


78. DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning


79. Planning with the Views via Scene Self-Exploration


80. ParaTool: Shifting Tool Representations from Context to Parameters


81. Battery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter Estimation


82. Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification


83. UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents


84. DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation


85. MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs


86. Xetrieval: Mechanistically Explaining Dense Retrieval


87. The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF


88. VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data


89. CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials


90. Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation


91. ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control


92. When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs


93. Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark


94. Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization


95. EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics


96. MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models


97. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet


98. PassNet: Scaling Large Language Models for Graph Compiler Pass Generation


99. ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression


100. Rubric-Guided Process Reward for Stepwise Model Routing


101. Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models


102. Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces


103. CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval


104. Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies


105. When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop


106. Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling


107. OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories


108. Provably Secure Agent Guardrail


109. DenseSteer: Steering Small Language Models towards Dense Math Reasoning


110. Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teachers and Students Using AI


111. Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth


112. Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility


113. BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents


114. GTA: Generating Long-Horizon Tasks for Web Agents at Scale


115. ReasonOps: Operator Segmentation for LLM Reasoning Traces


116. Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents


117. Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction


118. Governing Technical Debt in Agentic AI Systems


119. The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models


120. PRO-CUA: Process-Reward Optimization for Computer Use Agents


121. Beyond Consensus: Trace-Level Synthesis in Mixture of Agents



123. The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure


124. The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane


125. Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics


126. Robust and Efficient Guardrails with Latent Reasoning


127. Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching


128. Differentiable Belief-based Opponent Shaping


129. Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence


130. Mind Your Tone: Does Tone Alter LLM Performance?


131. When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis


132. Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild


133. BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation


134. VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis


135. Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes


136. Orthogonal Concept Erasure for Diffusion Models


137. Review Arcade: On the Human Alignment and Gameability of LLM Reviews


138. Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems


139. The Cognitive Categorical Transformer: Category-Theoretic Inductive Biases for Language Modeling


140. Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction


141. Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction


142. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion


143. LLMSurgeon: Diagnosing Data Mixture of Large Language Models


144. Unlocking the Working Memory of Large Language Models for Latent Reasoning


145. GPIC: A Giant Permissive Image Corpus for Visual Generation


146. Reasoning with Sampling: Cutting at Decision Points


147. RoboWits: Unexpected Challenges for Robotic Creative Problem Solving


148. On Language Generation in the Limit with Bounded Memory


149. In-Context Reward Adaptation for Robust Preference Modeling


150. Gram: Assessing sabotage propensities via automated alignment auditing


151. Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion


152. Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes


153. Archon: A Unified Multimodal Model for Holistic Digital Human Generation


154. City-Mesh3R: Simulation-Ready City-Scale 3D Mesh Reconstruction from Multi-View Images


155. MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings


156. Self-Trained Verification for Training- and Test-Time Self-Improvement


157. Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments


158. Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection


159. LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback


160. PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions


161. How LoRA Remembers? A Parametric Memory Law for LLM Finetuning


162. Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models


163. Reinforcement Learning with Robust Rubric Rewards


164. Do Language Models Track Entities Across State Changes?


165. Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning


166. Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization


167. BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models


168. Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency


169. HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime


170. What drives performance in molecular MPNNs? An operator-level factorial benchmark


171. Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection


172. CalArena: A Large-Scale Post-Hoc Calibration Benchmark


173. iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis


174. Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms


175. On Distributional Reinforcement Learning in Chaotic Dynamical Systems


176. Neural Network Verification using Partial Multi-Neuron Relaxation


177. Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?


178. Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies


179. DAMEL: Dual-Axis Multi-Expert Learning for Class-Imbalanced Learning


180. PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding


181. Beyond MSE: Improving Precipitation Nowcasting with Multi-Quantile Regression


182. No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval


183. Evolving Features vs Evolving Entire Trees with GP for Interpretable Survival Analysis


184. xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR


185. When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems


186. How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency


187. A Predictive Law for On-Policy Self-Distillation From World Feedback


188. Projectional Decoding: Towards Semantic-Aware LLM Generation


189. REPOT: Recoverable Program-of-Thought via Checkpoint Repair


190. Masked Diffusion Modeling for Anomaly Detection


191. Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage


192. Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models


193. Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation


194. Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders


195. Test Time Training for Supervised Causal Learning


196. VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies


197. Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas


198. Genetically Aligned Patient Representations Improve Hematological Diagnosis


199. Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations


200. Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots


201. Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction


202. HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding


203. CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving


204. Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs


205. Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents


206. Selection Hyper-heuristics Can Automatically Adjust the Learning Period to Optimally Solve Pseudo-Boolean Problems


207. Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents


208. Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate


209. LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training


210. CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation


211. Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering


212. Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension


213. Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions


214. Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation


215. ESPO: Early-Stopping Proximal Policy Optimization


216. HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization


217. CB-SLICE: Concept-Based Interpretable Error Slice Discovery


218. Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Models


219. Inferring Code Correctness from Specification


220. Data filtering methods for training language models


221. Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems


222. Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning


223. Energy-Aware NECO for Single-Pass Pixel-wise Out-of-Distribution Detection in Semantic Segmentation


224. A unified deeplearning framework for contrast-phase-specific virtual monochromatic imaging



226. The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer


227. Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies


228. Personalized Turn-Level User Conversation Satisfaction Benchmark


229. From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration



231. Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content


232. OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning


233. The Sample Complexity of Multiclass and Sparse Contextual Bandits


234. Predicting Causal Effects from Natural Language Queries using Structured Representations


235. Entity-Collision: A Stratified Protocol for Attributing Retrieval Lift in Agent Memory


236. COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings


237. DLM-SWAI: Steering Diffusion Language Models Before They Unmask


238. Learning Context-Conditioned Predicate Semantics via Prototype Feedback


239. Training Deliberative Monitors for Black-Box Scheming Detection


240. Brain-IT-VQA: From Brain Signals to Answers


241. VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models


242. Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization


243. SCOPE: A Lightweight-training LLM Framework for Air Traffic Control Readback Monitoring


244. GiPL: Generative augmented iterative Pseudo-Labeling for Cross-Domain Few-Shot Object Detection


245. GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing


246. Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection


247. KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing


248. Network Optimization Aspects of Autonomous Vehicles: Challenges and Future Directions


249. Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation


250. Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities


251. The New Pro Se: Generative AI and the Surge in Federal Civil Self-Representation


252. AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling


253. PhoneWorld: Scaling Phone-Use Agent Environments


254. Evolutionary Rule Extraction from Corporate Default Prediction Models


255. MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery


256. Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles


257. SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing


258. Composing Non-Conjugate Factor Graphs with Closed-Form Variational Inference


259. Honest Lying: Understanding Memory Confabulation in Reflexive Agents


260. Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset


261. Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment


262. Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs


263. How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions


264. How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions


265. SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents


266. AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing


267. DELOS: Detecting Shallow Transits in Kepler Photometry Using a Contrastive-Learning Framework


268. Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning


269. The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction


270. Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge


271. GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models


272. On the Optimizer Dependence of Neural Scaling Laws


273. Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies


274. TRACER: Persistent Regularization for Robust Multimodal Finetuning


275. SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow


276. Does Distributed Training Undermine Compute Governance?


277. Rethinking FID Through the Geometry of the Reference Dataset


278. GrepSeek: Training Search Agents for Direct Corpus Interaction


279. MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs


280. Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models


281. Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts


282. LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation


283. Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Level QA


284. Causal Label Recovery in Payment Networks


285. Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits


286. KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs


287. DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents


288. Extreme dynamic symmetry enables omnidirectional and multifunctional robots


289. OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources


290. Wait! There’s a Way Out: A Decision Mechanism for Forecasting Conversational Derailment


291. BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference


292. Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children’s Data


293. Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents


294. Stochastic Lifting for Generating Trajectories of Stochastic Physical Systems


295. Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback


296. TIMEGATE: Sustainable Time-Boxed Promotion Gates for Continual ML Adaptation Under Resource Constraints


297. Sustainable Metal-Organic Framework Water Harvesters in the Artificial Intelligence Era



299. Domain-Informed Representation for Evolutionary Sieving in Integral and Module Lattices


300. Evolutionary Refinement of Generative Graph Topologies: A Hybrid WGAN-GA Approach


301. Parallax: Parameterized Local Linear Attention for Language Modeling


302. CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control


303. Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Optimization


304. Real-rootedness of the Poincaré polynomials of $\overline{\mathcal M}_{0,n}$: an AI-assisted proof


305. SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation


306. Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback


307. Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving


308. When and How Long? The Readout-Mediator Angle in Temporal Reasoning


309. A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router


310. unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning


311. GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization


312. OISD: On-Policy Internal Self-Distillation of Language Models


313. Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG


314. Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text


315. SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers


316. Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning


317. Label-Free Reinforcement Learning via Cross-Model Entropy


318. LoRe: Adaptive Interaction-Evaluation Routing with Per-Step Interaction Budgets for Iterative Graph Solvers


319. FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks


320. Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening


321. The Hamilton-Jacobi Theory of Deep Learning


322. Comparing Post-Hoc Explainable AI Methods for Interpreting Black-Box EEG Models in Depression Detection


323. Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization


324. Conf-Gen: Conformal Uncertainty Quantification for Generative Models


325. CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models


326. First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope


327. AIRGuard: Guarding Agent Actions with Runtime Authority Control


328. Hallucination Detection-Guided Preference Optimization for Clinical Summarization


329. Quantum-Enhanced Adversarial Robustness in Artificial Intelligence


330. Context Distillation as Latent Memory Management


331. GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human


332. LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis


333. Representation Alignment Rests on Linear Structure


334. Balancing Multimodal Learning through Label Space Reshaping


335. TaxDistill: Improving Metagenomic Taxonomic Annotation via Distilled Genomic Foundation Models


336. PrismFlow: Residual Dynamics for Flow Matching in Time-Series Generation


337. Continuity and Ordinality Matter: Constraining Time Series Tokens for Effective Time Series Analysis with Large Language Models


338. Emergent Semantic Representations in World Models through Physical Interaction without Linguistic Supervision


339. Self-Play Reinforcement Learning under Imperfect Information in Big 2


340. Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?


341. GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models


342. Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning


343. How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines


344. Specialty-Specific Medical Language Model for Immune-Mediated Diseases


345. SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation


346. No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand


347. GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling


348. Assessing Dutch Syllabification Algorithms and Improving Accuracy by Combining Phonetic and Orthographic Information through Deep Learning


349. Transcribing Children’s Speech: ASR Performance and Obtaining Reliable Orthographic Transcriptions


350. A comparative study of transformer-based embeddings for topic coherence


351. S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering


352. Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation


353. Aryabhata 2: Scaling Reinforcement Learning for Advanced STEM Reasoning


354. Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models