전체 AI 논문 - 2026-09-23

1. CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents


2. SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving


3. Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents


4. Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It


5. The Delegation Blind Spot: Auditing Product Decisions from Agent Choices


6. Quantum-Aided Active Device Detection in Energy-Harvesting Symbiotic Radio Networks


7. Neutral-Atom-based Quantum Optimization for Resource Allocation in NOMA Networks


8. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure


9. REFLEX with Jev for Efficient Selective Control in LLM Agents


10. Reproducible AI Requires Reproducible Randomness


11. Recursive self-improvement of AI research agents


12. The Source of Disturbance Matters: External, Internal, and Control-Generated Noise in Adaptive Regulation


13. Reliability Theory for AI Control


14. Dual-Frontier: When Can an Agent Trust Its World Model?


15. FISSION: Label Augmentation for Bot Detection


16. Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction


17. Coding Agents are Strong Prompt Optimizers


18. A Multi-Timestep LSTM Ensemble regressor for Enhanced Short-Term Runoff Prediction


19. A Hybrid AI Framework for Academic Advising: Integrating Ensemble-Based Grade Prediction and a Rule-Based Expert System


20. Improved Multiplayer Bandit Algorithm for Bernoulli Rewards


21. RCShift: Certifying When Partial Linkage Suffices for Finite-Sample Decisions


22. Identifying Intelligent Processes via Online Sequential Testing


23. TREND-10K: A Comprehensive Dataset for Next-Generation Video Quality Assessment Based on Preference-Driven Media


24. EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models


25. The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke


26. Toward User-Mediated Self-Repair in Ubiquitous Robots Through Goal-Oriented Agentic AI


27. Unanimity Without Persuasion: A Single Round of Debate Erases the Disagreement That Verification Needs


28. When Verifiers Vote Backwards under Verdict Substitution: Signed Pivotal Value in Correlated Self-Consistency


29. VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning


30. The Cost of Conservation: Coordination-Memory Laws for Exact-Support Generation


31. When Big Data Becomes a Curse: Spatial Heterogeneity and the Limits of Learning from Passive Acoustic Monitoring Data


32. MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation


33. FairMon: A Tool for Monitoring and Visualizing Algorithmic Fairness


34. DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents


35. Early Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy Using Temporal Deep Learning on DWI


36. Neoadjuvant chemotherapy response prediction using pretreatment diffusion and contrast-enhanced magnetic resonance imaging with clinical variables


37. FusionMMT: A Unified Multimodal and Multitask Learning Framework for Nuclear Fusion


38. The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale


39. Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows


40. RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty


41. ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models


42. Adversarial Course-of-Action Generation: Game-Theoretic Multi-Agent Algorithms for COA matching & COA generation


43. FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents


44. Canonical locks that encode part-whole hierarchies


45. CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions


46. VideoX-Qwen: Data-Centric Instruction-Based Video Editing


47. CausalLoss-Fin: Attributing Financial-Agent Loss to Decisions and Infrastructure Faults


48. AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing


49. Prediction Is Not Detection: Evaluating Pre-Recognition Claims in Longitudinal Clinical AI


50. Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts


51. When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention


52. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks


53. Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction


54. Neurosymbolic Action Model Learning under Partial Observability


55. The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance


56. OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities


57. LingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data


58. TCMaster: Confidence-Aware Querying and Workload-Guided Physical Design for Multi-Source Traditional Chinese Medicine Knowledge Graphs


59. How Strongly Should Task State Influence an LLM Agent?


60. Toolcompass: Guiding Tool Trialing, Not Suppressing It


61. Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing


62. Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark


63. Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces


64. ChatT2: An Adaptive Framework for Developing a Large Language Model-Based Agent for Natural Product Domain Research


65. Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA


66. ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications


67. Evaluating Coding Agents on Kernel Exploit Generation


68. Transformer Heads Looking for Order


69. Gaze responses to false-positive computer-aided detection prompts during colonoscopy: a paired-video and real-time eye-tracking study


70. Direct Optimization of Generators for Search in Automated Theorem Proving


71. A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators


72. Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding


73. Weakly Supervised Quantum Error Mitigation


74. SMTB: Fast Structure-Mapping with Tight Bounds


75. Towards participatory speech dataset curation: A queer case study and conceptual framework


76. Queer inclusion in speech datasets: An audit and taxonomy of practical tensions


77. Spectra: A Rules-Driven LLM Pipeline for Automated KYC Document Processing


78. RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation


79. ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations


80. Real-Time Hand Gesture Recognition for OpenXR Using Transformer-Based Machine Learning


81. Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions


82. ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent Runtimes


83. From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI


84. Efficient Iterative Retrieval with Heterogeneous Batching


85. Robust Failure, Conservative Repair: Textual Knowledge Distillation from Cross-Model Failures


86. From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought


87. Clarification Is Not Correction: LLMs Fail to Let Go


88. Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing


89. Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks


90. Learned Enterprise Data Comprehension: Compression and Routing for Data Agents


91. Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass


92. When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning


93. MedGate-Fusion: Integrating First-Encounter Semantic Narratives and Physiological Biomarkers for Prospective Stroke Risk Stratification


94. The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis


95. Lean Pool: An AI-Maintained Archive of Formalized Mathematics


96. X-Planner: Event-Structured Task Planning for Embodied Intelligence


97. Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings


98. An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival Prediction


99. 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian Splatting


100. Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN Optimization


101. Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation


102. SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue


103. A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem


104. FleXray: Universal Clinical X-ray Segmentation


105. Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen


106. Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows


107. The Sirens’ Song: When Proximal Background Context Overshadows Distant Evidence


108. TraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits


109. Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning


110. Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning


111. Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation


112. From Alignment to Access Control: A Framework for GenAI Policy Enforcement


113. A Spectral Theory of Grokking: Weight Decay induces Feature Learning


114. Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models


115. Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference


116. Towards Hierarchical GNNs for multi-grid power flow: generalization across operating scenarios


117. Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models


118. The Disciplinary Language Transfer Problem: How Psychological Vocabulary Produces Governance Failures in AI Agent Deployment


119. Topology-Stratified Materials Discovery with A Flow-Based Generative Model


120. A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation


121. Do Vision Model See Like the Brain? A Comparison Across EEG Encoding Model


122. The Ethics of Artificial Intelligence in Military Operations


123. Radiomics-Conditioned Modulation of RenalCLIP Features for Clear Cell Renal Cell Carcinoma Classification


124. When Recursive Models Finish Computing


125. Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing


126. FeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation


127. PP-Net: A Hybrid Physical-Prior Neural Network for Scattered Light Removal in Biomedical Images on Embedded Devices


128. Complementary Roles of Radiomics and Foundation Representations in Renal Cell Carcinoma Classification: A Comparative Study of 2D and 3D CT Encodings


129. DeepFEAv2: Deep Learning for Transient Finite Element Analysis Beyond Structured Meshes


130. QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation


131. TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series


132. MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation


133. FairMean: Promoting Fairness in Distributed Learning under Label Poisoning Attacks


134. GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement


135. PACT: From Credit Assignment to Critic Alignment


136. TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling


137. Geometry-Aware Hyperbolic Residual Quantization


138. TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models


139. CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference


140. On the security and privacy of LLMs in Mobility


141. AIGC Video Detection based on the fusion of spatial-frequency-optical flow multimodal features


142. ABAI at COLIEE 2026 Task 1: Multi-Stage Retrieval with GraphRAG-Enhanced Meta-Learning, and a Post-Hoc Study of the Cross-Validation-to-Test Gap


143. Reducing Hallucinations in Large Language Models Through Integrated Self-Verification and Retrieval-Augmented Generation


144. When Unpaired Sets Support Shared-Corruption Calibration: Moment Geometry and Two-Sample Precision


145. WatchPoint: Executable User Feedback for Real-World Agentic Web Development


146. Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs


147. Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems


148. Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression


149. Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models


150. The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain


151. MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning


152. TTTIR: Unlocking Instance-Specific State Evolution via Test-Time Training for Image Restoration


153. StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots


154. Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods


155. TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models


156. Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication


157. Observing the Conduct of Systematic Reviews with Generative AI Support: An Experience Report from a Graduate Software Engineering Course


158. CricRAG: Retrieval Augmented Vision-Language Models for Personalized Cricket Coaching


159. EMERGE: Resolution-Agnostic Point Cloud Generation with Equivariant Graph-Based Diffusion


160. xWhyL: Causal Interactive Learning


161. REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States


162. Reciprocal Collaboration: how lessons from convergence in GLAMs can enhance interdisciplinary AI research


163. Compiling Sufficient Governance Context from Declared Losses and Reachable States: Exact Observation-Contract Synthesis with Cardinality and Cost Objectives


164. Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models


165. SE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schrödinger Bridges


166. Interweaving Marginals into Multivariate Sample Paths: Training-Free Dependence Construction for Probabilistic Time Series Foundation Models


167. Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction


168. Toward Responsible AI-Augmented Cyber Defense: Pattern Recognition, Defense-in-Depth, and the Case for Human-AI Collaboration


169. BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception


170. Risk-Aware Online Conformal State Probing


171. Evaluating the Effectiveness of SechKAN on 1D Data


172. In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization


173. CogenPVG: Cognitive-Enhanced Reflective Multi-Agent Framework for Persuasive Video Generation


174. MorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models


175. You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs


176. Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models


177. Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models


178. Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation


179. Beyond Class Marginals: Bounding Rehearsal Gaps without Freezing Class Co-occurrence


180. Self-Supervised Combinatorial Optimization with Constraints via Frank-Wolfe


181. Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages


182. Interpretable AI plus Handheld, Portable Retinal Photographs: A Low-Cost Glaucoma Screening Solution for West Africa


183. Teaching Reinforcement Learning and Humanoid Robotics to High-School Students: An Expert-Validated Curriculum Design on a Low-Cost Open Platform


184. From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs


185. When Quantum Meets AI: Quantum Methods for Machine Learning and Machine Learning Methods for Quantum Systems


186. An Exploratory Replica-Overlap Probe of the Grokking Transition


187. What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation


188. Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus


189. EMGBlend: Heterogeneity-Aware Self-Supervised Pretraining for Gesture and Force Decoding


190. RootQuantV2: Adapting a Vision Foundation Model for Root-Trait Regression from Minirhizotron Imagery


191. AkasicMEM: Governed Enterprise Memory for Agents


192. IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models


193. DefaultGNN: A Dual-Perspective GNN Framework for Predicting Corporate Default from Buyer-Seller Transaction Networks


194. A JEPA Recipe for Tabular Foundation Models


195. Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference


196. West-WRF AI 2-km: High-Resolution Prediction of Integrated Vapor Transport and Precipitation


197. Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training


198. Universal Fractal Natural Language Decision Map: Real-Time Edge Triage Across Heterogeneous Domains


199. RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models


200. Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining


201. A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization


202. Transformer-Informed Trajectory Optimization for Relative Motion in Cislunar Orbits


203. Lightweight Ranking Heads: Accelerating Multi-Task Experimentation in Production Recommender Systems


204. Predictive Uncertainty for Neural CAE Surrogates


205. Beyond Natural Language: An Agent-Native Language for Autonomous Science


206. Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing


207. Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development


208. PICPIs: Prediction-Interval-Conditional Prediction Intervals


209. VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models


210. Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes


211. How Children Design and Reason about Trustworthy AI Chatbots


212. Trains but Doesn’t Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers


213. Indirect tipping: a social attack surface in AI agent populations


214. GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings


215. From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health


216. Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction


217. Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction


218. Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning


219. Benchmarking Neural Defend ARCAS 1B: A Foundational Multimodal Deepfake Detection Model


220. Exposing Blind Spots in Deep Imbalanced Regression Evaluation


221. Towards Sustainable Magnetic Resonance Imaging: Insights from long-term, high-resolution energy recordings across an entire scanner fleet


222. Brain-Inspired Hierarchical Modularity for General Continual Learning


223. Stable Unsupervised Continual Chunking with Sheaf SyncMap


224. The Probabilistic Structure of Large Language Models


225. Impact Is Not Invalidation: Ask About the Claim, Not the Diff


226. WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation


227. Rachel: A general-purpose language model directs and revises retrosynthetic routes


228. You’ve Seen Enough: Quality-Constrained Image Coding for Machines


229. Federating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach


230. Physics-guided deep metric learning with continuous time embeddings for open-world radar pulse de-interleaving


231. LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay


232. Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures


233. LLM-Driven Training-free Location-Attribute Synergic Fusion: A Closed-Loop Paradigm for Dual-source Encrypted POIs and LULC Mapping


234. FrontierMath Erdős


235. Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione


236. Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation



238. “As a Language Model…”: Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It


239. Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models


240. Financial sentiment analysis using FinBERT with application in predicting stock movement