전체 AI 논문 - 2026-09-04

1. Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints


2. A Computationally Feasible Framework for Causal Probabilistic Explanation


3. Rethinking On-Policy Distillation of Large Language Models II: One Training Example


4. A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms


5. From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research


6. Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments


7. Efficient Test-Time Adaptation through Human-AI Interaction


8. The Natural Language Interaction Protocol and Standard for AI Agents


9. Environment Evolution for Terminal Agents


10. Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable


11. Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM


12. DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training


13. Spurious Advantage Hidden in GRPO


14. IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations


15. Instruction Duplication as an Inference-Time Control Primitive


16. FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models


17. InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models


18. LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening


19. The Dually Flat Geometry of Planning as Inference


20. Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing


21. Interface-Induced Trajectory Censoring


22. FiMI Banking: A Sovereign Model for Indian Retail Banking


23. More Criticism Does Not Make a Better Review: EquiReview-R


24. Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding


25. Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting


26. Value-Preserving Architectures for Agentic AI Systems


27. Lose the Order, Keep the Hierarchy: Deordering HTN Plans


28. Inferring Affective Consciousness in an Artificial Agent: A Case Study


29. Xiaomi-TabLDM: A Tabular Foundation Model Technical Report


30. STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation


31. Bioinfoysis Technical Report


32. Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations


33. Semantic Bayesian World Models


34. CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception


35. SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation


36. Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI


37. Transfiver: Human-AI Co-Inference through a Shared Editable State


38. DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions


39. Rethinking World Models for Safety-Critical Embodied Systems


40. SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation


41. Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation


42. Artificial Intelligence for Energy Optimization in Data Centers


43. Counterfactual Routing Using Integer Programming with Constraint Generation


44. Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study


45. Analysis of Prompt Engineering for Drug Toxicity Prediction


46. A computable representation of the physical laboratory enables verifiable workflows


47. KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents


48. The Attention Triangle in Audio-Video Models


49. HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews


50. GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis


51. Dalek: A Constructive Agent Machine


52. Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation


53. NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis


54. CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning


55. What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation


56. PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing


57. GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving


58. Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models


59. AutoGraphForge: Towards Automated Graph Theory Discovery


60. Beyond “Made with AI”: Visualizing Provenance Density to Mitigate the Transparency Penalty


61. Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents


62. DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents


63. Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection


64. Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation


65. A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant


66. Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory


67. Speculative Macro Commit for Faster Tool-Using Agents


68. MasterControl Seventeen Every Time


69. Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence


70. Compile by Training: Turning Natural-Language Specifications into Local Neural Functions


71. ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize


72. One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing


73. Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning


74. Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views


75. SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents


76. SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center


77. A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle


78. Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR


79. Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis


80. A Non-Formulable Theorem: A Fundamental Limit of Finite Syntactic Systems and Its Consequences for Security and AI


81. CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation


82. PatchBench: Evaluating AI Agents for Vulnerability Patching


83. TAP-Path: Task-Adaptive Structural and Token Pruning for Efficient and Trustworthy Pathology Foundation Models


84. Subspace Inference Enables Efficient Active Reward Learning from Preferences


85. When Models Edit Too Much: On the Fidelity of Minimal Code Edits


86. Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation


87. Influence of Extruded Filament Shape on Buildability in 3D Concrete Printing: A Geometry-Informed Deep Learning-FEM Approach


88. Representational alignment yields generalizable safety in language models


89. The Blind Spot in 2D Infants’ Pose Estimation:Robust Learning from Noisy Annotations


90. Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition


91. Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes


92. RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting


93. Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO


94. Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations


95. RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting


96. GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs


97. FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation


98. A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors


99. Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data


100. GazeFS: Target-Centered Gaze-Trajectory Forecasting and Stabilization from Gaze-Head History


101. The impact of phase information for few-shot fine-grained image classification


102. Witnesses Explain Anomalies


103. Free Pause Tokens


104. LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes


105. IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks


106. ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation


107. Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks


108. Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study


109. Symmetries and Causality: Causal Effect Identification Beyond IID Data


110. Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning


111. Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography


112. Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners


113. Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks’ financial statements


114. </think> Doesn’t Stop Reasoning: Analysis of Spurious CoT Termination


115. EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders


116. Test-time adaptation for speech enhancement with an autoregressive speech prior


117. ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection


118. Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation


119. FailBench: How Reliable are VLMs at Judging Robot Task Success?


120. On the Interaction Between Model Compression and Test-Time Adaptation


121. How Far Can Synthetic Data Take Thai OCR?


122. LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks


123. From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control


124. Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning


125. WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval


126. TruncGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates


127. LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL


128. Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion


129. LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues


130. Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech


131. BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI


132. Pattern Over-Generalization of Knowledge Graph Embedding


133. Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird’s-Eye Maps


134. Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings


135. When Users Don’t Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents


136. The Psychological Costs of Artificial Intelligence Adoption in Software Engineering


137. Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory


138. It’s the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories


139. TraveL: Transformer-based Multi-view Path Distributional Representation Learning


140. The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems


141. Privacy, Robustness, and Fairness Trade-offs in Federated Intrusion Detection: Geometric Indistinguishability at the Aggregation Interface


142. StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios


143. Spectral Convergence of Random Feature Method in Multiple Dimensions


144. TabScope: Question-Adaptive Scope Selection for Table Question Answering


145. Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data


146. FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience


147. Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression


148. SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking


149. ObserverBench: Testing Mechanistic Estimates for Intervention and Control


150. Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation


151. Evaluating Graph Neural Networks for Change-Criticality Classification in Maritime Navigation Charts


152. Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy


153. Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning


154. When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization


155. The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors


156. PrivateHub: Contrastive Diffusion Model for Private Sensor-Intensive Environment Data Generation


157. Privacy-Preserving Heterogeneous Multi-LLM Federated Inference for Cognitive Diagnosis


158. Reflect-SQL: A Self-Reflection Based Framework for Text-to-SQL


159. Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation


160. Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions


161. Counterexamples as Feedback for Agent Self-Correction


162. ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval


163. X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System


164. Anonymization, Not Elimination: Utility-Preserved Speech Anonymization


165. Traceable TTS: Toward Watermark-Free TTS with Strong Traceability