전체 AI 논문 - 2026-09-24

1. StudentBench: AI and human tutoring yield equivalent GRE learning gains


2. An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act’s Code of Practice


3. Learning the Cost of Reliable Inference


4. Shutdown Sabotage Propensities in Multi-Agent Systems


5. PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety


6. Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination


7. Discovery of fully efficient fault indicators along a data-based diagnosis process


8. SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference


9. A Resilience Recovery Method for Complex Traffic Network Security Based on Trend Forecasting


10. Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents


11. A hierarchy of faithfulness criteria for knowledge base completion


12. Reachable Global Optimization in AI Systems: How Global Is Global?


13. Agentic Governance and Adversarial Verification for Policy-Constrained LLM Healthcare Appeal Generation


14. Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM


15. Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure


16. Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis


17. Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions


18. Categorical Internalisation of Environmental Groupoids for Generalisable POMDP Solving


19. Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence


20. SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design


21. BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport


22. State-Grounded Conditioning: Wrapping User-Facing LLM Agents Where Direction Depends on Live State


23. Not What You Meant: Can LLMs Follow a Specified Negation Semantics?


24. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents


25. Emergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution


26. MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design


27. CART: Closed-Loop Adaptive Red Teaming for Large Language Models


28. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents


29. Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance


30. Stable Geometry with Divergent Task Evidence for Efficient Long-Horizon Agent Compression


31. Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms


32. Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents


33. StateComp: Learning When to Compress History in Long Horizon Agents


34. Large Knowledge Model: From Papers to a Scientific Reasoning Landscape


35. Sparse-Observation Atmospheric Thermal Forecasting with Physics-Informed Neural Networks for Climate-Aware Digital Twins


36. PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks


37. Memory Control Signals Emerge Before Action in Long Horizon Agents


38. Hunyuan-A13B Technical Report


39. EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory


40. TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent


41. DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents


42. CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments


43. XLOG: A CUDA-Native Engine for Neurosymbolic Integration


44. Enhancing Small Language Models for Power Outage Report Generation via Minimum Risk Training


45. Do We Need Complex Topology Control? Distinct-Peer Random Routing Improves Cost-Efficiency in Sparse Multi-Agent Debate


46. Provably Complete Generalized Planning with LLMs


47. Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit


48. Propose, Don’t Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors


49. Math Reasoning in LLMs is Organized by Approach, Not Topic


50. Are Stated Reasoning Steps Causally Load-Bearing?


51. Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations


52. Reinforcement Learning with Decomposed Subtasks


53. Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts


54. Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution


55. Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment


56. Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations


57. TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents


58. Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity


59. Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse


60. Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction


61. Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark


62. Order-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning


63. Agent-Editing World Model: Rethinking World Modeling for LLM Agents


64. Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model


65. Learning Holographic Reduced Representations with Clifford Variational Autoencoders


66. When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment


67. Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer


68. AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios


69. MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference


70. MemBodied: Recurrent Associative Memory for Vision-Language-Action Models


71. Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers


72. Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?


73. Do Center Biases Propagate? Robustness of Pathology Foundation Models in Whole-Slide Image Classification


74. From Agent Output to Authorized Transition


75. “We’ll Fix It Later”: Education, AI, and the Deferral of Privacy in EdTech


76. Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation


77. Field-of-View Extension in Dental Cone-Beam CT via Implicit Neural Representations and Diffusion Model-Based Refinement


78. Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching


79. Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness


80. Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark


81. LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations


82. Curriculum Learning with GNN-based Reinforcement Learning for Job Shop Scheduling


83. Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings


84. TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval


85. Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing


86. PISCES: Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather Anomaly Detection and Early Warning


87. Controlled Attribute-Specific Summarization of Interrogative Dialogues


88. Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance


89. Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices


90. Spread and Scale: What Determines Whether Test-Time Budget Allocation Pays


91. Schrödinger’s Code Repository: Have LLMs Learned SWE-bench or Memorized It?


92. RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction


93. False-science induction in autonomous scientific discovery


94. Bounded Loops: Pre-Run Spend Bounds, Proved Termination, and Verified Completion for Agent Harnesses


95. TopoGS: Topology-Aware Anchor Feature Aggregation for Large-Scale 3D Gaussian Splatting


96. A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools


97. What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis


98. Query Implied Generative Engine Optimization


99. A Non-Invasive Cloud-Based Migration Strategy for Post-Quantum Cybersecurity in Smart HVAC Systems: Architecture, Implementation, and Empirical Evaluation


100. Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis


101. Evaluating ADC-only deep learning pipelines for breast cancer detection and segmentation using standalone diffusion-weighted MRI


102. LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law


103. EidosDoc: Implicit Structure Encoding for Cost-Effective Semi-Structured Document QA


104. Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures


105. Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning


106. Hard Negatives Reveal What Easy Negatives Hide: Cross-Lingual Harmfulness Representations Degrade with Resource Tier Under Hard Negatives


107. InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies


108. AI-Driven Neural Surrogates for In Silico Design of Cognitive-Affective Neuromodulation Targets


109. The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA


110. FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation


111. InternW0: A Foundational Physical World Model for Efficient Real-World Interactions


112. Learning Local Heterogeneity and Cross-Region Context for Large-Scale Traffic Forecasting


113. Compliant AI Infrastructure for Regulated Finance: A tiered multi-agent framework with DLT audit trails for financial operations in DACH


114. InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation


115. Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality


116. When Context Misleads: In-context Learning with Jurisdiction in Large Language Models


117. Hidden not Deleted: How Networks Suppress Entangled Features


118. The Capability Manifold and ML Scaling Laws


119. DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment


120. FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration


121. TNLearn: An Open Source Python Package for Task-based Neurons


122. PhyMo: A Physical-Field Modality for Multimodal AI4Physics


123. Behaviora - A Conceptual Architecture for External and Internal Behavior of Robots and Agents


124. NV-Reason-CT: 3D Visual Language Model for CT Analysis


125. Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models


126. Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound


127. Kairos: Grounded Forecasting of Presence and Directional Flow in 4D Scene Graphs


128. Issuer-Sovereign Agentic Payments


129. BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models


130. Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration


131. What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit


132. Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech


133. Automotive mmWave Spinning Radar Place Recognition with Spatially Gated Feature-Correlation Representation


134. Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools


135. Psychoacoustically Aligned Latent Smoothing for Adversarial Robustness of Full-Duplex Speech-to-Speech Dialogue Models


136. Planned Test-Time Scaling with Coordinated Reasoning Paths


137. Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models


138. Geometry-Conditioned Visual Place Recognition in Natural Environments


139. Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction


140. Constraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems


141. Evolving Inspectable O-RAN Slicing xApps with LLMs


142. Breaking Weather-Content Coupling: Type-Severity Guided Progressive Disentanglement for All-in-One Infrared Restoration


143. Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning


144. Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls


145. Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS


146. Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation


147. KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling


148. Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition


149. The Risk-Sensitive Schrödinger Bridge: Is Not a KL Projection


150. Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR


151. Combining LLMs and Genetic Search for ARC-AGI-2


152. Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices


153. Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation


154. KATOsuper: Surrogate-accelerated neural topology optimization with sensitivity-consistent Fourier neural operators


155. Scalable Subgraph Sampling via Resistance Curvature


156. Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach


157. Self-Evolving Multimedia Verification through Memory Consolidation of Contestation Experiences


158. Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement


159. The Linear Representation Hypothesis Needs a Group Action


160. The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems


161. A Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery


162. When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense


163. Intelligence Across Embodiments


164. Local Evidence and Geometric Readout Repair in Trained GNNs


165. Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving


166. The Gaussian Is Enough: Flow-Matching Priors Do Not Help When Fine-Tuning Large Behavior Models


167. EMA: Elastic and Performance Transparent Memory Across GPUs


168. An open benchmark for machine learning-based polymer property prediction


169. Loss Choice or Model Choice? The Role of Forecast Level in Cryptocurrency Volatility Forecasting


170. Topological Signatures of Cyber-Attack Classes in Natural Visibility Graph Representations of Network Traffic


171. Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms


172. Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation


173. A 3D Pose-Based Ensemble Framework for Cricket Shot Classification and Automated Biomechanical Analysis


174. Cross-Modal Contrastive Learning from Histopathology and CT for Automated Renal Cell Carcinoma Grading


175. On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning


176. COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference


177. Ajar: Measuring Open Privilege in Agent Defenses


178. Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness


179. Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection


180. FLINT: Fast Lightweight Inference for Traversability


181. QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs


182. SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms


183. COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation


184. A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction


185. LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels


186. Spec2COBOLRot: An Agentic-AI Degradation Loop for Realistic COBOL Corpus Generation


187. Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits


188. Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management


189. What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus


190. Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection


191. Learning Stiffness Dependent Fluid Structure Dynamics from Coarse Flow Representations


192. Gödel’s and Scott’s Variants of the Ontological Argument in Lean 4


193. Attention-based representations for multi-task computation