전체 AI 논문 - 2026-09-18

1. An Empirical Study of Harness Design for Coding Agents


2. RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents


3. Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure


4. Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models


5. Ownership in AI-Assisted Everyday Tasks


6. PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen’s Interval Relations


7. Limits of Confidence in Diffusion


8. Language-model groups overstate consensus when replaying human deliberation on a reasoning task


9. Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation


10. FreqCondNorm: Towards Cross-domain Predictive Maintenance through a Frequency-Conditioned Transformer Foundation Model


11. SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness


12. How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents


13. SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback


14. The Organization of Inference: Information, Resource Constraints, and AI Production


15. Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model


16. A Qualitative Model for Reasoning about Path and Support



18. NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction


19. Diagnose, Recover, Certify: Task Readiness under Hidden Dynamics Changes


20. AgentPProf: Semantic Profiler for Long Horizon AI Agents


21. JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations


22. AI-Driven Real-Time Relay Optimisation in Smart Urban NR-V2X Networks via Learning-to-Optimise Graph Neural Networks


23. When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority


24. JointMatch: A Unified Heterogeneous Graph Neural Solver for Large-Scale Ride-Sharing Matching


25. Sequential Contextual Fit Predicts Human Behavioural and Neural Dynamics Across Domains


26. PaGNet: A Panel-Aware GBDT–Neural Network for Multi-Target Corporate Tax Avoidance Proxy Forecasting


27. MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents


28. Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents


29. Solving Minimum Span Antibandwidth and Cyclic Antibandwidth Labeling Problems


30. UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning


31. A Proposal for an Agentic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces


32. Tailored to you: longitudinal effects of personalising language models


33. Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference


34. FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity


35. WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement


36. MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution


37. DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models


38. Can Data Attribution Filter Out Subliminal Learning? Not Reliably


39. FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction


40. Geopolitical Divisions Across Languages in Large Language Models


41. E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews



43. Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs


44. Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics


45. MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation


46. Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models


47. From “Who Is This User?” to “What Does This Purchase Mean?”: A Deployed Pipeline for Semantic User Profiling at Bank Scale


48. TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives


49. Physical knowledge on historical data matters more than enforcing physical constraints on the forecast


50. Reproducibility is not construct validity: LLM measurement of institutionally situated communication


51. Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles


52. A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents


53. MetaRTL: Meta-path Attention Enhanced Relational Table Learning


54. Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization


55. Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy


56. Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems


57. Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary


58. TorchCraft: Unified binder design by inverting an all-atom structure predictor


59. Rethinking Multi-Agent Collaboration: When More Is Less


60. AutoData: Agentic Search for Pre-training Data Selection


61. LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents


62. FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA


63. When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models


64. Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions


65. ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI


66. Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs


67. From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization


68. SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership


69. Continual Enterprise World Model Discovery in Dynamic Systems


70. Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks


71. When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent Résumé Screening



73. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems


74. EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data


75. An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence


76. LLM-as-an-Improver: Turning Verification into Better Candidates


77. QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training


78. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models


79. Compositional Reasoning in Language Models under Reinforcement Learning Post-Training


80. The syntax and semantics of goals


81. Closed-World Resolution Against Tool Hallucination in LLM Agents


82. MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs


83. Do AI Agents Understand Computer Architecture?


84. Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses


85. What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis


86. Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer


87. What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks


88. BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research


89. Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes


90. Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation


91. Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision


92. FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations


93. Paint-Anything: Unified Any-Color Control for Image Generation and Editing


94. ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis


95. Quantifying Overclaiming Propensity in Frontier LLM Agents


96. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning


97. Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations


98. GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies


99. Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights


100. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation


101. Large Language Models as Falsifiers for Cyber-Physical Systems


102. Don’t Mask the Environment: Observation Supervision Changes How Agents Explore Under RL


103. HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface


104. Chronicle: Cut-Point Replay for Regression Testing of LLM Agents


105. A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies


106. Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape


107. Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control


108. Mitigating Retaliatory Algorithmic Collusion in Repeated Games


109. Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain


110. Edustories: A Collection of Real-world Case Studies from Classroom Practices


111. greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI


112. Deep Learning-Based Classification of Cognitive and Resting States Using Electroencephalography Signals


113. Fingerprinting Multimodal Large Language Models


114. A Mathematical Model of Motivated Emotional Mind - Cognitive Embodied System


115. When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain


116. SCGFM-ART: Amortized Relational Transport for Structure-Centric Graph Foundation Models


117. TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation


118. Stress-testing Alignment Midtraining


119. Xeno-Interpretability: Investigating the Alien Minds of LLMs


120. Accelerating Sharded Data Parallelism at Scale with Federated Learning


121. STR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite Networks


122. LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality


123. Human and AI-generated texts between modal logic and statistics


124. Labeled Incidence Structures for Native Transformer Modeling of Text, Knowledge Graphs, and Hypergraphs


125. A Multi-Objective Optimisation Framework for Corticomuscular EEG-EMG Pair Selection in Hybrid BCI


126. A Hybrid Gaze-Motor Imagery BCI Framework for Effective Decision Communication


127. CleanVideo: Adaptive Concept Erasure for Text-to-Video Diffusion Models


128. Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment under Staggered Treatment Adoption


129. Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning


130. Self-complementary completions on six vertices


131. Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data


132. Scene-Conditioned Relation Routing for urban cellular activity forecasting


133. Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study


134. SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems


135. VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots


136. FacetCRS: Multi-Faceted Preference Learning for Pricking Filter Bubbles in Conversational Recommender System


137. QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization


138. Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge


139. Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants


140. Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models


141. AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair


142. Local Sparsity Enables Unsupervised LLM Safety Detection


143. Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis


144. A Scalable Trust Discovery Architecture for the Internet of Agents


145. MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards


146. Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition


147. PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation


148. Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection


149. AI Should Facilitate Democratic Deliberation at Scale


150. The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents


151. Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression


152. Astronex-World 1.0: Real-Time Interactive World Model Foundation


153. Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems


154. Dynamic Generalized Gromov-Wasserstein Optimal Transport


155. EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning


156. AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models


157. Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification


158. MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation


159. Efficiently Distributed Federated Learning


160. KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms


161. Learning and Transferring Closed-Loop Robot Software


162. ClashBench: Conflicts Leading Agents to Seize and Harm


163. PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces


164. Zarya: A Hybrid Autoregressive–Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference


165. A Functional Pilot for Certified Freshness-Aware Semantic–Spatial Range Retrieval


166. PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance


167. Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision


168. Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies


169. Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles


170. CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection


171. Long-horizon autoformalization of a core theorem underlying MIP* = RE



173. Learn Your Own Thoughts: Abstract Token Curriculum


174. SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes


175. Self-Evolving Search Index


176. DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education


177. Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction


178. TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation


179. DeltaSelect: Affordable A/B Testing for Coding Agents


180. Form Over Content In Gradient-Based Data Attribution Methods


181. Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents


182. CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives


183. Large Language Model Agents for Evidence Based Genetic Disease Severity Classification


184. A Multi-Modal Generative Model for Tomato Disease Leaves Understanding


185. Compressed Active Subspaces for Scalable Bayesian Inference


186. Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication


187. AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation


188. CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions


189. For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances


190. Efficiently Linking Unstructured Data for Multi-step Reasoning


191. From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning


192. Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models


193. From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation


194. Efficient Nash Equilibrium Computation for Cybersecurity Games


195. Riemannian–Lorentz Fusion of Vision Transformers and State-Space Models


196. LinePilot Digitizer: Line-Plot Recovery with Manual and Automatic Calibration


197. How to Guide Your Language Flow


198. Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment


199. Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning


200. AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment


201. GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning


202. Why Pretraining Fails to Share Cross-Lingual Knowledge


203. Physics-Informed Hemodynamic Modeling for Data-Free Prediction and Sparse-Data Assimilation


204. Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices


205. Robust Conformal Intrusion Detection via Traffic-Aware Calibration and Attack-Orbit Invariance


206. YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers


207. The AR Fairness Metamodel: A Structured Framework for Fairness Measures


208. PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows


209. Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers


210. Layer-wise Curriculum Learning for Efficient LLM Compression


211. Not All Nodes Are Created Equal: Homophily-Aware Stratification for Stable GNN Evaluation


212. MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators


213. REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception


214. Code-as-Auditor: Executable Compliance Reasoning via Regulation-to-Code


215. Optimal Transport Metric Learning for Feature Alignment in Partially Supervised Segmentation