전체 AI 논문 - 2026-09-28

1. Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency


2. DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education


3. Multi-agent Scaling Across Disjunctive and Compensatory Tasks


4. A Flow Matching Framework for Neural Representational Dissimilarity


5. Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity


6. UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting


7. “AI is (not) the new…”: A Diagnostic Analogy Framework for Generative AI’s Cultural Impacts


8. Game Arena: Strategic LLM Evaluation in Competitive Environments


9. Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency


10. Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents


11. Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection


12. Programs-of-Layers in LLMs through the Lens of Cortical Areas


13. Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State


14. The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models


15. G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies


16. MA-WAM: Multi-Agent World-Action Model for Test-Time Planning


17. Purin: A Biology-inspired Mechanism for Artificial Neural Networks


18. DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration


19. Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution


20. Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture


21. Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation


22. Accounting for Bias Enables Sustainable LLM Evaluation


23. SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization


24. Semantic Navigation for Issue Localization in Code Repository


25. Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models


26. Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence


27. Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?


28. Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge


29. AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution


30. Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning


31. OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas


32. Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents


33. Externalized CPDAG Summaries Improve LLM Causal Deduction


34. Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders


35. Cheap, open agents make LLM pollution harder to mitigate


36. Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance


37. Same Text, Different Numbers: The Divergence of LLM-Based Measures


38. Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings


39. SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents


40. MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens


41. FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models


42. LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting


43. Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms


44. MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory


45. Self-Play Search Distillation for Large Language Model Reasoning


46. JevSoup: System-One Routing for Training-Free LoRA Composition


47. Training Graph Foundation Models on The Web Graph



49. EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting


50. TISD: On-Policy Self-Distillation with Trajectory Intervention


51. SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting


52. Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis


53. PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices


54. A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory


55. Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes


56. HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents


57. ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning


58. Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More


59. Insurance Reserve Intelligence Platform


60. HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models


61. Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication


62. Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging


63. ORCA: Evaluating LLMs on Data Science Code Translation


64. From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia


65. Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows


66. Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents


67. CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems


68. LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information


69. The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?


70. LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents


71. Audio LLMs Know When They Can’t Hear You


72. T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation


73. HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases


74. Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks


75. Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content


76. Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study


77. Benchy: towards a universal language for task-oriented AI benchmarks


78. BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering


79. Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework


80. Pretrained ASR Pseudo-labeling for Noisy Police Audio


81. Spectral Feedback for Test-Time Alignment of Protein Diffusion Models


82. Predicting Transmembrane Protein Topology from 3D Structure


83. A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods


84. Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems


85. Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol


86. When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess


87. ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?


88. Bringing AI to Autonomous Systems – From Cognition to Collective Intelligence


89. Statistical attribute alignment for black-box generative AI via output post-processing


90. Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer


91. OC-GS: Gaussian Splatting for Irregular Turntable Capture


92. Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum


93. Can You Check That? The Checkability Boundary for Local LLM Network Automation


94. ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos


95. Evaluating Cultural Awareness of LLMs for Haitian Creole


96. PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents


97. Uncertainty-Aware Federated Learning for Infant Movement Analysis


98. Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality


99. From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training


100. ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs


101. Implicit Neural Representation for Hyperspectral Video Compression


102. Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis


103. Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers


104. Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models


105. ActKV: Efficient LLM Agents through Action-Guided KV Cache Management


106. Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization


107. Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding


108. A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents


109. DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models


110. CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support


111. AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents


112. Towards VLA-Dreamer: Refining VLA Behavior Using World Models


113. Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows


114. UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning


115. Softmax Reparameterization for Output-Head Quantization


116. Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains


117. Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions


118. MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries


119. Agentic Limit Order Books: Phase Transitions and Market Impact


120. Geometric Inconsistency Localization in Multi-View Image Sets


121. Acoustic-to-Text KV Compression for Full-Duplex Speech Models


122. Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews


123. BAT-CLIP: Trimodal Alignment of Brain, Audio and Text


124. Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting


125. AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side


126. SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations


127. ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation


128. Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift


129. FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification


130. JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models


131. Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding


132. From Shortcut Learning to Discrete Neural Insertion Sort


133. Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods


134. DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models


135. Quantum Diffusion Models for Medical Image Analysis


136. DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving


137. G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation


138. Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance


139. The Linear Representation Hypothesis for Vision-Language-Action Models


140. FLIP: Final Layer Inference-Time Probing for Vision-Language Models


141. FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators


142. Does Uniform Discrete Diffusion Need Time?


143. MVVBench: Benchmarking 4D Reasoning in Vision-Language Models


144. PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem


145. OneWorld: Learning Consistent Physics Across Actions in World Models


146. Spackle: Completing Large View Single Image NVS with Adaptive Gaussians


147. Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models


148. UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound


149. Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations


150. Adaptive Pilot Selection for Unified Semantic Communication and Semantic Sensing in ISAC


151. Warned alike, AI agents avoid the less-crowded road while people take it


152. Persistent Negatives for Adversarial Black-Box On-Policy Distillation


153. Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development


154. MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation


155. Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings


156. Evaluation Is All You Need for Multi-Modal Autonomous Driving


157. XPhysICS: Cross-Physical-Domain Threat Grounding for Industrial Control Systems Security


158. Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models


159. NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation


160. Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI


161. Anatomy-Aware Dexterity-Driven Design Optimization of Surgical Continuum Robots


162. TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation


163. Werracle: Sub-Cent Intra-Block AI Reflex Oracles and Flash-Loan Circuit Breakers for EVM Smart Contracts


164. Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4


165. VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control


166. Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation


167. SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion


168. Threat-Aware Energy-Efficient Deployment for Dynamic UAV Networks: A Multi-Agent RL Approach


169. Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation


170. A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code


171. Subjects, Not Authors: The Authorship Hazard in Agentic Dataspaces


172. MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification


173. The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge


174. Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases


175. Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms


176. Auditing Latent-Space Monitors for Autonomous Driving


177. Proportional Representation in Temporal Voting with Ranked Preferences


178. Convergence guarantees for Muon: New parameter regimes and generalizations


179. Inquesto Score: A reliability Protocol For Voice Agents


180. PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control


181. Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs


182. CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production


183. A Benchmarking Framework for Context-aware XR Interfaces


184. Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language


185. Understanding Perturbed Parameter Ensemble Sensitivities Using A Contrastive Learning Approach


186. A Unified Account of Concepts and Chunks


187. What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study


188. DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning


189. Cost-Aware Best-LLM Identification using Dueling Feedback


190. Strategic Self-Consistency


191. Adaptive multi-resolution Gaussian processes: Scalable exact inference with naturally data-sparse covariance matrices


192. Coding Agents Aren’t Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations


193. What Will Remain Human in Software Architecture? A Focus Group Report


194. Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops


195. SignTrace: Describe a Sign, Find the Word


196. SlideLab: Audience-Centered Scientific Slide Generation and Evaluation


197. Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents


198. A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models


199. A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID


200. When Does Advection-Aware Graph Nowcasting Help? A Controlled Study of Distributed Solar Ramp Forecasting with a Self-Supervised Cloud-Motion Estimator


201. ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers