전체 AI 논문 - 2026-08-26

1. Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses


2. SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL


3. FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs


4. A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments


5. Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA


6. Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core


7. StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments


8. CAFE: Self-Improving Search Agents Need Co-Evolving Feedback


9. Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought


10. StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing


11. Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav


12. RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons


13. Meta$^n$: Recursive Self-Improvement through Emergent Depth


14. Lifted Model Construction under Approximate Commutativity


15. Confident at the moment of action: belief miscalibration in LLM play under hidden information


16. The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models


17. Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning


18. Causal Modelling of Support Interventions for Student Competency Assessment


19. Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms


20. PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos


21. Joint Optimization of Tool Creation and Use for Large Language Model Agents


22. EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents


23. When “Must” Becomes “Maybe”: Constraint Weakening in LLM Agent Workflows


24. Discovering Adaptive Transmission Programs for Collective Innovation


25. Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models


26. PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents


27. Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites


28. Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling


29. HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning


30. Mahalanobis-Based Multi-Head Attention for Complex State Propagation


31. A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads


32. Partial Identification under Causal Orders by Linear Programming


33. A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation


34. ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping


35. Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs


36. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use


37. Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems


38. The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents


39. Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning


40. SonarLLM: A Native Sonar–Optical Multimodal Large Language Model for Underwater Perception


41. Can a Dynamic Internal Field Govern a Transformer’s Cognition? Certifiability, not Superiority, in Homeostatic Compute Control


42. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight


43. OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning


44. VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models


45. ReproAgent: Contract-Guided Paper-to-Code Reproduction


46. RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards


47. Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling


48. Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding


49. Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing


50. Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks


51. SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction


52. STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation


53. TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models


54. Evaluating Multiple LLM Generations with Validated Task Coverage


55. Constraint-Guided Enterprise Data Mapping with Large Language Models


56. MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG


57. Preference Data Selection for Mitigating the Alignment Tax in Large Language Models


58. Paritok-4B: Intent-Conditioned Context Compression for Coding Agents


59. Task-Adaptive Rubrics for GUI Reward Modeling


60. OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses


61. Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping


62. AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL


63. Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing


64. ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation


65. Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments


66. EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals


67. AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval


68. Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression


69. Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems


70. Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment


71. Relative Time Intervals Representation for Word-level Timestamping with Masked Training


72. Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding


73. Reflection with Action-Induced Visual Differences for Desktop GUI Agents


74. Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing


75. Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction


76. Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning


77. Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling


78. When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs


79. Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design


80. More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving


81. Recursive Agentic Reasoning


82. More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight


83. Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors


84. Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining


85. MARS: Multi-Specialist LLM Relay System for Competitive Programming


86. PROOF-Gen: From Optimized Data to Better Distillation


87. Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment


88. Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems


89. BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification


90. Provenance Guided Incremental Learning Under Evolving Concept Definitions


91. AI Finds A Way


92. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors


93. Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications


94. In-Context Inpainting for Time Series Forecasting



96. SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models


97. Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention


98. A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification



100. Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware


101. AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace


102. Do LLMs Understand Limit Order Book Dynamics?


103. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment


104. Automata from Agent Traces: Failure and Next-Step Prediction


105. Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering


106. MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models


107. Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value


108. FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare


109. AI Agents Push Humans Out of the Loop


110. How much of a measured AI preference is the model, and how much is the instrument?


111. Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes


112. Function-Level Execution Feedback for Code Preference Optimization


113. TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery


114. A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts


115. LLM Agents Perform Controlled Experiments Using Simulation Models


116. ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence


117. RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation


118. LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training


119. Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows


120. Automatic Model Card Generation Using an LLM


121. Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy


122. Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks


123. Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations


124. The RAT: A Unified Bayesian Model for RAG Evaluation


125. Method, Mind, and Morality: How People Make Sense of Artificial Intelligence


126. Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets


127. Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity


128. Deep Learning Super Resolution for Satellite Cloud Mask Downscaling


129. Constrained Hyperparameter Optimization for Streaming Data


130. On-policy Distillation with Verifiable Reward


131. Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration


132. Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems


133. A Literate Programming Environment for Human and Machine Agents


134. Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks


135. $\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions


136. Across the Loss Landscape with Progressive Growth


137. COCI: Conference Organisers and Content Identifier


138. StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment


139. FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment


140. LumiXAI: A Modular Full-Stack Framework for Feature Attribution


141. Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR


142. When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study


143. Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning


144. Evaluating Deep Multivariate Imputation Models on Wearable Device Data


145. Multilevel Fair Allocation under Additive Preferences


146. Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning


147. Markerless Pose Estimation for Resistance Training Technique Assessment


148. Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs


149. FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision


150. Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis


151. Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning


152. SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling


153. Contrastive Branch Policy Optimization


154. ‘Ghaib in Translation’ aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with ‘Missed-in-Urdu’ Scores in LLM Hate Speech Detection


155. Tlow: Flow-based Item Tokenizer for Recommendation


156. Preference Optimization for Non-Verbal Vocalization Synthesis


157. LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes


158. Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection


159. PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment


160. From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender


161. Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking


162. TransPhy: Visual In-Context Learning for Physically Grounded Image Editing


163. PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control


164. Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection


165. MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes


166. Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents


167. PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding


168. When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs


169. ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal


170. Mechanistic Circuit Identification for Controllable Data Generation


171. VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference


172. Don’t Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding


173. Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models


174. Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings


175. ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning


176. What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions


177. IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views


178. WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents


179. SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding


180. Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation


181. The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem


182. RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation


183. Evaluating Language Models on Cross-Language Code Functional Equivalence


184. NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution


185. The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses


186. STAIN-FL: Stealthy Targeted Attack Injection with Contextual Triggers in Federated Learning


187. Luce: Relightable Gaussians for 3D Asset Generation


188. QML for Quantum Sensing under Measurement-Induced Information Loss


189. RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding


190. Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs


191. Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory


192. A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource


193. A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization


194. Revelation Control


195. Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)


196. Coronavirus Optimization Algorithm: A Success-History Adaptive Evolutionary Framework with Archive-Assisted Search and Stagnation Recovery for Global Optimization


197. Automated Synthesis of Cloud Emulators


198. ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork


199. Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization


200. Infant Care Video Dataset for Classification of Interventions Using Transformers


201. Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation


202. Place, Slice and Schedule: Hierarchical O-RAN Control of a Tethered mmWave UAV-gNB


203. Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring


204. Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders


205. LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology


206. Restoring Without Forgetting: Continual Learning Across Image Degradations


207. EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis


208. When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk


209. Disentangled Skill Representations for Predictive Human Modeling


210. What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development


211. TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers


212. EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$


213. Too much of a good thing – when knowledge distillation promotes overfitting, and how to avoid it


214. The Limits of Automatic Evaluation of Creativity in Large Language Models


215. Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model


216. From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers


217. Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap


218. Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling


219. Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail


220. ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents


221. Macro-Operator Generation and Predicate Selection for TAMP Operator Learning


222. When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs


223. Identifying Latent Declarative Representations of Code for Assisting Repository Migration


224. Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal


225. REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring


226. Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers


227. A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision


228. Progressively Learning Heterogeneous Skills in a Unified Latent Space