전체 AI 논문 - 2026-08-12

1. Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration


2. sLTN: Structural Logic Tensor Networks


3. Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding


4. RTSKG: Building a Rail Transit Station Knowledge Graph Dataset


5. SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure


6. V-FiLLM: Verified Financial LLM Reasoning Benchmark


7. XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving


8. FedCGR: Federated Cross-Domain Generative Recommendation


9. ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling


10. IO Factory: Simulating AI-Enabled Influence Campaigns at Scale


11. ComBodied Agents: a New Paradigm of Human-Centric Agentic AI


12. Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming


13. Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction


14. EvoMem: Memory-Augmented Evolution for Code Optimization


15. ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation


16. SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation


17. Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information


18. Compositional Benchmark Synthesis for Hierarchical Human Action Recognition


19. Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution


20. Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence


21. Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory


22. REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems


23. FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs


24. VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus


25. Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome


26. Decision-Aware Approximation of Belief Functions for Evidential Combinatorial Optimization


27. Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph


28. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment


29. Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent


30. DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation


31. Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization


32. SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models


33. Measuring Semantic Abstractness of SAE Features via Nonlocality


34. MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows


35. RadFusion: Towards Threshold-Controllable Radiology Report Generation


36. MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph


37. From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents


38. GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning


39. INSIDE the Student’s Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators


40. Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning


41. Multi-Granular Rationale-Guided Molecular LLM for Property Prediction


42. Evaluating Rational Contracting in Natural Language


43. RLMOpt: Adaptive Prompt Optimization via Recursive Language Models


44. Quantum Incremental Learning with Mixed State Prototypes


45. Rationale-Guided Learning for Multimodal Emotion Recognition


46. Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning


47. Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance


48. Recovering Wasted Compute in Autoresearch Agents


49. Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects


50. Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning


51. Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models


52. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?


53. Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research


54. Hierarchical Compositionality for An Assistive AI Agent


55. Toward a Theory of Value in AI Alignment


56. Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures


57. Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability


58. Interpreting Language Model Hidden States at Scale


59. Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction


60. Beyond Decision Boundaries: Relational Geometry Attacks on Contrastive Embedding Manifolds


61. Self-evolving Agentic Customer Support System at LinkedIn


62. Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems


63. Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models


64. Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes


65. Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding


66. Edge Phoneme Recognition for Children’s Speech through Age-Aware Training


67. Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents


68. TRACE: Trustworthy Retrieval-Augmented Conversational Engine


69. Generating Attacks for LLMs with GFlowNets


70. SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents


71. The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI


72. MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory


73. CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation


74. Automating and Scaling Behavioral Scientific Research on AI Agents


75. ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models


76. Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models’ Carbon Footprint


77. MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis


78. SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning


79. Closed-Loop LLM Co-Pilots for Digital Agriculture


80. Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning


81. ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls


82. Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation


83. How to Verify Consistency of Probabilistic Claims


84. From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop


85. Attention-Path Fragility as an Uncertainty Signal in Large Language Models


86. Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting


87. Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory


88. Entropy-Centric Explainable AI for Remote Sensing Image Segmentation


89. A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa


90. 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment


91. Multiclass Sentiment Analysis for Identifying Political Viewpoints


92. Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data


93. R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video


94. Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis


95. On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation


96. Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers


97. TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation


98. ReLTEx: Reliable LLM-based Taxonomy Expansion


99. CARE: Confidence-Aware Reasoning for Reliable Medical VQA


100. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes


101. A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models


102. Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation


103. FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation


104. GitSkills: A Dataset of Agent Skills on GitHub


105. VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?


106. TACTICL: Task-Aware Compression of Tabular ICL Models


107. Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition


108. MIRA: Medical Image Reflection for Agentic Diagnosis


109. Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation


110. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse


111. Modelling Geographic Atrophy Progression using Implicit Neural Representations


112. Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation


113. BPG: Balancing Plasticity and Generalization for Domain Incremental Learning


114. Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization


115. MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams


116. The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election


117. A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem


118. Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI


119. Optimal Stopping of Self-Refining Foundation Models


120. DuplexWorld: Can voice agents help you get through the day?


121. Most biomedical publications show signs of LLM-assisted writing


122. Conversational Orchestration for Organic 6G


123. Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control


124. ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes


125. Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization


126. SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information


127. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering


128. Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics


129. Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement


130. Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving


131. MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models


132. Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts



134. DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction


135. $π$-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement


136. A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language


137. Retrieval-Corrected Conformal Prediction for Time Series


138. ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover


139. Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration


140. On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models


141. Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry


142. Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits


143. Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models


144. Rethinking Text-Based Image Retrieval in Specific Domain


145. Improving TensorSketch Using Complex Random Variables


146. Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training


147. SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning


148. Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation


149. Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models


150. Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning


151. MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection


152. Persistent Recursive Worlds Enable Autonomous Software Evolution


153. Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging


154. From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models


155. FUSE: Frame-Unified Stress Estimation from Facial Video


156. What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research


157. Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique


158. Causality Sum Rules in Conventional Scattering Matrices


159. Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry


160. Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models


161. ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation


162. Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation


163. A Single Atom in Front of a Mirror is a Universal Reservoir Computer


164. Beyond Forecasting: Recasting Volatility Control as a Routing Problem


165. MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices


166. MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model


167. Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks


168. Towards Unified Dynamic Face Landmark Detection


169. Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement


170. Narrative Keyframing for Generative Creative Writing


171. Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories


172. Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility


173. Frozen Brain-MRI Foundation Models Are Site Fingerprints


174. MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models


175. Comprendia: AI-Augmented Code Comprehension


176. Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output


177. Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies


178. Toward Human Rights Benchmarking for LLMs: A Pilot Methodology


179. TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent


180. Unsupervised Detection of Groundwater Storage Anomalies in Ghana Using GRACE Satellite Data


181. FACT: Failure-Aware Causal Training for World-Action Models


182. Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems


183. ELMER: Evolutionary Language Model that Explores and Refines


184. The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse


185. From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation


186. MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation


187. Multimodal Item Parameter Estimation using Simulated Response Probabilitie


188. Procedural Fairness Failures in RLHF from Preference Averaging


189. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4


190. Exploring Semantic Stability Across Reviews in the Linux Kernel


191. Status Association Does Not Reliably Predict Decision Leakage


192. Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds


193. Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review


194. Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons


195. UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs


196. DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents


197. Sheaf-Based Federated Representation Learning


198. Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification


199. Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection


200. Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy


201. Do AI weather models miss extremes?


202. Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator


203. Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting


204. Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems


205. HoosierHelp: Benchmarking LLM Agents for Social Service Navigation


206. Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents


207. When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning


208. How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation


209. LLM Agents Factory: Retrieval of Domain-Specific LLM Agents


210. “YES! YES! I absolutely love this insight!” Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots


211. The Gaussian-Multinoulli Restricted Boltzmann Machine: A Potts Model Extension of the GRBM