전체 AI 논문 - 2026-08-24

1. VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences


2. Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation


3. Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets


4. From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry


5. AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization


6. CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment


7. Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems


8. Ontology-supported AI Model and Dataset Management


9. Enhancing LLMs in Predictive Political QA with Semi-Structured Data


10. Personalized Privacy Control in LLMs via Attention Head Intervention


11. SENTRY: Deterministic, Intelligent Risk Assessment for IT Change Management


12. From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics


13. Root cause analysis via difference graph discovery from linear time-series data


14. Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda


15. ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models


16. When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge



18. CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models


19. The Cost of a Physics Prior Is Bounded by the Ablation Gap


20. Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts


21. Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance


22. Don’t Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents


23. Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models


24. Deep Learning Models Also Recall Features


25. Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry


26. TreeWY: Speculative Verification for Gated DeltaNet Hybrids


27. Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning


28. TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming


29. The Logic of Machine Self-Preservation


30. No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators


31. Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control


32. UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists


33. ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries


34. Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems


35. MGAL: A Multilingual Granularity-Aware Long-Context Benchmark


36. RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation


37. TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding


38. Foundation Models for Partial Causal Identification


39. Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress


40. Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence


41. Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context


42. SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers


43. Dynamic Context Scheduling: Learning Beyond the Static Universe


44. Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation


45. Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation


46. Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring


47. CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting


48. Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization


49. Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design


50. Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis


51. Continuous-Time Quantum Walks based Graph Neural Network


52. ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation


53. Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol


54. DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning


55. VortexChat: An agentic framework for autonomous multi-objective integrated photonic design


56. CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery


57. Why2Speak: Faithful Reasoning for Abstaining Action Policies


58. DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents


59. Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance


60. Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions


61. Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents


62. SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL


63. Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work


64. Dual-Cache Latent Space Communication between Heterogeneous Language Models


65. Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills


66. Difficulty-Aware Semantic-ID Optimization for Generative Recommendation


67. FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth


68. Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation


69. Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning


70. Volumetric Radiology AI in the Era of Multimodal Large Language Models


71. FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning


72. A Temporal Planning Approach for Intelligent Flood Response


73. Lost in Translation: How Universal Ethical Values Fail to Translate Across Global Contexts


74. Terminal Agents: A Survey of AI Agents in Command-Line Environments


75. STCO: Conditional Neural Operators for Time-Dependent PDEs


76. Who Delegates to AI? Evidence from 53,000 Agent Configurations


77. Categorical AI phenomenology: A first-person approach


78. StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models


79. World models of environment, agent and joint agent-environment systems


80. When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory


81. Environmental Slow AI: Design Principles for Generative Systems


82. Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory


83. Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness


84. Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles


85. A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications


86. Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification


87. PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure


88. SDAD: Spec-Driven Agentic Development for the AI-Native SDLC


89. Primal Acceleration of Newton’s Method


90. AI with Authority, from Application to Silicon


91. TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems


92. Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning


93. EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering


94. Adapting Knowledge Graphs for Behavior Denoising in Sequential Recommendation


95. Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking


96. Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration


97. Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness


98. No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation


99. Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset


100. Anchored Regularized Direct Least Squares (ARDLS): Integrating Established Prioritization Operators for Priority Elicitation in the Analytic Hierarchy Process


101. DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization


102. SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control


103. Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds


104. AID-Guard: Stateful Authorization for Delegated Agent Effects


105. HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization


106. Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence


107. A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans


108. CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents


109. Atom Learning Model (ALM): how a real classroom got tokenised


110. ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents


111. A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration


112. Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems


113. PromptResponse: Optimizing Prompts for LLM Coding Tasks


114. TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics


115. AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images


116. CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors


117. $Z^2$-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN


118. CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment


119. Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge


120. Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models


121. Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models


122. WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving


123. Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making


124. Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric


125. Vibe Coding and Web Application Security: A Twin-Prompt Study


126. Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs


127. Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight


128. MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation


129. Source-Free MT Evaluation Is Not MT Evaluation


130. Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization


131. KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs


132. BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP


133. Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment


134. STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction


135. Scaling Muon for Diffusion Transformers


136. When Generated Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception


137. TRACE: Training-time Report-guided and Clinically Ordered Concept Editing


138. Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation


139. Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders


140. CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models


141. Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding


142. CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models


143. Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting


144. PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering


145. Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation


146. Identity-Aware Human-Object Interaction Motion Captioning


147. Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes


148. Lightweight Adaptive ReduNet via Hyperspherical Manifold Learning


149. C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination


150. Amplifying the imaging power of digital sky surveys with space telescopes data and generative AI


151. RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction


152. One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation


153. Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic


154. ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection


155. AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale


156. When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation


157. JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification


158. Testing and Evaluation of Agentic AI Systems In Military Command and Control


159. Aggregate, Don’t Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity


160. Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents


161. Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising


162. ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations


163. Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes


164. An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy


165. Making Deployments Safe at Meta: Health Checks for Continuous Change-Safety


166. Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology


167. AEGIS: Preventing Cross-Domain Resource Abuse in MCP


168. Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach


169. Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources


170. An LLM agent for end-to-end computational materials discovery


171. ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib


172. Approximate Homomorphisms and Convergent Representations in Transducers


173. BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers


174. From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing


175. Six misconceptions about large language models: A minimal model and diagnostic taxonomy


176. Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility


177. LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine


178. Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI


179. Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants


180. Ansari: A Retrieval-Grounded Islamic AI Assistant – Architecture, Deployment, and Lessons from 140,000 Conversations


181. Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions


182. Infrared Hotspot-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Under Mechanical Abuse


183. EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators


184. VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models


185. Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance


186. When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems


187. ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora


188. A Hybrid Edge Cloud Digital Twin for Welfare-Constrained Control in Poultry Production


189. Trilingual Topic Modeling of Sri Lankan Parliamentary Debates


190. Hadith computational science in the age of large language models: a critical narrative review


191. Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure


192. Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration


193. ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models


194. NeuroStrata: An Electroencephalographic Connectivity-Aware Deep Representation Learning Framework for Dynamic Brain Network Analysis of Mental Stress


195. The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP


196. How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel


197. Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality


198. Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing


199. Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias


200. When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots’ Safety Risks for Generation Alpha