전체 AI 논문 - 2026-08-03

1. ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction


2. Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics


3. AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers


4. DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat


5. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback


6. COntExt: Towards Context-Aware Ontology Extension from Operational Metrics


7. AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction


8. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember


9. Beyond Retrieval: Analytic Memory for Multimodal Agents


10. ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models


11. Beyond Component Testing: Validating Agentic AI Systems


12. MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation


13. Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents


14. Don’t Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL


15. MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft


16. CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents


17. Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration


18. A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation


19. On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness


20. Evidence-Grounded Constraint Checking in Construction Documents


21. MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents


22. Scaling Scientific Discovery Environments for Turn-Level Agentic RL


23. MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations


24. NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability


25. Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design


26. Fragility of Value under Imperfect Alignment


27. Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions


28. Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures


29. EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses


30. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition


31. Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks


32. Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery


33. Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO


34. ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding


35. Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support


36. How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories


37. An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents


38. Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding


39. TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter


40. ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning


41. LLM Framework for Discovering Major Mathematical Conjectures: AI’s Quest for the Next Riemann Hypothesis


42. Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review


43. OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems


44. The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations


45. CENDRe: Concept Extraction with Natural Domain Representations


46. When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning


47. A Human-Centered Validation of the Explainability-Performance Coefficient


48. FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models


49. TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning


50. MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models


51. ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation


52. TerraNova: A Foundation Model for the Anthropocene


53. From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale



55. TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion


56. QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models


57. AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair


58. Explore Beyond the Boundary Using Entropic Information


59. Dense Temporal Contrast Synthesis via Conditioned Latent Transport


60. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens


61. Cross-Lingual Transfer for Machine Translation in Turkic Languages


62. Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning


63. SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery


64. DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation


65. The persuasive power of large language models does not depend on their perceived national origin


66. Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation


67. OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation


68. CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation


69. TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation


70. RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems


71. When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration


72. Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters


73. FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution


74. Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution


75. MOSAIC: Masked Outsourcing of Secure AI Computations


76. SAF-OPD: Stable Advantage Fusion for On-Policy Distillation


77. SERUM: State Extraction and Refinement for User Modeling


78. MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation


79. MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation


80. CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning


81. ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency


82. Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory


83. Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations


84. Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership


85. HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators


86. InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation



88. DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs


89. metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making


90. Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients



92. Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives



94. Learning Lookahead Lemmas for Neural Network Verification


95. Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance


96. Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving


97. Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds


98. Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning


99. PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits


100. RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images


101. A robust association between LLM use and scientific productivity: Assessing stopping-time selection


102. Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates


103. Retrieval-Driven Training-Free AI-Generated Video Attribution


104. DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models


105. FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation


106. Gated Q-learning: Add Off-Policy Bias to Taste


107. Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers


108. TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models


109. Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth


110. Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use


111. To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing


112. RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data


113. Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?


114. TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text


115. A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging


116. Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity


117. Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing


118. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation


119. Stratified Negation in RDF Rules: A Correct Approach (Extended Version)


120. WaiT for the Signal: Simple Frequency-Aware Flow-Matching


121. SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction


122. DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing


123. A user’s guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem


124. WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization


125. Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning


126. SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM


127. Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent


128. Reflected UAS: Corrected Deterministic Stability and Direct CTMC Drift Calculation



130. Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation


131. MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer’s Disease Classification


132. LAWFUL: Law-Aligned Witness for Faithful Use of Latents


133. Guarantees on Dynamical System Distinguishability for LLM Token Generation


134. Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems


135. Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation


136. HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring in Laying Hens


137. Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration


138. COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention


139. Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations


140. ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning


141. Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation


142. Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework


143. The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?


144. The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models


145. Topology-Aware Data Movement for Disaggregated GPU Inference


146. Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students’ Collaborative Discourse in Prompt Engineering Tasks