전체 AI 논문 - 2026-07-30

1. Can AI agents conduct open-ended AI research? Early evidence from two case studies


2. Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork


3. OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding


4. Linguistic Monoculture in LLM-Assisted Language Use


5. AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching


6. On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment


7. Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data


8. Belief-Guided Decision Making with Uncertainty Gating in the Game of Go


9. What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation


10. From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence


11. Property-driven Causal Abstractions for Markov Decision Processes


12. Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM


13. UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks


14. AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution


15. Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting


16. AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining


17. Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants


18. Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models


19. Evidence-Ledger Adjudication for Claim-Evidence Traceability


20. EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks


21. MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning


22. CG-World: A Large-Scale World-State Dataset and Protocol for World Models


23. CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games


24. Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?


25. TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning


26. Position: Evaluation Scores Are Perishable Knowledge Claims


27. GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure


28. GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning


29. When benchmark inferences do not compose: Projectibility in AI evaluation


30. ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science


31. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems


32. Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models


33. APEX-Accounting


34. The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making



36. Anatomy Contextualized Adaption of CT Foundation Models


37. Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark


38. DLAM: Distributional Latent Actions with Temporal Constraints


39. MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning


40. SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context


41. Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents


42. MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair


43. Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise


44. Visual Credit Audit for Multimodal Spatial Reasoning


45. SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence


46. ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection


47. CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation


48. BayesAME: Bayesian Active Model Evaluation


49. SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception


50. Progressive Multimodal Alignment for Continual Instruction Tuning


51. Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning


52. BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories


53. Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing


54. Human diversity fuels collective creativity that large language models cannot simulate or sustain


55. Hearsay: Vision-Language Medical Diagnoses Without an Image


56. Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents


57. ReCo: Reweighting GRPO Against Distributional Concentration


58. From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs


59. Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility


60. AI as Friction for Reflection Support in Ideation


61. A First Look at Coding Agents’ Compliance with AI Contribution Rules in Open-Source Communities


62. FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning


63. Crossing-Free Probabilistic K-Line Forecasts Without Retraining


64. SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response


65. SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution


66. Journey Operators for Structured Multi-Axis Composition


67. See2Think: Do Multimodal Models Really Use Intermediate Visual States?


68. MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities


69. Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification


70. Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text


71. An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI


72. Multimodal fusion of visual and morphometric features for avian bone classification


73. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model


74. Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives


75. FARI: Robust One-Step Inversion for Watermarking in Diffusion Models


76. Automated Multilabel Mpox Research Classification with Explainable Transformer Models


77. MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation


78. Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL


79. Scientific Knowledge Discovery in the Age of Large Language Models


80. Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection


81. Constitutional Midtraining: Content Presence Drives Alignment Gains


82. Physically Real-time Infrared Attack against Optical Flow Estimation Networks


83. FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows


84. FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking


85. Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses


86. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability


87. Guarding Organizations Against Malware Risk: A Novel Graph-Based Malware Detection Method


88. Understanding Context Sampling in TabPFN on Small Tabular Datasets


89. WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models


90. Living-Harness Is an Interactive-Agent Evolver


91. Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution


92. A Physics-Informed Framework for PID Tuning of Chemical Processes Using Large Language Model Agents


93. One Run Is Not an Idea: The Implementation Lottery in Automated Research


94. Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models


95. Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks


96. ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform


97. A Persona-based Rate Action Index


98. AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control


99. Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression


100. The Art of Not Forgetting A Local Learning Architecture for Continual Learning


101. A Graph-Native Bitemporal Memory Store for Conversational AI Agents


102. HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models


103. Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning


104. LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving


105. Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection


106. PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories


107. ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models


108. Mergeable Model-Side Aggregation States for Long-Context Language Models


109. Reinforcement Learning on Cost-Constrained Quadrupedal Hardware


110. FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing


111. Voice Memory for Agentic Speech Recognition


112. Misalignment Has a Personality: A Big Five Account of Emergent Misalignment


113. Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction


114. Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment


115. Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text


116. Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning


117. High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption


118. Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research


119. When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses


120. Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks


121. Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach


122. StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents


123. SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation


124. AgentGUI: An Interface for Observing and Steering Long-Running AI Agents


125. Entity Resolution in Practice: Lessons from a Self-Serve Pipeline


126. Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection


127. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech


128. Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring


129. Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges


130. (EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations


131. Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition


132. A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment


133. Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels


134. Try Again, Don’t Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models


135. GPT-Red: Automated Red Teaming via Self-Play at Scale


136. TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions


137. Weight and Height Estimation from a Single Human Image Captured in the Wild


138. A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models


139. Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate


140. FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents


141. IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations


142. IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval


143. GuidedRAG: Semantic Steering of Retrieval-Augmented Generation



145. AI Security Priorities: A Field-Wide Agenda


146. The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem


147. The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty


148. Do Methods Support the Claims? Intra-Paper Verification for Peer Review


149. The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science


150. Archetypes or ability? Clustering for modelling student mathematical competence


151. Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities


152. Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football


153. Large-Scale ChatBot Validation Through Customer Digital Twin Simulations


154. Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning


155. Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact


156. Predict before you train: Scaling Laws for particle physics foundation models


157. A Methodology for Designing Knowledge-Driven Missions for Robots