전체 AI 논문 - 2026-08-28

1. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution


2. Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation


3. Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study


4. CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases


5. Sophistication in GenAI Use: Field Evidence from a Large Firm


6. Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance


7. Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification


8. LLMs Can Design Near-Optimal OR Algorithms


9. BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models



11. What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents


12. Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable


13. BPMN4CAI: A BPMN Extension for Modeling Dynamic Conversational AI


14. Thomson: Continual Learning of Frontier Models for SovereignAI


15. When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents


16. Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection


17. GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL


18. TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation


19. LAAF: A Layered Accountability Architecture Framework for LLM Applications


20. pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning


21. A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes


22. Omni-Interactive Universal Embedder


23. A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems


24. ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions


25. DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research


26. GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory


27. From Atomic to Agentic: Towards Interpretable Evaluation of LLMs’ Agentic Mathematical Capabilities


28. A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering


29. Counterfactual Bias Testing for Application Tracking System


30. AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion


31. Learning-Augmented Online Allocation under Unreliable Advice: Robustness, Exposure Fairness, and Distribution Shift


32. Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall–workload trade-offs and run-to-run consistency


33. C-Unseen: Weak Signal Detection in Dynamic Temporal Knowledge Graphs via LLM Reasoning


34. BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click


35. LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems


36. SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers


37. Decoupling Planning and Control for Instructable Agents


38. AI Control Scientist: LLM-driven Agentic System for Automated Control Design


39. Categorizer Automata for Discounted-Sum Payoffs


40. DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?


41. Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case


42. AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design


43. Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds


44. Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training


45. Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing


46. Accelerating Scientific Research with Gemini in the Real-World


47. Five Primitives for Governing Autonomous AI Agents at Runtime


48. Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation


49. SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation


50. AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling


51. DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows


52. Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation


53. PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents


54. Don’t Overthink, Don’t Underthink: Toward Adaptive Reasoning in Agentic AI


55. Fine-Tuning of Transformer models with Frames


56. ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving


57. FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence


58. Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems


59. Assessing mentalization in humans and large language models


60. SKILL.state: Scalable Long-Horizon Agent Skills


61. 6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation


62. The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts


63. LLM Agents for Time-Series: A Survey


64. Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy


65. Same Model, Different Harness: Different Coding-Agent Results


66. GameWAM: A World Action Model for Video Games


67. Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling


68. Agentic AI for operating scientific instruments for nanoscale characterization



70. Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs


71. Predicting Consequences and Reinforcing Navigation Policies with Latent World Models


72. Invocation-Level Reliability of Tool-Using Agents


73. Is Your Neighborhood Safe? Place-based Stigma in Large Language Models’ Urban Safety Judgments


74. Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion


75. TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education


76. Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI


77. AI Revealed Preferences


78. Knowledge Cards: Structured Knowledge for AI Systems


79. Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript


80. A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving


81. A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian


82. SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum


83. GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions


84. Selection Bias Correction in Retail Intelligence


85. EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG


86. Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration


87. Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models


88. Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse


89. LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs


90. The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting


91. The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning


92. CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering


93. PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices


94. Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap


95. Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset


96. EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction


97. SWE-Prime: Fewer Trajectories, Better Performance


98. From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench


99. RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution


100. Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit


101. Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners


102. CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators


103. How Language Models Organize and Structure Moral Knowledge


104. Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction


105. LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics


106. Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions


107. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models


108. Stageboost: Recommending Signals Based on Counterfactual Estimation


109. KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations


110. RCMN: Understanding Misleadingness in Influential Public Discourse


111. PAWBench: How Far Are We from Probabilistically Aligned World Modeling?


112. Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit


113. TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection


114. Compositional Online Learning for Semantic Data Processing Systems


115. STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration


116. PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference


117. LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration


118. When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue


119. ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification


120. Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents


121. Active sensing to characterize the heterogeneity of plant stress


122. Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors


123. Learning Transverse Momentum Distributions from Raw Scattering Events via Conditional Diffusion


124. Emotional Preferences as Goal-Priority Regulation


125. Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations


126. Performance Foundations of Parallel & Distributed Reasoning Language Models


127. Multi-Person Human Motion Forecasting in Complex Scenes


128. FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets


129. Magnon-induced phononic Chern insulator


130. Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS


131. When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems


132. PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?


133. Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic


134. From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation


135. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA


136. Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay


137. Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory


138. Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research


139. FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs


140. Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction


141. Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning


142. LiveVVT: High-Fidelity Video Virtual Try-On in Real Time


143. AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability


144. FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models


145. Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference


146. PailitaoGR: Latent Think-with-Images for Generative Image Retrieval


147. CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes


148. Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries


149. J-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data


150. Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations


151. RTNav: Towards Real-Time Zero-Shot Object Navigation


152. Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance


153. Diff Mining: Logit Differences Reveal Finetuning Objectives


154. SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning


155. The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection


156. Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI


157. Simultaneous Envy and Equitability Guarantees


158. Co-Evolving Structured Knowledge and Reasoning in Language Models


159. Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries


160. CG4AI: A Column Generation Framework for Training AI Models Under Constraints


161. Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives


162. Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds


163. How Unlikely Is “Unlikely”? Assessing Verbal Probability Perception Across Large Language Models


164. Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models


165. MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models


166. On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study


167. How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation


168. NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation


169. Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model


170. Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs


171. ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices


172. A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers


173. Comparing Chunking and Embedding Strategies for Turkish RAG Systems


174. When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models


175. Investigating the Influence of Prompt and Response Languages on LLM Content Generation


176. PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants


177. A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs


178. Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors


179. ClassVision: AI-Powered Classroom Attendance System


180. DRL: A Deterministic Relational Middleware Layer for Transaction-Safe Enterprise NL2SQL Under Schema-Graph Scaling


181. Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention


182. Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning


183. Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment


184. From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents


185. Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models


186. Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation


187. VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation


188. Evaluating AI Generated Summaries for Cancer Patients



190. Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes


191. Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation


192. Syntax vs. Semantics: How Transformers Learn Deep Dependencies


193. FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes


194. Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales


195. From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement


196. Exploring the Role of LLMs in HPC Programming: A Survey