전체 AI 논문 - 2026-08-14

1. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist


2. QuoteBench: How Matched Scores Can Hide Command-Path Failures


3. AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)


4. MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination


5. A Unifying Perspective on Causal World Models: From Observations to Representations to Structure


6. Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension


7. RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level


8. Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes


9. Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development


10. Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings


11. Jointly Predicting Courses and Grades Using a Transformer-Based Model


12. TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies


13. Rules or Character? Scaling Laws for AI Safety Design


14. LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning


15. LLM-Guided Graph Generation for Structure-Based Local Improvement Methods


16. StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems


17. NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space


18. Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision


19. Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability


20. vToken: Token-Level Virtualization for Reclaimable KV Caches


21. Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test


22. TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems


23. Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents


24. SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents


25. Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing


26. Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement


27. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback


28. Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting


29. Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment


30. SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference


31. EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding


32. Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds


33. Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)


34. Uniform Herding: Exemplar Replay with Representation Refresh


35. VALG: An Agentic System for ML Theory Research


36. DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition


37. BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs


38. From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion


39. Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic


40. OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways


41. Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$


42. Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses


43. FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving


44. Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence


45. Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence


46. Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals


47. ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification


48. AI and Consumer Rights in India Working Paper


49. Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents


50. Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories


51. CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation


52. ARAC: Benchmarking Auto-Research’s Alignment and Completeness on End-to-End Researchs


53. PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs


54. Correct Is Not Governed: Provenance Integrity in Agentic Workflows


55. Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence


56. Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies


57. The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis


58. Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs


59. Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy


60. On the Expressive Power of Transformers


61. Designing AI Pipelines for Decision-Ready ITSM Intelligence


62. General Probabilities of Causation with Causal Knowledge


63. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries


64. Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence


65. @skills: Attention is all you have


66. Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues


67. DiG-bench: Discovery in Games


68. Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting


69. Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces


70. Trie Automata for Constrained Decoding over Large Finite Sets


71. CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence


72. $\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution


73. Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents


74. MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents


75. Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction


76. Research Assistant: AstraZeneca’s Agentic System for R&D


77. Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization


78. Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation


79. Don’t Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese


80. Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning


81. Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing


82. Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments


83. Position: The Alignment Community is Unintentionally Building a Censor’s Toolkit


84. Diagnostic Foundation for Evaluating LLMs’ Research Integrity as Co-Scientists


85. Position: Reasoning is a Learnable Rule-Based Process


86. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design


87. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark


88. LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure


89. Vero: Can AI Agents Build Formally Verified Software Repositories?


90. The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity


91. DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data


92. Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity


93. Synthetic Persona Pretraining: Alignment from Token Zero


94. AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models


95. Concept Drift Detection and Adaptive Retraining of Malware Classification Models


96. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification


97. CAPRI: Contract-Aware Proof Repair for Isabelle


98. UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models


99. ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models


100. Algebraic Decomposition Theory for Transformer Length Generalization


101. Are You Sure You’re Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity


102. Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference


103. Deliberate Practice: Learning Robot Skills under a Budget


104. Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks


105. Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs


106. Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples


107. Training AI Scientists to Replicate Research


108. It’s How You Ask: Gender-Associated Linguistic Bias in LLMs


109. Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services


110. Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy


111. Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks


112. Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model


113. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures


114. Into the ORBIT for Time Series: Training Regimes for Foundation Models


115. Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models


116. Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data


117. GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport


118. Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales


119. CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport


120. NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video


121. GEM: A Generative Embedding Model Bridging Reasoning and Retrieval


122. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint


123. Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering


124. LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service


125. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation


126. EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory


127. Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization


128. LOB-ID: Evaluating Synthetic Market Data by Inception Distances


129. TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes


130. Operationalizing Cyber Threat Intelligence with GraphRAG


131. UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations


132. Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language


133. Generative Universal Multimodal Retrieval with Dual-role Identifiers


134. Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents


135. The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use


136. AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization


137. H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities


138. Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference


139. InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers


140. EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction


141. NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents


142. Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation


143. SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data


144. A Compositional Theory of Curvature in Probabilistic Circuits


145. BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving


146. Falsehood and Impossibility Are Different Directions in an AI’s Representation of Language


147. FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation


148. Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval


149. AQuA: Recursively Self-Improving Quantitative Trading Research Agents


150. From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options


151. Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing


152. Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors


153. PIPES: Securing Agent Perception with Provenance and Priors


154. CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives


155. Memorization Diagnostics for Code LLMs Should be Scale-Aware


156. Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents


157. SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization


158. PatientAct: Theory-Grounded Mental Health Client Simulation


159. ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval


160. Error-Aware Reverse Auction Mechanism for Large Language Model Routing


161. HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement


162. Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks


163. Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging


164. Demand Transfer Estimation at Scale via Restricted Logit Modeling


165. Interpretable Causal Discovery via Causal-Effect Constraints


166. Novels generated by language models show compressed formal variation


167. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory


168. LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning


169. PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping


170. Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks


171. What Makes a Peer? Valuation-Anchored Similarity in Private Markets


172. Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling



174. Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts


175. SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization


176. Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection


177. Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review


178. SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents


179. A Hierarchical Energy-Based Model for Multimodal Cognition


180. Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models


181. Query Timing Produces Opposite Positional Biases Between LLMs and Humans


182. From Observation to Intervention: Memory in Brains and Large Language Models


183. Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities


184. FluctlightDB: A Memory Model of Data for AI Agents


185. EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector


186. Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles


187. Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio


188. Humans are Missing from AI Coding Agent Research


189. Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF


190. Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students’ AI Literacy, Learning Outcomes, and Reflection


191. From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks


192. StreamReason-Bench: Can Large Language Models Reason about Event-Time Stream-Processing Semantics?


193. Mimicry without understanding: the origins of decision bias in large language models


194. StorySpark: Module-wise Evolutionary Search for Story Premise Generation


195. Steering the Language Axis: From Linear Decodability to Causal Control


196. Vision-Language Models are Fragile Multilingual Associators


197. Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching


198. AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement


199. Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition


200. When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models


201. Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance


202. What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting


203. LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning


204. The AI Accountability Ecosystem in the Era of Language Models