전체 AI 논문 - 2026-08-27

1. Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings


2. SwarmWorld: Stigmergic technological evolution in societies of language-model agents


3. Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems


4. Imitation Learning for Connection-Tableau Construction


5. AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs


6. ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs


7. Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs


8. SciMIF: Understanding Multimodal Instruction Following in Scientific Domains


9. Quantitative Analysis of $ω$-Regular Robust MDPs


10. LivingRAG: Augmenting Graph RAG with Experience


11. Candidate supply and answer selection shape the value of LLM judging in multi-agent systems


12. How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation


13. Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures


14. Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems


15. Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions


16. LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents


17. ToST: A Tree-of-Thought Socratic Teaching Framework for Multi-Path Guidance and Parallel Thinking


18. Narcissus: Program Synthesis Using Context-Aware LLM Approximations


19. Using profiles of cognitive capability to assess AI suitability for workplace tasks


20. Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models


21. CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval


22. PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning


23. Training Alignment Auditors via Reinforcement Learning


24. Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness


25. Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents


26. Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks


27. Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs


28. Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory


29. FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review


30. BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks


31. PhaseShift: Topology-Aware Data Harmonization and Model Consolidation Across Signalized Intersections


32. Hierarchical MoE for Multi-Modal ILD Diagnosis


33. FLARE: Verifying MILP Reformulations with LLM-Based Theorem Proving


34. LLM-Driven, Datasheet-Aware Automated Hardware Compatibility Verification for Early-Stage, Pre-Schematic Embedded System Design


35. Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows


36. Tunable Tool-Call Rates in LLM Agents via Representation Steering


37. FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs


38. Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning


39. PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?


40. Auto-Policy, not Auto-Skill: Compiled Agent Skills for the Physical World


41. SimVerity: When Does Simulated Agent Success Survive Physical Deployment?


42. LifePlanner: Evaluating LLM Agents for Geo-spatial Planning with Social Media Data


43. Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs



45. Solving Robust POMDPs with Omega-regular Objectives via Partially Observable Stochastic Games


46. FrontierChallenge: Evaluating Scientific Workflow Completion


47. post-graph-rag: A PostgreSQL-Native Graph RAG Engine


48. Semantic Graph Unification for Industrial Digital Threads: Bridging 11 Heterogeneous Manufacturing Systems Through Ontology-Driven Knowledge Graphs


49. Same-Player Verification for Account Consistency in Counter-Strike 2


50. Natural Language Input, Semantic Track Representation, and LLM Inference: Making the Maritime Information Exchange Model Tractable


51. Measurement-Budget Allocation in Quantum Learning with Finite-Shot Generalization Guarantees


52. AI-Powered Mental Health Chatbots in Africa: A Systematic Review and Culturally Adaptive Framework


53. Reliable LLM-Powered Decision Engines for Large-Scale Supply Chain Operations: Architecture, Safety, and Performance Guarantees


54. SIMGUIDE: Procedurally Grounded Multi-Context Representations for Personalized Agent Planning


55. VLM-based automatic multi-granularity graph representation of building layouts for design informatics


56. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning


57. A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training


58. MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching


59. Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders


60. TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development


61. ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing


62. Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving


63. Prefix Sliding for efficient test-time scaling


64. $R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning


65. How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention


66. The Value of Human Expertise


67. DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation


68. VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction


69. FRAME: separating sampling variation from representational cause in medical imaging fairness


70. PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology


71. A Statistical Audit of Physical AI Benchmark Redundancy


72. One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation


73. TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding


74. When Composition Doesn’t Add Up: Humans Identifying Defects in AI-Generated Images


75. Code World Model: Coding Agent as World Brain


76. Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation


77. Towards A Unified Information Bottleneck Framework for Time Series Explanations


78. Unlocking Multimodal Protein Language Models at Inference Time


79. Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening


80. VINCENT: Validated Interaction Network for Cross-drug Explanation of Therapeutics


81. Missing the Butterfly and Predicting the Past: Features or Bugs of Accurate AI Weather Models?


82. Skill Issue: Are Skills Language-Invariant in LLMs?


83. Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training


84. EVOMAL: Self-Poisoning in Self-Evolving Coding Agents


85. MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum


86. Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models


87. TailSFT: Filtered Fine-Tuning Improves Post-Training Performance


88. MeMark: Membrane-Space Watermarking for Spiking Neural Networks


89. Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale


90. It’s a matter of timescale: non-linear utility in successor features and multi-objective planning and learning


91. When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies


92. Difficulty-Aware Sample Allocation for Adaptive Data Augmentation in Semantic Segmentation


93. Unsupervised Anatomical Feature Learning via Diffusion Models: Enhanced Medical Image Segmentation with Denoising Diffusion Probabilistic Models


94. Learning New Facts with QLoRA: An Acquisition-Retention Frontier


95. AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation


96. Data Citation for Large Language Models: A Challenge


97. From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis


98. Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty


99. Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context


100. Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking


101. LDAC-Net: A Learnable Multi-Lag Differencing Attention-Convolution Network for Drift-Robust Recognition with Low-Cost MOX Gas Sensors


102. Leveraging Inter-object Affordances for Efficient Planning in Contact-rich Tasks


103. Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding


104. A Dual-Transformer for Multi-Camera View Recommendation


105. A Hierarchical Synergistic Deep Learning Framework Integrating Composition, Structure, and Ionic Transport for Solid-State Electrolyte Discovery


106. Are Concept Bottleneck Models Effective as Decision-Support Systems?


107. V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning


108. PolyMemDB: A Polyglot Database System for AI Memory Management


109. MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations


110. ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models


111. Controllable Affective Generation via Latent Vector Steering


112. CrossMambaTuning: Synergistic Spatial and Cross-Layer Adaptation for Machine Vision Compression


113. Physics-Informed Foresight Pruning for Sparse PINN Solvers of Nonlinear PDEs


114. AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research


115. When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory


116. A Hybrid Usability Approach for Rating Evaluation of M-Commerce Applications


117. A Tendon-Driven Five-Fingered Hand with Distributed Tactile Perception for Dexterous Manipulation


118. Goodput Maximization for Large Language Model Edge Inference: A Two-Phase Maskable PPO Approach


119. Syn2Logic: End-to-End Neuromorphic Design Automation


120. Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming


121. MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities


122. DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation


123. 4DStreamCtrl: Interactive Video Generation with Online 4D Control


124. Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction


125. Automatic weld seam segmentation for industrial quality control: a comparison of RGB and polarimetric imaging with CNN and transformer architectures


126. MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration


127. VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality


128. MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize


129. Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery


130. DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models


131. RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps


132. Token-Oriented Semantic Communication with Pretrained Vision Transformers


133. PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction


134. Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks


135. Q&A or Document-Based? The Effects of Interface Type on How Screen Reader Users Access Interconnected Documents


136. Adaptive Triggering for Bias Correction in LLM Reasoning


137. CRAMER: Control via Request-Aware Masking for Editing Recommenders


138. Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement


139. CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos


140. Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks


141. LLMscope: Extracting LLM Assets from Edge AI Chips via Optical Probing


142. InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance


143. Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs


144. A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography


145. Generative Action-Chunk Sampling for Adaptive Stiffness Control in Physical Human-Robot Collaboration


146. SHSP: Structure-Aware Hierarchical Solution Prediction for Mixed-Integer Linear Programming


147. Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation


148. Neural-Bayesian Structure Learning for Discrete Choice Modeling


149. What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift


150. A Few Pages of Markdown: Committed AI Configuration and Lower Quality Cost after Coding-Agent Adoption


151. Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making


152. Output Dilution: Redundant but Fragile Representations in MoE Models


153. SPECMINE: A Large-Scale Corpus of Spec-Driven Development Artifacts


154. Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment


155. Hyperbolic Latent Geometry for Tree-Structured Prototype Networks: A Local-vs-Global Trade-off


156. Self-Explanation Tutor for Active Study of CS1 Worked Examples


157. Lightweight Machine Learning-Driven Monocular Sidewalk Path Extraction for Embedded Micromobility Navigation


158. AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models


159. Bayesian Flow Networks for Offline Trajectory Planning


160. Belief Cascades Drive Persuasion in LLM Agent Networks


161. Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?


162. SkyDrive: Learning to Drive in a New City from Aerial Traffic Monitoring


163. Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching


164. When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting


165. SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation


166. Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events


167. GRAPE: Gradient Refinement and Progress-Aware Exploitation for Query-Efficient High-Dimensional Bayesian Optimization


168. Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment


169. Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures


170. The Von-Neumann State-Space Transformer for neural decoding


171. NVExplain: Explaining Time Series Forecasting with Latent Trajectory Analysis and Structure-Preserving Surrogates


172. DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning


173. HealthBench-Psych: A Mental Health Subset of OpenAI’s HealthBench


174. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration


175. DataKernelBench: Can LLMs Optimize Database Queries on GPUs?


176. SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification


177. Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels


178. ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies


179. A Primer on Computational Semantics for Artificial Intelligence Systems


180. Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal


181. D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation


182. Unsupervised Post-Training of Foundation Models: A Survey


183. Clearing the Underbrush: AI-Enhanced RF Interference Suppression


184. Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation


185. Targeting the Attention Heads Behind Object Hallucination in LLaVA


186. Evaluating and Preventing Security Smells in AI-Generated Ansible Code


187. Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace


188. The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline


189. Demystifying Reinforcement Learning Post-Training of Language Models


190. CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery


191. FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference


192. When Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study


193. ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration


194. Multi-Modal Anomaly Detection: A Survey



196. A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards


197. Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation


198. Modality Contribution Score - A Per-Patient Framework for Quantifying the Relative Diagnostic Contribution of Structural MRI and Amyloid PET in Alzheimer’s Disease


199. Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment


200. Domain-Adaptive ASR for Telephony AI Agents: Fine-tuning Canary Flash Models for Enterprise Contact Center Applications


201. Analyzing and Correcting Benevolence Bias in Large Language Models


202. From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction


203. PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting


204. Beyond the Chatbot: Co-Learning and Co-Teaching through a Dual-Persona Generative-AI Assistant


205. aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI


206. MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents


207. Real-time closed-loop protocol to assess neural variability in temporal coding


208. What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals


209. VisDocAgentBench: Benchmarking Agents for Visually Rich Document Retrieval