전체 AI 논문 - 2026-08-17

1. Handover of In-Context Learning State Across Session Boundaries


2. Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers


3. Split the Labor: Separating Evidence Interpretation from Decision Aggregation


4. Twin: Playing an Unknown Game with a Test-Time Digital Twin


5. Shift Aware Transfer Learning with Adaptive Dual-Encoder Fusion for PM Forecasting in Data-Limited Environments


6. SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning


7. Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports


8. PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments


9. The Dynamics of Intelligence Explosions


10. Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations


11. The Past and Future of AI Scientists


12. LLMs Don’t Pay for the Jump


13. Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons


14. AgentRewind: Recoverable Execution for Long-Horizon LLM Agents


15. Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages



17. Disentangled Shared Representations Improve Morpho-Transcriptomic Integration


18. ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond


19. Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents


20. Program-space Diffusion for Morphology-to-Transcriptomics Prediction


21. AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs


22. Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture with Hardware-Enforced Safety


23. Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning


24. TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments


25. Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models


26. Polaris : Multi Agentic System for Conversational Enterprise Analytics


27. Attributing Preprocessing Invariance in Spectral Foundation Models


28. MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement


29. A Generalized Parallelogram Rule for Proportional Analogies on Riemannian Manifolds


30. APTER: Adaptive Post-Training with Expert-Grounded Rubrics


31. FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction


32. Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding


33. BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs


34. Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine


35. Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach


36. QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation


37. Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost


38. Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground


39. A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents


40. Retrieval Grounding Latent Reasoning for Dense Retrieval


41. Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers


42. A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images


43. Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails


44. Scaling Domain Data Repetition in LLM Pretraining


45. Benchmarking data-driven material models on the classic Treloar dataset


46. Demystifying Agent Skills: Why They Work-Until They Don’t


47. Agent-Orchestration in Autonomous Chip Design


48. Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders


49. Buy the Rumor, Sell the News: When Is News Priced In?


50. Simulation-Driven Vehicular Traffic Data Augmentation: Extending Sensor Coverage Through Virtual Sensing


51. Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy


52. Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead


53. Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence


54. HELIX: Model-Harness Co-evolution for Recursive Self-Improvement


55. AI Research Preference Models


56. Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact


57. When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict


58. MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends


59. Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference


60. SDO: Subspace Deconflicting Operator for Multi-Adapter Composition


61. From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL


62. FLARE MCMC: Fidelity-based Layer-Adaptive REcursive proposals for MCMC


63. Simulation-Aware In-Context Policy Improvement for LLM-Aided Analog Layout Refinement


64. Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation


65. Coverage Aware Active Evaluation for Failure Discovery with Paired Systems


66. Learning to Assemble Novel Structures with Unfamiliar Parts under Semantic Constraints


67. Second Thought: Reasoning in Parallel as LLM Agents Act and Observe


68. Ontology-Grounded Project Memory for Coding Agents


69. Exploring ESC Winners with Nested Diagrams


70. A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure


71. Reward Machines for Signal Temporal Logic


72. ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction


73. Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning


74. Algorithm Design and Physician Liability


75. How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights


76. SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data


77. Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis


78. No Universal Signal Predicts Sample-Level LLM Regression under Version Updates


79. MobileMem: Learning from a Year of Mobile Experiences


80. Active Perception for Embodied Disambiguation


81. Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents


82. Measuring Cross-Task Behavioral Consistency in Language Model Agents


83. Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors


84. AI Evaluation Should Work With Humans


85. Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents


86. A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing


87. Modular Cognitive Architecture Emerges in Large Language Models


88. Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking


89. Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation


90. Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils


91. Marionette: Predicting World States, Rendering Geometry, Painting Appearance


92. Learning-to-Transition for Large-scale and High-Order MIMO Detection


93. RecipeNet: A Hierarchical Transformer for Recipe Data


94. Universal Thermodynamic Interatomic Potentials for Crystalline Materials


95. Generating Benchmark Health Data Using a Tabular Diffusion Transformer


96. Optimal Scheduling of Road Maintenance Jobs Considering Impact on Traffic Flows


97. Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes



99. Designing Compact Neural Architectures via Neuron Gating and Mixed Activation


100. From Style Replication to Style Exploration: Enabling Art Style Exploration with Analyze-Experiment-Resituate Framework


101. Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice


102. AI-Assisted Discovery and Construction of a Counterexample to the Convergence of Three-Block ADMM with the Identity Matrix as its Third Constraint Block


103. Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination


104. GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection


105. DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding


106. Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation


107. A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models


108. Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes


109. A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation


110. Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans


111. Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels


112. Seeing Red, Thinking Bad: Color Bias in Vision Language Models


113. SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning


114. Multi-Objective Bayesian Optimization for Model Merging


115. Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response


116. Training Fair Tabular Foundation Models



118. Structure-Guided Spatiotemporal Attention Graph Neural Network for Traffic Flow Prediction


119. Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation


120. Self-Supervised Visual On-Policy Distillation


121. SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation


122. HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting


123. Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions


124. AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations


125. BGA: A noise-immune neural distillation framework for malicious signature extraction in high-entropy encrypted flows


126. Overcoming Shortcut Learning in Graph Neural Networks through Active Explanation Guidance


127. From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics


128. Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation


129. Forecast Collapse in Time-Series Foundation Models


130. Rewrite Once, Validate Anywhere: Producing OWL-Aware SHACL Constraints (Extended Version)


131. P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems


132. Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation


133. MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation


134. Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency


135. Voxel-based 3D Facies Segmentation from Seismic Data: A Comparative Study


136. Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use


137. HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation


138. AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning


139. ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models


140. Content Based Video Narration of Gameplay with Vision Language Models


141. MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning


142. EchoRec: Multi-Item Prediction-Empowered Generative Recommendation via Cycle-Consistent Preference Alignment


143. Musical Mirrors: The LLM as Sounding Board in Songwriting


144. CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification


145. CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing


146. Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning


147. CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts


148. Agentic Transaction: Towards ACID-Compliant Agent Systems


149. Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen


150. Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model


151. Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions


152. ASSERT: A Measurement Pipeline for GenAI Audits


153. AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution


154. Optimal Power Allocation and AI Receiver Design for Superimposed DMRS and Data Transmission


155. PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization


156. Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions


157. CutClean: Neural Network Pruning for Privacy-Preserving Inference


158. Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models


159. Data-driven techniques for translational neuroscience and personalized neuro-health


160. Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline


161. Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost


162. Capacity-Dependent Effects of Data Selection for Reasoning



164. TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials


165. CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA


166. SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers


167. MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation


168. Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT


169. From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models


170. Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation


171. Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography with Diffusion Models


172. Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning


173. Context Aware AI Assistant and AR Interface for Lunar Extravehicular Activity (EVA) Procedural Guidance


174. IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering


175. UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction


176. From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent


177. Jais 2: A Family of Arabic-Centric Open Large Language Models


178. BCMT: Blockwise Causal Memory Transformer


179. Interactive Analysis of Global Explanations using Aggregated Class Activation Maps for Network Data


180. The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microsoft Excel


181. Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems


182. Think in Latent, Explain in Language: Self-Explainable Latent Reasoning


183. Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study


184. Don’t Claim Benchmark-Oriented Optimization Improves General Coding Capability – Diverse Evaluation Is Required


185. Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support