전체 AI 논문 - 2026-07-17

1. Pretraining Data Can Be Poisoned through Computational Propaganda


2. SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration


3. teLLMe Why (Ain’t Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data


4. AutoSynthesis: An agentic system for automated meta-analysis


5. When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space


6. Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation


7. Plover: Steering GUI Agents through Plan-Centric Interaction


8. Can We Trust Item Response Theory for AI Evaluation?


9. Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy


10. MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection


11. The Industrialization of Research ; On AI-Driven Science and Its Consequences


12. Concept-Guided Spatial Regularization for World Models in Atari Pong


13. Long-Context Fine-Tuning with Limited VRAM


14. BrainPilot: Automating Brain Discovery with Agentic Research


15. Man, Machine, and Masterpiece: Artistic Ownership in the AI Era


16. SMC-ES: Automated synthesis of formally verified control policies


17. Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development


18. Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers


19. CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models


20. Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation


21. Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach


22. Proof-or-Stop: Don’t Trust the Agent, Trust the Evidence – Loop Engineering for Verifiable Evidence-Gated Lifecycle Control


23. Reachability-Aware Pretraining for Efficient Target-Oriented Path Exploration in Temporal Knowledge Graph Reasoning


24. CrimeNER Demo: Named-Entity Recognition in the Crime Domain


25. Transcoders for Investigating Deception in Language Models


26. Global Index on Responsible AI: 2026 Report


27. AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery


28. InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring


29. Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment


30. Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications


31. SmartRAG: Native Graph-Based RAG for Mobile Device


32. TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning


33. MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers


34. Analytic Abduction: Causal Decomposition and Governed Commitment for Human–AI Coordination


35. Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models


36. SportD: Can VLMs Physically Strategize?


37. MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research


38. Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology


39. Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments


40. Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents


41. Democratizing Agent Deployment Safety: A Structural Monitoring Approach


42. Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation


43. Towards an Intention Abstraction Layer for Autonomous Industrial Systems


44. Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent


45. WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays


46. RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning


47. VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence


48. Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions


49. SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation


50. Step-Level Preference Learning for Generative Agents in Social Simulations


51. Tactile: Giving Computer-Using Agents Hands and Feet


52. Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers


53. CausalGraphX: A Counterfactual Graph Neural Network Framework for Explainable Systemic Risk Assessment


54. Reward-Free Evolving Agents via Pairwise Validator


55. Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration


56. CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models


57. A Comparative Analysis of Machine Learning Models for Long and Short-Term Forecasting of the Egyptian Stock Market: A Focus on EGX30


58. Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving


59. CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents


60. Traccia: An OpenTelemetry-Based Governance Platform for AI Systems


61. Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions


62. Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)


63. AI Agents Do Not Fail Alone:The Context Fails First


64. Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation


65. The Steering Budget: Examples beat Knobs


66. Align AI to Dynamic Human-AI Workflows


67. How Artificial Intelligence LLM Engines Shape the Global Conflict Information Environment


68. RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination


69. ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System


70. When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models


71. MemoHarness: Agent Harnesses That Learn from Experience


72. Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers


73. Enhancing Small Language Models Reasoning through Knowledge Graph Grounding


74. ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability


75. Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models


76. Human AI Construction of Bayesian Networks for Operational Decision Support – A Virtual Survey Approach


77. Interpretable Language Model for Closed-Loop Type 1 Diabetes Control


78. DialogueVPR: Towards Conversational Visual Place Recognition


79. RegNetAgents: A Multi-Agent Framework for Cross-Network Regulatory Driver Identification in Cancer Genomics


80. IMEX Interaction-Based Model Explanation


81. HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs


82. Intelligent Three Level Learning Architecture for Autonomous UAV Swarms in Search and Rescue


83. RoboTTT: Context Scaling for Robot Policies


84. SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions


85. SceneBind: Binding What and Where Across Vision, Audio and Language


86. Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents


87. In-Place Tokenizer Expansion for Pre-trained LLMs


88. Symbal: Detecting Systematic Misalignments in Model-Generated Captions


89. MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization


90. Mask-Aware Policy Gradients for Diffusion Language Models


91. Subjective Risk Decomposition: A New View for Uncertainty Quantification


92. T^2MLR: Transformer with Temporal Middle-Layer Recurrence


93. Scaling Behavior Foundation Model for Humanoid Robots


94. NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference


95. Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents


96. Towards Hierarchical Structure Understanding of Newspaper Images


97. ANet Patu-1: The Value of Connection in the Agent Network


98. Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening


99. When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration


100. LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research


101. OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios


102. Latent Trajectory Discrimination for AI-Generated Text Detection


103. Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation


104. Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control


105. A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems


106. Benchmarking Face Recognition without Real Faces


107. Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks


108. Show Me How You Reason and I’ll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution


109. FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers


110. StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows


111. Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs


112. Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature


113. Asymmetric Peak-Aware Loss for Peak-Critical Time Series Forecasting


114. RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems


115. Interventional Causal Circuits for Safe Robot Action Testing and Failure Recovery


116. Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience


117. Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning


118. Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality


119. Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition


120. Large Audio Language Models for Spoofing-Aware Speaker Verification


121. FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models


122. The Misclassification of Autistic Writing as AI-Generated


123. VideoSEMA: a scalable and efficient Mamba-like attention for video understanding


124. Harnessing LLMs for Reliable Academic Supervision: A Comparative Study


125. Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models


126. Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach


127. An Intelligent-Cloud Edge Multimodal Interaction System for Robots


128. LLM-Driven Approach to Modeling Tool Interoperability in Automotive Domain


129. MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents


130. Knowing You at First Glance: Inferring Apparent Personality from Faces


131. Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization


132. Angular Gaussian Supervised Contrastive Learning for Long-Tailed Electrocardiogram Arrhythmia Diagnosis


133. Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems


134. Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms


135. Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction


136. Governing Artificial Intelligence: Public Preferences and Regulatory Options


137. Gate-Zero Growth: A Geometric Framework for Function-Preserving Continual Learning


138. A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi


139. Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models


140. SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents


141. Controlled Reformulation Testing for Logical Consistency in Large Language Models


142. VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation


143. Multi-Scale ViT Inference with Habitat-Fit Priors and kNN Retrieval for Multi-Species Plant Identification


144. Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards


145. EdgeFaaS: A Function-based Framework for Edge Computing


146. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026


147. Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution


148. Global drivers and barriers to the public acceptance of autonomous vehicles: Evidence from 17 countries


149. Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel


150. ConFlow: Constraints-Guided Learning with Flow Matching for Motion Generation


151. Decision Making Needs Uncertainty Quantification [Lecture Notes]


152. Integration Matters: Rollout-Based Training for Constrained Diffusion Models


153. An offline approach to fNIRS-guided reinforcement learning for robot behavior


154. Why Git Is the Memory Solution for the Agentic Development Lifecycle


155. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI


156. HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization


157. Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values


158. Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution


159. The Prover Is the Judge: Verified Security Software from AI Coding Agents in Ada/SPARK


160. Beyond scalar losses: calibrating segmentation models via gradient vector field surgery


161. Copy-on-Write Scoring: Application-Specific Agent Evaluations


162. ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model


163. PReM: Learning What to Preserve and When to Refresh for Context Compression


164. Accounting for Hysteresis and Eddy Currents in Finite Element Simulations of Ferromagnetic Laminated Cores using a Recurrent Neural Network


165. Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models


166. Assessing AI in Introductory Physics Problem Solving


167. ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs


168. Local Additive Feature Attribution: A Mathematical Taxonomy and Reporting Checklist


169. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning


170. Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection


171. LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks


172. SeeSE3: Emergence of 3D Space in Vision Features


173. Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation


174. NexForge: Scaling Executable Agent Tasks via Requirement-First Synthesis


175. Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape


176. Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control


177. RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences


178. The Cost and Network Limits of Space-Based AI Compute


179. Structured Feedback Improves Repair in an LLM Agent Loop


180. Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation


181. Certified Domain Consistency for Multi-Domain Retrieval: Label-Free Per-Domain Contamination Control with Conformal Risk Guarantees


182. “Trust Junk” Leads to Unjustified Support for Highly Discriminatory Predictive Models


183. Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak


184. The Planar Case of Thomas Positive Circuits Conjecture


185. Volition Elicitation: Operational Semantics for People and Their Machines


186. Explainable Geospatial AI for Satellite Ground Station Siting Using LiDAR-Derived Terrain Intelligence


187. Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods


188. Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning


189. ReportMedSAM: Guiding Segmentation Through Radiology Reports


190. CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning


191. T5-CSBoost: Adversarial Perturbation Resistant LLM Fingerprinting


192. Information-Theoretic Limits of Reliability and Scaling in Language Models


193. Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect


194. MAPS: Modeling Co-Existing Subjective Perspectives and Shared Meaning in Multi-Agent Cognitive Dialogue


195. Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation


196. Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility


197. Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs


198. Token Time Continuous Diffusion for Language Modeling


199. Automatically Evolving Prompt Guidelines for Task-Specific Optimization


200. LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets


201. Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs


202. Falsifiable Release Gates for Self-Improving Systems


203. Fast-Fading Channel and Power Optimization of the Magnetic Inductive Cellular Network


204. All Polarized but Still Different: a Multi-factorial Metric to Discriminate between Polarization Behaviors on Social Media