전체 AI 논문 - 2026-09-03

1. Discriminative World Models for Web Agents


2. AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application


3. Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis


4. SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment


5. Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents


6. Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems


7. Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills


8. Door-in-the-Face Requests and Refusal Behaviour in Large Language Models


9. Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting


10. Collective creativity in hybrid societies


11. CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI


12. UTP-Bench: Uncertainty-aware Travel Planning Benchmark


13. Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks


14. Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions


15. SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning


16. Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds


17. SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology


18. CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging


19. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems


20. APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering


21. LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails


22. Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training


23. Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality


24. PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks


25. PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment


26. SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams


27. PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion


28. ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models


29. Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning


30. FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs


31. EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision


32. Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents


33. Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics


34. READY or Not: Reliable Enterprise Agent Deployment


35. MASkills: Continual Skills Optimization for Multi-Agent LLM Systems


36. Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems


37. CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning


38. ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction


39. MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity


40. DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents


41. Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision


42. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models


43. ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations


44. When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor


45. Benchmarking Language Models for Statistical Problem Formulation


46. Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment


47. Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?


48. The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction


49. Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence


50. Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization


51. The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents


52. SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval


53. Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern


54. Induction and Inquiry via Probabilistic Reasoning over Language and Code


55. When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection



57. Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI


58. EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models


59. Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework


60. Post-Training Language Models for Gold-Medal Performance in Coding Competitions


61. frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study


62. Dutch Books for Language Models


63. From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution


64. Untangling the Mechanisms of Misleading Context in Medical Question Answering


65. HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design


66. Language Models Can Control Their Own Attention


67. RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models


68. DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models


69. From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs


70. TaRA: Training-Aware Low-Rank Adaptation Initialization


71. Automated Vulnerability Injection in Smart Contracts Using Large Language Models


72. Competitive Market Behavior of LLMs


73. ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction


74. Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs


75. Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework


76. Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion


77. Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models


78. RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection


79. ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering


80. DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models


81. Addressing Trust in AI Systems through Education: A Didactic Perspective


82. Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression


83. Towards One-for-All Robustness Across a Continuum of Threat Levels


84. Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment


85. Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking


86. Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts


87. PolERo: Studying Political Evasion in Romanian


88. MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts


89. Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance


90. NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning


91. Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions


92. Fair Stable Matching: A Nash Social Welfare Approach


93. Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information


94. AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers


95. ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans


96. What Is Worth Representing? Representational Empowerment for Continual Model Construction


97. DiffIE: Diffusion-based Open Information Extraction


98. SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment


99. VoRTeC: Taming Foundation Flow for One-step Real time Video Compression


100. RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification


101. Auditory Illusion Benchmark for Large Audio Language Models


102. Do Large Language Models Capture the Diversity in their Training Data?


103. PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation


104. CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation


105. Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics


106. DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space


107. SAUF-Net: Structure–Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation


108. InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models


109. Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging


110. SMart: A Multi-source Multi-phase Time Series Representation Transfer Framework


111. Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation


112. GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories



114. Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation


115. OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations


116. Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor



118. C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees


119. text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation


120. Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap


121. MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs


122. Predict, Don’t Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models


123. Git4Data: Database-Native Version Control for AI Agents


124. Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts


125. Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models


126. Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation


127. Seed-Anchored Budget-Bounded Graph Rendering for Question Answering on Industry-Standard Power-Grid Information and Exchange Models


128. InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation


129. InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation


130. Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight


131. Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network


132. On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers


133. Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens


134. Accurate in space, unreliable in time: how LLMs represent national cultural change


135. OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation


136. Thinking effort aligns between humans and reasoning models in abductive reasoning


137. Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge


138. Agent Memory Is a Surface for Endogenous Authorization Laundering


139. Interpretable Symptom Vectors for Depression in a Large Language Model


140. Zeta-Lite: A Concurrent, Branchable In-Browser SQL Database for Agentic Memory


141. hLLM: Single Pass Decoding for Generative Reranking


142. VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages


143. Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks


144. Dictionary-Guided Mutation Operators for Automated HDL Repair


145. Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics


146. Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives


147. HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation


148. RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis


149. Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study


150. CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction


151. Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports


152. Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration


153. How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making


154. PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation


155. NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference


156. From Feature Interaction to Feature Transport - A Unified Block for Scalable Recommendation Models


157. A Data-Driven Multimodal Method for Early Detection of Coordinated Abnormal Behaviors in Live-Streaming Platforms


158. The Utility of LLMs in Recommender Systems Explanation Evaluation


159. RecEvolve: A Knowledge-Driven Autonomous Agent System for Recommender Systems



161. WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling


162. Two Centuries of Sexism in British Parliament: A Computational Analysis of Women’s Representation in the Hansard Corpus