전체 AI 논문 - 2026-09-08

1. A Deep Generative Model for Synthesizing Labeled Wireless Signals


2. Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe


3. Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence


4. Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models


5. CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents


6. Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education


7. Does Your Agent’s Memory Survive a Model Upgrade? A Controlled Study of Memory Portability


8. Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models


9. LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams


10. Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness


11. RISE: Recursive Improvement via Self-Extrapolating Policy Distillation


12. Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods


13. GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity


14. Testing Interchangeability in LLM Agent Teams


15. Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference


16. AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance


17. Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents


18. Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions


19. A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability


20. Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory


21. Uncensored Open-weight Models: Redistribution as the Persistence Layer


22. Substrate-Aware AI Agents: Execution Context as a First-Class Input


23. ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs


24. CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review


25. What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection


26. The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior


27. A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment


28. SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding


29. Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens


30. Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation


31. ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding


32. LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28


33. Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent


34. Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?


35. TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents


36. MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning


37. Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications


38. Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment


39. TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing


40. Language models judge war differently when tested for alignment


41. A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering


42. Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball


43. Why We Care About Understanding: Competence through Predictive Compression


44. Solving Hard XAI Queries Based on a Compiled Dual-Rail Encoding


45. Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing


46. Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing


47. From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments


48. Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach


49. MARLA: A Conceptual Scaffold for Regulatory Learning under the EU AI Act


50. AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems


51. CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games


52. From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents


53. LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models


54. CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution


55. MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting


56. MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models


57. ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults


58. Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance


59. CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric


60. When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models


61. MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis


62. Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection


63. Whose record is this? Diagnosing and authorizing record use in personalized multimodal models


64. ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing


65. DODR: Deterministic Operator-Driven Reasoning in Latent Space


66. Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges


67. Shadow Queries for Private Retrieval in Vector Databases


68. DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems


69. Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM


70. PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces


71. FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality


72. Model Retirement Creates Reproducibility Risk in Biomedical AI Publications


73. SQL-Zero: Self-Evolving Text-to-SQL


74. Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network


75. Train What You Deploy:Token-Faithful Post-Training of a Production Coding


76. ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies


77. Harness-agnostic detection and immunization of reward hacking in self-evolving language models


78. Continual Graph Memory for Adaptive Recommendation under Intent Drift


79. A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark


80. SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents


81. Leveraging Imperfect Restoration for Data Availability Attack


82. $τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction


83. Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines


84. Extremely Sparse Supervision Incentivizes Reasoning Ability


85. La Agente Óptima: Towards Agentic Self-Driving Laboratories


86. Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection


87. IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion


88. From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs


89. Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials


90. Towards a universal language of concepts: A survey


91. MaxKernel: Agentic Kernel Generation for TPUs


92. What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents


93. BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker


94. Rethinking Indirect Prompt Injection as a Test-Time Search Problem


95. ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality


96. When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference


97. PerfReasoning: How Well Do LLMs Reason on Hardware Performance?


98. HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals


99. Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer


100. Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets


101. A Removal Based Approach to Improve LLM Faithfulness at Test-Time


102. Iris: Climbing to the Search Frontier


103. Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security


104. Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation


105. From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance


106. EXAONE Forecast for Finance


107. Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction


108. RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments


109. Reflection-aware Generative Novel View Synthesis


110. What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies


111. When LLM Decompilers Recompile More and Preserve Less


112. Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool


113. The History Is the Detector: Executing CVE Patch History, End-to-End


114. Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions


115. RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?


116. How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing


117. CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls


118. Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization


119. PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting


120. A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR


121. Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets


122. AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics


123. Beyond Stationarity in Time Series: Discovering Causal Structures and Latent Regimes via Markov Blankets


124. A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment


125. A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning


126. TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors


127. NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer


128. A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support


129. Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection


130. Beyond Co-purchase Relation: Evolution of Complementary Recommendations at Allegro


131. Qlippy: A Retrieval-Augmented GenAI Assistant for Reproducible Quantum Workflows and Experiment Tracking


132. How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions


133. Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection


134. How a Chatbot’s Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI


135. Amortizing Scaling Law Construction Costs


136. VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition


137. MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression


138. One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation


139. ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems


140. TreeFI: Value-Aware Statistical Fault Injection for Deep Neural Networks


141. Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair


142. Methane Detection On Board Satellites from Unorthorectified Imagery


143. Sound-based Multi-Person 3D Pose Estimation


144. Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction


145. RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents


146. Attention-guided super-resolution of 4D flow MRI in carotid arteries


147. SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection


148. ReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk Classification


149. Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents


150. PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation


151. Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning


152. CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation


153. MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain


154. MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate


155. Reinforcement Learning for improving Large Language Models’ Catalan text simplification capabilities


156. Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution


157. Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3


158. Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents


159. Can Activation Steering Capture Multidimensional Authorship Style?


160. Dynamic Heterogeneous Graph Representation Learning: A Survey


161. Persistent Teacher Anchoring for Tool-Using Agents


162. When Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI


163. Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models


164. Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models


165. Building a research-software catalog with a coding agent: from hackathon prototype to public deployment


166. Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty


167. Wireless Foundation Models: State-of-the-Art and Open Challenges


168. Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion


169. Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle


170. Tracing Audio Grounding and Answer Selection in Audio LLMs


171. SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds


172. PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning


173. Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms


174. When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models


175. Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models


176. Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics


177. Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI


178. Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection


179. Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes


180. Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters


181. A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap


182. Hakken: Predicting future discoveries to fill the gaps in today’s knowledge


183. Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective


184. Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning


185. Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation


186. Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning


187. A Roadmap for MEG Foundation Models


188. When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models


189. GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion


190. REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation


191. A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models


192. You Really Didn’t Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments


193. What Moves? Localized Motion Representations for Compositional Scene Control


194. Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models


195. Ultrasound-Based Prediction of Cirrhosis Decompensation Using Large-Scale Computer Vision Models


196. Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs


197. Blockchain-Enabled Secure Logging for Fiscal Electronic Mechanisms: Evaluation of the Greek eSEND and myDATA Tax Systems


198. VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models


199. Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors


200. Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis


201. Abstraction Agent


202. Evidence Integration in Large Language Models


203. Scalable Context Orchestration for Serving LLMs Over Voice


204. When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs


205. AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks