전체 AI 논문 - 2026-09-22

1. Harness-Zero: Harness Distillation via Agent-as-Harness


2. Emergent Collusion in Long-Horizon LLM Agent Interaction


3. Et Tu, Brute? Economic Misalignment in Personal AI Agents


4. BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction


5. A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories


6. Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models


7. Partner-Specific Affective Precision in Social Active Inference


8. Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection


9. MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution


10. GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes


11. Convex AI Compositionality and the Governance of AI System Populations


12. Construting Reverse Thinking: Developing Large Language Models’ Reverse Thingking Ability


13. Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents


14. World State Generator


15. TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction


16. Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents


17. DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security


18. Custom Named Entity Recognition and Topic Classification for Global Health Publications


19. Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis


20. The Endless Exam: Mathematical Constructions from Today’s Models toward Superintelligence


21. Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging


22. Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards


23. Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information


24. VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning


25. Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs


26. LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning


27. Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling


28. When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain


29. How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation


30. Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior


31. Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement


32. Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting


33. SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision–Language Models


34. LIMIT: Less Is More for Instruction Tuning in Text-to-SQL


35. CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment


36. APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction


37. Self-Healing Harness for Runtime Oversight of Agent Self-Modification


38. EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation



40. DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents


41. Incremental Consistency Execution for Autonomous Intelligent Systems


42. Representation-guided in-context learning for medical image interpretation with multimodal large language models


43. Structured Decomposition for Reliable LLM-Generated Access Control Policies



45. Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech


46. Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure


47. FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering


48. ACLArena: Agent Continue Learning in Multi-stage Post-training


49. Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents


50. LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning


51. UniK: Universal Knowledge Perception for Digital and Physical AI


52. Divergent strategies and convergent outcomes in autonomous materials discovery


53. Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control


54. Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems


55. Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer


56. Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery


57. Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models


58. WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks


59. Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows


60. On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks


61. ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents


62. PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI


63. Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment


64. AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction


65. Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting


66. TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents


67. Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau


68. Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation


69. CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine


70. From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving


71. Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events


72. Tutoring Large Language Models to be Domain-adaptive, Precise and Safe


73. FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics


74. LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs


75. Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees


76. Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World


77. PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models


78. OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling


79. R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction


80. A Compact Stance-Indexed Anterior-Posterior COP Representation for Parkinson’s Disease Classification from Plantar VGRF


81. AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows


82. Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model


83. When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used


84. ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures


85. From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis


86. DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks


87. CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy


88. ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control


89. Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems


90. Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale


91. Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework


92. A Survey on the Linear Representation Hypothesis



94. Generative Embodied Multiple Behavior Control Systems for Human-like Agents


95. Self-Organizing Agent Teams Learn to Reason Together


96. Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA


97. Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation


98. GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment


99. MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators


100. AutoGym: Blueprint-First Generation of Verifiable Agent Gyms


101. EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability


102. IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law


103. Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus


104. The Wisdom of Artificial Deliberative Crowds


105. Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation


106. Goal-driven Variant Categorization


107. Learning 3D biophysical cell properties from 2D images and cell-population statistics


108. Social Influence and the Allocation of Scientific Attention in AI Populations


109. PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation


110. An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users


111. Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models


112. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay


113. WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory


114. DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation


115. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses


116. DolphinBench: Mapping the Pareto Frontier of Agent Memory


117. Rare Event Estimation via Iterative Unalignment


118. Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences


119. Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks


120. Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization


121. Visuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning


122. OSWorld-Pro: Process-based Evaluation for Computer Use Agents


123. SE(3) Neural Potential Fields for 6-DoF Trajectory Planning Directly from Images Without Explicit 3D Reconstruction


124. When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting


125. Small-world Networks of Agents Brainstorm AI Risks to Support Ideation


126. SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture


127. Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI



129. Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection


130. When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs


131. PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning


132. NPU Accelerator: Quantized Real-Time Vehicle Detection on PYNQ-Z1 Using FINN


133. Enhancing Transformer Representations of Symbolic ODE Expressions


134. LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot


135. A digital-twin framework for forecasting treatment-day imaging with contour uncertainty in adaptive proton radiotherapy


136. Reasoning Topology Matters: A Controlled Study of LLM-Based Cybersecurity Analysis


137. “MeBo Leaves a Piece of You Behind”: Designing a Relational Voice-Based Memory Companion for Older Adults


138. Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference


139. What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization


140. Understanding Hyperspherical Geometry of ECAPA-TDNN Embedding and Its Impact on Zero-Shot Voice Conversion


141. Trust in Edge-Enabled IoT Security: Features, Challenges and Research Directions


142. Touch2Robot: Robot Touch in the Human Demonstration Loop


143. Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement


144. iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs


145. Annie, Are You Okay? How Style- and Context-Based Personalization Shape AI-Assisted Decision-Making


146. From Semantic Decisions to Feasible Trajectories: Self-Evolving LLM-Guided Optimal Control for Narrow-Space Parking


147. Augmented Hypothesis Testing with Persona-Based LLM Simulations


148. FedMust: Semi-supervised Multi-task Student-Teacher Federated Learning for Multi-organ CT Segmentation


149. GraphToolbox: A Configurable Python Framework for Graph Neural Network Forecasting


150. Overlay_dx - Automating forecasting evaluation


151. $t_0$: A Time-Series Foundation Model for Forecasting with Context


152. QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation


153. Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents


154. On Emergent Capabilities and Model Merging


155. Lifted Bellman Linear Programming for Offline Reinforcement Learning


156. AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos


157. VPRune: Efficient Training-free Pre-LLM Visual Token Pruning


158. Conduit: An Experience Data Plane for Distributed Reinforcement Learning


159. Do LiDAR Language Models Really Understand Spatio-temporal Relationships?


160. ActGov: Governing LLM Agent Actions via Policy-Constrained Validation


161. WPBench: A Comprehensive Benchmark for Wind Power Forecasting


162. FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding


163. Estimating Accurate Hand Pose in Camera Space with Vision Transformer


164. ARM: Attention with Routed-Memory for Learnable Sparse Control


165. Artificial Structure Function Search: Preserving Artificial Functional Connectivity for Structured Pruning


166. Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors


167. Information-Time Proximal Policy Optimization


168. URA-NER: A Unified Retrieval-Augmented Framework with Retrieval Alignment and Uncertainty Reduction for Low-Resource NER


169. DeceptionAnalyser: A Web-Based AI Tool for Performing Structured Deception Analysis with Argumentation Schemes and LLMs


170. Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection


171. Mitigating Entity Type Confusion in Cross-Domain NER via Multidimensional Quantification and Reasoning Enhancement


172. A Lean and Spec-Driven AI-Assisted Software Development Lifecycle for Applied AI Education: The AI-SDLC Approach


173. The Undetected Damage of Quantization on Retrieval and How to Fix It


174. Adapting Boltz-2 with limited experimental activity data improves early enrichment in virtual screening


175. KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation


176. TTSE: A Two-Track Online Self-Evolution Framework


177. Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection


178. vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation


179. MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents


180. Taramandal-GPT: Enhancing Astrodynamics Problem-Solving with Knowledge Retrieval and Structured Thinking


181. Hessian Rank Constraint for Learning Structure of Nonlinear Latent Variable Models


182. Memory vs. Context? Influential Factors of Factual Recall in Language Models


183. Opinion Leader Dynamics: How Sparse Attention Shapes Token Clustering


184. MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents


185. TAC-Time: Texts as Channels For Multimodal Time Series Forecasting


186. Graded-Relevance Composed Multimodal Retrieval for E-commerce Visual Search at Scale


187. STAR: Scene- and Task-Aware 4D Radar Preprocessing Towards End-to-End Cognitive Radar


188. Data Agents: Agentic Data Systems


189. Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding


190. ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation


191. Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective


192. WidgetVA: A Widget-Centric Framework and Benchmark for Agentic Visual Analytics


193. FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention


194. From Bits to Beliefs: Recoverable Semantic Fingerprints for Black-Box Verification of Large Language Models


195. From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning


196. MECT: Mixture of Experts with CNN-Transformer Network for Speaker verification


197. InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection


198. Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models


199. RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations


200. MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes


201. Djinnlang: Higher-Level Programming by Unambiguous Specification with an LLM in the Compiler


202. HaikuS2S: A Cascaded System For Responding In Verse


203. MCPGen: Benchmarking LLMs on Executable MCPWorkflow Development


204. ReVeal: A Reconstruction-Aware Real-to-Sim Framework for VLA Policy Evaluation


205. Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges


206. Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs


207. SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses


208. GLR-MM: Graph-Based Global-Local Reconstruction for Robust Multimodal Chest X-ray and EHR Representation Learning under Missing Modalities


209. VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking


210. From UNDRR Reports to Event Records: Schema-Constrained LLM Extraction of Georeferenced Disasters


211. Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking


212. FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model


213. Belted Engression: Sufficient Dimension Reduction for Generative Distributional Regression


214. TriFleetRCA: On-Premise LLM Root Cause Analysis for Kubernetes


215. GRACE: Grounded Adversarial Reasoning over Canadian Law


216. STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification


217. When the Agent Becomes the Kernel: A Systematization of Security on the Path to AI-Native Operating Systems


218. Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification


219. A multi-temporal dataset for mapping burned areas in the Brazilian Cerrado using time series of remote sensing imagery


220. Smoothed Analysis of Inconsistent A*


221. Which Constraints Are Missing? Ask the Verifier: Graded Rewards for Constraint-Following Music Generation


222. Spiking Neural Network Actor-Critic Proximal Policy Optimization Control for Autonomous UAV Navigation Through Constrained Openings in Civil Infrastructure and Buildings


223. PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding


224. PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models


225. StyleAT: Defending Face Recognition Against Semantic Attacks


226. Cost-Aware Reinforcement Learning with Action Masking and Projection for Battery Energy Storage Dispatch under Suppressed-Spread Market Shifts


227. Listen Then Reason: Perception-Grounded Test-Time Reinforcement Learning for Large Audio-Language Models


228. On the Efficiency-Safety Dilemma in Large Reasoning Models


229. ARID: A Deployable Edge AI System for Structured Information Extraction from Industrial Maintenance Work Orders


230. MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model


231. Modeling Clinical Workflow for SYNTAX Scoring from Coronary Angiography Videos


232. Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements


233. SemDHT: Certified Semantic Discovery for Peer-to-Peer Agent Networks over Exact-Key DHTs


234. Predicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition


235. Heating in human-HVAC interaction for smart homes: An interdisciplinary overview


236. CE$^4$L: Continual Ego, Exo, and Ego-Exo Learning


237. RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents


238. Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations


239. RLVR$^{2}$: Reinforcement Learning with Verifiable Rubric-based Ranking


240. WaveletECO: A Closed-Loop Physical ECO Platform and a Specialized Local Language Model


241. PhysReflect: Geometry and Perception Guided Diffusion for Physically-Plausible Mirror Reflections


242. RiverVLN: Phase-Grounded Temporal Vision–Language Navigation for Unmanned Surface Vehicles


243. Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability


244. OmniEcho: Spatial Audio Understanding for Embodied Agents


245. Enhancing Shrimp Disease Detection via Deep Learning and Data Refinement for Resilient Aquaculture


246. Blind Thermodynamic Ontology Discovery from Anonymous Experiments


247. Bayesian Filtering in Physical Systems via Test-time Trained Flow Matching


248. Discovering Physical Representation Languages


249. Human-guided physics-constrained AI agents construct an auditable model of soil-plug evolution


250. A Patient World Model for Early Forecasting of Digital Health Campaign Outcomes: Capabilities and Limits


251. Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference Services


252. Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines


253. ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs


254. Semantic Candidate-Job Matching: A Comparative Evaluation of Dense Embedding Models in Hybrid Retrieval


255. Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment


256. Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning


257. Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents


258. LLMs as Linguistic Chameleons: Decoupling Semantics and Structure for Privacy-Preserving Communication


259. GrapeSplat: Geometry-Grounded Reconstruction via Amalgamated Pose-Free Encoding for Feed-Forward 3D Gaussian Splatting


260. EquiSELD: Efficient training of equivariant sound event localization and detection networks


261. Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics


262. QwenVLConnector: A Fast, Unified Medical VLM Chatbot for Fine-Grained Clinical Perception and Text Generation


263. MM-ContextFold: Context Folding for Multimodal Agentic Retrieval


264. Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures


265. DiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic Simulation


266. Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone


267. MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs


268. From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness


269. Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity


270. Watching Quantum Models Think: Hilbert-Space Interpretability in Quantum Transformer Blocks


271. Interpretable Multi-Hypersphere Deep Anomaly Detection for Open-set Supervised Anomaly Detection


272. Dual-Locking Learned AI Models: A PIN-Based Sparse QIM Watermarking and Adaptive Index Permutation Approach


273. A Horizon-slicing Approach to Minimum Obstacle Displacement Planning for Robot Navigation


274. Automatic multimodal UX improvement recommendations from LLM agent user simulations


275. General Collaborative Intelligence: Architecting Cognition for Resilient Multi-Agent Ecosystems


276. When Agentic Trust Crosses Organizational Boundaries: Structural Externalization and a Reference Model for Trust Evidence


277. Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems


278. RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling


279. NostrAgent: A Decentralized Identity and Delegation Architecture for Sovereign Agentic Systems


280. An Evolutionary Agentic Approach for Open-ended Image Quality Perception


281. Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling


282. An Iterative LangGraph Agent for Text-to-SQL: Natural Language Access to the Chicago Crime Database


283. AVTR-1: Open Stack for Real-Time Interactive Avatars


284. Are Coreset Selection Methods Worth Their Cost?


285. Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion


286. Block-Sparse Attention with Semantic-Geometric Decoupled Routing


287. The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI


288. Per-Query Gating of LLM Rerankers for Multi-Hop Retrieval


289. Towards Full Pipeline FP8 Reinforcement Learning for LLMs


290. The Moral Check: Strategic AI Governance for the Pacing Problem


291. Planning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous Driving


292. Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs


293. Testing the Construct Validity of a Functional Valence Axis in LLM Agents


294. SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery


295. The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents


296. Commonsense-Grounded Path Planning from Abstract Instructions


297. AlexandriaX 2026: The First Shared Task on Dialectal Arabic Machine Translation


298. Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation


299. SelfOp: An Optimization Algorithm for Self-Improving Security Agents


300. ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding


301. Parameterized Dense-Sparse Fusion for Hybrid Retrieval: Tuning a Rank-Score Mix on BEIR SciFact with Qdrant


302. Beyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix


303. MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning


304. From Code to Requirements: Agentic Reverse Engineering of Business Rules at Enterprise Scale


305. Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments


306. LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator


307. Vision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD Modeling


308. From Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference Implementation


309. Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation


310. HIGenNTO: Scalable Humanoid Interaction Generation via Noise-Space Trajectory Optimization


311. From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly


312. Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation


313. Fairly Compensated Distributed Information Retrieval and Augmentation for AI Agents


314. Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act


315. SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning


316. Beyond the Leaderboard: Counterfactual Diagnosis of End-to-End and VLA Driving Policies Under Domain Shift


317. Zero-Trust Authorization and Discovery for Enterprise MCP


318. Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer


319. The Ups and Downs of Backprop Weights


320. FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation


321. A Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support


322. Toward Personalized Sleep Guidance from Wearable Data Using Language Models


323. Contextual Causality with Large Language Models: A Survey


324. MarsRecon: Self-Supervised and Multimodal Surface Representations for Mars


325. Forecasting Intrathecal Tracer Enhancement from Pre-Contrast Brain MRI: Direct Regression versus Flow Matching


326. Initial Evaluation of Potential Bias in Coverage of Humans in Wikidata


327. Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch


328. Dimensionality reduction for AI based hyperspectral image classification based on XAI


329. AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation


330. Artificial Neural Networks as Surrogate Models in Black Box Optimization


331. Visual Graph Reasoning via Knowledge Compilation


332. GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents


333. Authority-Preserving Evaluation of Medical Vision-Language Assistants


334. On The Robustness-Resolution Tradeoff In Temporal Quantization Of Event Streams


335. Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models


336. ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI


337. Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection


338. Rethinking Streaming Video Diffusion Model: Context, Execution, and Training


339. Brain-to-Image Generation: Reconstructing Visual Stimuli from EEG using Generative Adversarial Networks


340. Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening


341. Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework


342. Hi-Singers: A Comprehensive High-Quality Dataset for Expressive Audio-Driven Singing Head Synthesis


343. Large language models in medical time series analysis


344. Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection


345. Which Part of the Context Layer Does the Work? Separating Semantic Content from Retrieval Scaffolding in Text-to-SQL Agents


346. RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks


347. Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning


348. DIPLOMAT: Dialogue-Span-Aware Direct Preference Optimization for Polite Persuasive Workplace Negotiation Dialogues


349. Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations


350. Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation


351. CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models


352. CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation


353. Predictors and Orchestrators: Parsimonious Machine Learning within an Agentic AI Harness for Multi-Horizon Karst Aquifer Forecasting


354. A Tutorial on Prompt Engineering: From Messy Thoughts to AI Workflows


355. Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews


356. CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents


357. The Corroboration Illusion: When More News Makes LLM Forecasts Less True


358. Do Chess Explanations Reflect Model Decisions? Behavioral and Token-Level Tests of LLM Reasoning Faithfulness


359. Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing


360. Replay-Gated Neural Execution: Decoupling Persistent Behavioral Specifications from Neural Realizations in Frozen Language Models


361. H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution


362. Knowledge Graph-Augmented Ambient AI for Clinical Note Generation


363. Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks


364. Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark


365. Team DArgk at the 2026 ELOQUENT lab for evaluating generative language model quality: Residuals of Humanity: AI Detection Evasion via GRPO Fine-Tuning


366. Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B


367. Toollery: Scaling LLM Agents to Thousands of Skills and Tools


368. Replicating the Geometry of Emotion Representations in a Base Open-Weights Model


369. An Implant-to-Wearable IR-UWB Transmitter-Receiver Architecture and Layered Protocol for High-Density Brain-Computer Interfaces


370. Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems


371. PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations


372. Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers


373. The Role of AI in Online Reviews


374. Dissecting Hierarchical Reasoning Models: A Mechanistic Study


375. EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines


376. The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation


377. A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents


378. SegTSim: A Big Data Driven Segmented Temporal Simulation Framework for Heterogeneous Multivariate Systems


379. DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction


380. Beyond Task Completion: Training Capable and Safe Computer-Use Agents


381. Contrastive World Models


382. Multiple latent orderings better predict language model preferences


383. Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs


384. Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models