전체 AI 논문 - 2026-09-02

1. Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers


2. Can LLMs Discover Scientific Laws in Real and Parallel Worlds?


3. EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation


4. When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation


5. Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement


6. Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers


7. EdiTikZ: Scientific Figure Editing from Revision Trajectories


8. Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations


9. EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems


10. SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding


11. Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades


12. LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting


13. Automated Event Log Generation from Unstructured Text Using Finetuned LLMs


14. A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation


15. Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems


16. Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents


17. Dual Process Motion Planning


18. Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations


19. Prompt-Robust Language Models: Which Training Strategies Work?


20. H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning


21. FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue


22. Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate


23. Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs


24. Space Generative AI with Solar Energy Harvesting


25. ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning


26. User Representation via Cross Multi-source Behavior Pre-training for Mobile Games


27. WorldBench: Culturally Grounded Benchmark for Multilingual Agents


28. QILP-0: Constructing Observational Declarative Twins of Quantum Circuits


29. AgentFactory: Towards Automated Agentic System Design and Optimization


30. Data-Driven Persona-Conditioned Agents for A/B Test Simulation


31. Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees


32. Figures as Programs: Recursive Generation of Editable Scientific Figures


33. CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins


34. Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance


35. VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don’t Mean Preferences


36. RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation


37. In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?


38. CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training


39. CacheBridge: Efficient Cross-Model KV Cache Transfer


40. Denoising Diffusion Generative Models Secretly Calculate Attentions


41. Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning


42. FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study


43. Beyond the Clock: Measuring the Value of Adaptive Revision


44. Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems


45. Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources


46. Towards Generalizable Visually Grounded Exploration of Household Devices


47. FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation


48. Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents


49. AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation


50. One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning


51. Towards a Reliable and Practical Eval Pipeline


52. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?


53. When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection


54. DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory


55. Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks


56. S^3martCirc: Self-supervised Smart Circuit Discovery


57. ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents


58. Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs


59. Agentic Empirical Asset Pricing: Methodological Foundations


60. SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification


61. A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies


62. ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything


63. Value Over Language Model: Detecting Original Contribution in Writing


64. Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs


65. Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets


66. SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task



68. DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation


69. REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows


70. Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs


71. Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing


72. Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts


73. Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random


74. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs


75. VoiceLongMemEval: Do Assistants Remember How You Sounded?


76. Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation


77. ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation


78. When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency


79. CoVer: Conflict-Aware Claim Verification


80. Wave Function Backpropagation with Explicit Temporal-Interval Dynamics


81. Validity-Aware Jailbreak Evaluation for Large Language Models


82. The Privacy-Hallucination Tradeoff in Differentially Private Language Models


83. EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models


84. Towards a Belief-Based World Model for LLM Agents


85. mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers


86. Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations


87. SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents


88. SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation


89. Dependency-Aware Chain-of-Thought Compression for Financial Reasoning


90. RestoreBench: Can AI Agents Restore Power Flow Convergence?


91. Dr. Claw: An AI Scientist Workspace for Vibe Research


92. A Stable Aggregation Method for Quantum Federated Learning


93. Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models


94. SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning


95. Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective


96. The Assistant’s Ideal Self


97. The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems


98. Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching


99. The Answer Is Not the Argument


100. Hypotheses-Guided Self Distillation for Continual Personalization


101. Authority Bias in Conversational Search Engines for Academic Paper Recommendation


102. Invalidation Contracts for Cross-Episode Agent Memory


103. Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems


104. ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback


105. AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning


106. ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation


107. LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark


108. Asymmetries in Spontaneous and Instructed Deception


109. IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training


110. Recursive Criticality of AI Self-Improvement


111. Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations


112. Different representation learning objectives recover distinct latent structures from the same psychometric data


113. AI Morbidity and Mortality: A Framework for Clinical AI Failure Review


114. MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts


115. When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation



117. UI-Venus-2 Technical Report


118. SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces


119. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets


120. Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls


121. Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models


122. Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing



124. HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models


125. Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation


126. Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation


127. CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?


128. The Rise of Verbal Reinforcement Learning


129. Mechanism Design for Alignment and Control


130. Designing Proactive Thought Partners for Writing


131. Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs


132. From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification


133. H3-World: Turning Language Understanding into World Control


134. Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories


135. BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net


136. A Mathematical Theory of Reusable Neural Bases for Network Compression


137. Can LLMs Design Video Coding Tools? A Case Study on Planar Mode


138. Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data


139. TempCloze: Can Video-LLMs Identify the Missing Middle?


140. LatentPress: Context Compression Beyond Text and Vision


141. Optimizing Byzantine Node Placement in Decentralized Federated Learning


142. Rethinking Learnability in Offline Data-driven Optimization


143. GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions


144. Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents


145. When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning



147. Learning Sparse Decision Trees via Transformer Variational Auto-Encoders


148. Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading


149. Provably Safe Sim-to-Real Transfer


150. Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching


151. Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity


152. PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction


153. CHARM: Character Hallucination for Multicultural Role Play Benchmark


154. Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs


155. Probing Factual Knowledge Transfer with Training Data Interventions


156. Bandits in Prod: Hyperparameter Optimization at Inference Time


157. MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval


158. GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation


159. HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives


160. EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents


161. Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models


162. TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models


163. The Constitutional Coverage Trilemma in AI Governance


164. One Prompt Is Enough: Watermark Laundering Through Foundation Image Models


165. Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents


166. From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs


167. MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence


168. Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling


169. REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs


170. Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges


171. Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening


172. Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment


173. Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion


174. Superposed Latent Autoencoder


175. StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization


176. Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set


177. DNC-IMM: Early Lane-Change Intention Recognition via Neural Calibration Based on Driving Context Information


178. EDRAC: Benchmarking Arabic Dialect Reading Comprehension


179. Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation


180. StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions


181. Text-guided flow matching enables sample-efficient crystal structure generation


182. Lagged Coupling: Internal Representations Become Readable Before They Become Causal


183. HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation


184. From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion


185. ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives


186. Causal Evidentiary Governance for High-Risk Machine Learning Systems


187. On Synthesis of Metric Interval Temporal Logics


188. A Network Science Perspective on Evaluating Deep Graph Generative Models


189. SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models


190. Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close


191. Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages


192. On the Human and Computer Alignment of Attribute-Based Music Matches


193. Semi-Supervised Virtual Staining via Morphology Preservation and Histopathological Realism Constraints


194. Disclosure-Gated User Simulation for Companion-Agent Evaluation


195. The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research


196. Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling


197. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding


198. Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches


199. DualStake: Dual-Path Confidence Calibration in Deep Research Agents


200. Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO


201. Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking


202. Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting


203. Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation


204. Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair


205. ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection


206. A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling


207. Probabilistic Model Checking of Autoregressive Neural Sequence Models


208. Replacing Training with Memory: Listwise Selection for Text-to-SQL


209. Visual Attention Faithfulness in Vision-Language Models is Heterogeneous


210. HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution


211. Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online


212. Agentic programs: an emerging form of scientific software in computational materials science


213. MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries


214. Instella-MoE Technical Report


215. Solaris: Towards Interfaces That Are Generated, Not Coded


216. VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM


217. Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures


218. Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models


219. Differentially Private Paired Table-Image Multimodal Synthesis


220. A Study of Hidden-State Optimization Order in Predictive Coding Networks


221. Visual Framing for News Stance Detection via Image Generation


222. EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction


223. TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data


224. Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity


225. Restrict, Don’t Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation



227. Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning


228. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems


229. Predicting Program Exit Code with LLMs and Programming Language Semantics


230. GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning


231. A Mathematical Framework for Legacy, Governance, and Decision Integrity in Enterprise AI


232. WiseSpec: Requirements-Driven Agents for Code Generation


233. EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection


234. EM^2Mem: Event-Centric Multimodal Memory for Large Language Models


235. Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers


236. Are We There Yet? Assessing Computer-Use Agents for Blind Users’ Accessible Interaction with Desktop Applications


237. The Safeguard Worked. Is the LLM System Safer?


238. The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space


239. RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces


240. Independent Reinforcement Learning in Discounted Markov Games


241. Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models


242. EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities


243. Exploring Collaboration between a language and a non-language agent


244. Higher Structures in Deep Learning


245. Operational Regimes in Non-Convex Optimization: A Multiplier-Based Taxonomy


246. Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective


247. HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference


248. Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach


249. Capability-Gated Language Models: Security Composes, Utility Does Not


250. (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement


251. Risk-Aware Decision-Making for Autonomous Overtaking: A World Model-Based Mixture-of-Experts Framework


252. FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos


253. Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You


254. Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning


255. Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure


256. Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning


257. Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces


258. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts


259. The Curse of Multilinguality in Lexical Normalization


260. A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension


261. Workload Identification with Physical Side Channels for AI Governance


262. Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains


263. WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization


264. Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer’s Disease Detection


265. Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems


266. CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships


267. CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction


268. Don’t Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes


269. QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization


270. Rock, Paper, Scissors, … Dynamite - A Model of Disruption from New Technologies


271. Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation


272. WHALE: A Simple Recipe for Joint Harness-Weight Optimization


273. Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost


274. Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models


275. Intelligent Edge Computing


276. Do General NLP Embeddings Capture Ontological Reasoning?


277. Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs


278. Flawed in Nature, Perfect through Evolution


279. Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy


280. Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding


281. Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence


282. Commit-first LLM judging inherits the judge’s own errors


283. Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation


284. KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training


285. RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks


286. AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis


287. Auditing Harness Tampering in Self-Improving Agents


288. Life Operators: a self-evolving framework for multiscale life modelling


289. Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy


290. OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization


291. Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents


292. Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning


293. Medical Causal Hypothesis Verification with Large Language Models


294. RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving


295. ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration


296. A Formal Analysis of Agent Payment Protocols


297. DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction


298. CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language


299. ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation


300. Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment


301. From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling


302. Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness


303. REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent


304. GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments


305. Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training


306. RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery


307. trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories


308. Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning


309. InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information