전체 AI 논문 - 2026-05-28

1. Calibrating Conservatism for Scalable Oversight


2. CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models


3. SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks


4. CubePart: An Open-Vocabulary Part-Controllable 3D Generator


5. CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning


6. Utility-Aware Multimodal Contrastive Learning for Product Image Generation


7. AlphaTransit: Learning to Design City-scale Transit Routes


8. Multi-Adapter Representation Interventions via Energy Calibration


9. LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?


10. OpenURMA: A Clean-Room Open Implementation of the Unified Bus Protocol


11. Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor


12. Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI


13. The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic


14. TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning


15. VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora


16. DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution


17. An LLM-Based Assistance System for Intuitive and Flexible Capability-Based Planning


18. AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation


19. The Ethics of LLM Sandbox and Persona Dynamics


20. Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation


21. LACUNA: Safe Agents as Recursive Program Holes


22. Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution


23. Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability


24. MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation


25. Continual Model Routing in Evolving Model Hubs


26. A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enhancing Stability in Multimodal Sentiment Analysis


27. Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns


28. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks


29. Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations


30. Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning


31. Cultural Binding Heads in Language Models


32. Do Agents Know What They Can’t Do? Evaluating Feasibility Awareness in Tool-Using Agents


33. Entropy-aware Masking for Masked Language Modeling


34. Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection


35. GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven Financial Forecasting


36. Benchmarking AI for low-resource contexts: Thinking beyond leaderboards


37. ProvMind: Provenance-grounded reasoning for materials synthesis


38. From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints


39. Diffusion Large Language Models for Visual Speech Recognition


40. GONDOR to the Rescue: Satisficing Planning with Low Memory


41. DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes


42. Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning


43. Measuring Progress Toward AGI: A Cognitive Framework


44. HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs


45. You Live More Than Once: Towards Hierarchical Skill Meta-Evolving


46. Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs


47. From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence


48. CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict


49. Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning


50. Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement


51. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets


52. Plan Before Search: Search Agents Need Plan


53. FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models


54. Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains


55. SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models


56. An Enhanced Large Neighborhood Search Approach for the Capacitated Facility Location Problem with Incompatible Customers


57. From Fact Overwriting to Knowledge Evolution: Causal Editing via On-Policy Self-Distillation


58. Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation


59. REED: Post-Training Representation Editing for Cross-Domain Linguistic Steganalysis


60. Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR


61. ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research


62. Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning


63. Global Policy-Space Response Oracles for Two-Player Zero-Sum Games


64. Entropy Distribution as a Fingerprint for Hallucinations in Generative Models


65. AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?


66. PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management


67. When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?


68. Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers


69. Learning When to Optimize: Verified Optimization Skills from Expert GPU-Kernel Lineages


70. The Illusion of Opting in AI-Mediated Consequential Decisions


71. Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents


72. Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning


73. Localizing Input Uncertainty Quantification for Large Language Models via Shapley Values


74. OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings


75. Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning


76. OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents


77. Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting


78. Deconstructing Spatial Complexity: Hierarchical Decomposition for LLM Spatial Reasoning


79. Data-Efficient On-Policy Distillation for Automatic Speech Recognition


80. Do Clinical Models Change Treatment Decisions?


81. Gradient Step Plug-and-Play Model for Dental Cone-Beam CT Reconstruction


82. CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models


83. Human-like in-group bias in instruction-tuned language model agents


84. Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification


85. Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction


86. Examining Agents’ Bias Amplification versus Suppression in Multi-Agent Systems


87. BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization


88. MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing


89. Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information


90. ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay


91. BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models


92. Verifiable Benchmarking of Long-Horizon Spatial Biology


93. MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents


94. Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG


95. MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation


96. Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings


97. PetroBench: A Benchmark for Large Language Models in Petroleum Engineering


98. MIRA: A Bilingual Benchmark for Medical Information Response Audit


99. Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback


100. Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training


101. An Empirical Audit of k-NAF Budget Accounting for Anchored Decoding


102. Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure


103. AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios


104. STAB: Specification-driven Testing for Algorithmic Bottlenecks


105. Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations


106. The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces


107. From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection


108. Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential Planning


109. DiagramRAG: A Lightweight Framework to Retrieve Scientific Diagram for Figure Generation


110. Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows


111. Show, Don’t TELL: Explainable AI-Generated Text Detection


112. SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats


113. Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization


114. Dr-CiK: A Testbed for Foresight-Driven Agents


115. SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment


116. A Unified Framework for the Evaluation of LLM Agentic Capabilities


117. PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management


118. Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness


119. AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models


120. FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research


121. C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning


122. MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents


123. When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models


124. TCP-MCP: Landscape-Guided Co-Evolution of Prompts and Communication Topologies for Multi-Agent Systems


125. EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA


126. Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions – A Governance Framework for High-Stakes AI Systems


127. Revealing Algorithmic Deductive Circuits for Logical Reasoning


128. EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents


129. Constrained Auto-Bidding via Generative Response Modeling


130. GraD-IBD: Graph Representation Learning from Diagnosis Trajectories for Early Detection of Inflammatory Bowel Disease


131. A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test


132. A Query Engine for the Agents


133. Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles


134. Auditable Decision Models with Learned Abstention and Real-Time Steering


135. Got a Secret? LLM Agents Can’t Keep It: Evaluating Privacy in Multi-Agent Systems


136. PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft


137. SkillGrad: Optimizing Agent Skills Like Gradient Descent


138. Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration


139. A Policy-Driven Runtime Layer for Agentic LLM Serving


140. Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking


141. DeepSciVerify: Verifying Scientific Claim–Citation Alignment via LLM-Driven Evidence Escalation


142. Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models


143. Cross-Entropy Games and Frost Training


144. Behavioural Analysis of Alignment Faking


145. Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems


146. Reasoning and Planning with Dynamically Changing Norms


147. Laguna M.1/XS.2 Technical Report


148. Voluntary Collusion with Secret Tools in Competing LLM Agents


149. Cyberbullying Governance on Social Media: A Unified Framework from Content Identification to Intervention


150. You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal State Intervention


151. Agyn: An Open-Source Platform for AI Agents with Scalable On-Demand Execution, Agent Definition as a Code, and Zero-Trust Access


152. Discovery Agents for Real-Time Analytics: Toward Proactive Insight Systems


153. LaneRoPE: Positional Encoding for Collaborative Parallel Reasoning and Generation


154. RULER: Representation-Level Verification of Machine Unlearning


155. Why LLMs Fail at Causal Discovery and How Interventional Agents Escape


156. DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents


157. On the Origin of Synthetic Information by Means of Steganographic Inheritance


158. Soro: A Lightweight Foundation Model and Chatbot for Tajik


159. Identifying and Understanding Human Values in Text: A Tailorable LLM-based Architecture


160. Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation


161. OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration


162. Skill-Conditioned Gated Self-Distillation for LLM Reasoning


163. Do Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval


164. Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents


165. Rethinking Memory as Continuously Evolving Connectivity


166. Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL


167. Preference-Shaped Expected Hypervolume and R2 Improvement: Exact Computation and Monotonicity


168. Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text


169. BIRDNet: Mining and Encoding Boolean Implication Knowledge Graphs as Interpretable Deep Neural Networks


170. MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems


171. IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents


172. Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study


173. A Fresh Look at Lamarckian Evolution and the Baldwin Effect


174. Deep Learning Strain Estimation: Is Physics-Based Simulation the Solution?


175. Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images


176. AI in the Workplace: The Impact of AI on Perceived Job Decency and Meaningfulness


177. Sense Representations Are Inducible Interfaces


178. The Attentional White Bear Effect in Transformer Language Models


179. Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking


180. Measuring Form and Function in Language Models


181. Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification


182. Online Irregular Multivariate Time Series Forecasting via Uncertainty-Driven Dual-Expert Calibration


183. Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News


184. Position: Retire the “Positive Backdoor” Label – Secret Alignment Requires Strict and Systematic Evaluation


185. Thermodynamic properties of chemically disordered compounds via AI-driven estimation of partition function with the PULSE method


186. Models That Know How Evaluations Are Designed Score Safer


187. Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem


188. SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving


189. Efficient Pre-Training of LLMs through Truncated SVD Layers


190. Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression


191. Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs


192. A Multi-dimensional Framework for Evaluating Generalization in EEG Foundation Models


193. Token Optimization Strategies for LLM-Based Oracle-to-PostgreSQL Migration


194. Stochastic Gradient Descent with Momentum is Algorithmically Stable


195. Do LLMs Favor Their Providers? Measuring Vertical Integration Bias in Code Generation


196. Learning Theory of the SVRG: Generalization and Convergence Analysis


197. Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets


198. Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification



200. SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs


201. The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment


202. BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers


203. Bayesian Gated Non-Negative Contrastive Learning


204. Anomaly as Non-Conformity via Training-Free Graph Laplacian Energy Minimization


205. VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs


206. ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation


207. CLANE: Continual Learning of Actions on Neuromorphic Hardware from Event Cameras


208. Score Based Error Correcting Code Decoder


209. Improving Evaluation of Recombination-based Cartesian Genetic Programming


210. Learning the Error Patterns of Language Models


211. Multi-Agent LLM-based Metamorphic Testing for REST APIs


212. Identifying Explicit Parsimonious Piece-wise Polynomial Relationships in Industrial time-series: Application to manipulator robots


213. Hybrid Neural World Models


214. Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models


215. Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning


216. How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving


217. ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation


218. PrunePath: Towards Highly Structured Sparse Language Models


219. GUI Agents for Continual Game Generation


220. IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage


221. VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer


222. SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping


223. Pruning and Distilling Mixture-of-Experts into Dense Language Models


224. Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation


225. Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension



227. FLORO: A Multimodal Geospatial Foundation Model for Ecological Remote Sensing Across Sensors and Scales


228. QuITE: Query-Based Irregular Time Series Embedding


229. Performance and Explainability Requirements of Evolutionary Algorithms in Real-World Physics-Informed Optimization


230. DEPART: DEcomposing PARiTy across Multilingual LLMs


231. DeltaMCP: Incremental Regeneration via Spec-Aware Transformation for MCP servers


232. SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents



234. MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content


235. EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction


236. Revisiting Change Detection Methods for their Application to Serac Fall Time-Lapse Monitoring


237. SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter


238. Mind the Gap: Mixtures of Gaussians in Approximate Differential Privacy


239. StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment


240. PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting


241. I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors


242. Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts


243. On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective


244. Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts


245. SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection


246. VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning


247. MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models


248. Learning Compositional Latent Structure with Vector Networks


249. Integrated and Cross-Architecture Interpretation of LLM Reasoning


250. Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution


251. Learning to Assign Prediction Tasks to Agents with Capacity Constraints


252. Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models


253. Geometry-Correct Diffusion Posterior Sampling with Denoiser-Pullback Curvature Guidance and Manifold-Aligned Damping


254. KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs


255. Periodic RoPE for Infinite Context LLMs


256. Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses


257. Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors


258. ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning


259. Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations


260. When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?


261. Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study


262. Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking


263. ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations


264. The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages


265. SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control


266. VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild


267. SPAR: Support-Preserving Action Rectification


268. From Detection to Mechanism: Cross-Attention Graph Neural Networks Enable Drug-Drug Interaction Type Prediction An Ablation Study with Acetylsalicylic Acid Validation


269. DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification


270. Fine-Tuned LLM as a Complementary Predictor Improving Ads System


271. FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation


272. Snippet-Driven Supply Chain Discovery with LLMs: Scaling Visibility in China


273. LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation


274. Symmetry Defeats Auditing


275. Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security


276. ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions


277. Turning Video Models into Generalist Robot Policies


278. Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models


279. ChildEval: When large language models meet children’s personalities


280. Locality-Aware Redundancy Pruning for LLM Depth Compression


281. Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict


282. UniMaia: Steering Chess Policies with Language for Human-like Play


283. Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning


284. Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought


285. High-Fidelity Industrial Crash Dynamics Prediction via Geometry-Aware Operator Learning with Memory-Efficient Low-Rank Attention


286. Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions


287. Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection


288. Worker Disagreement Reveals Sharp Directions in Local SGD


289. HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning


290. UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind


291. CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text


292. Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning


293. Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers


294. Backdoor Attacks on Fault Detection and Localization in Cyber-Physical Systems


295. Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting


296. How the Optimizer Shapes Learned Solutions in Equivariant Neural Networks


297. Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment


298. Developing an Intelligent Job Recommendation System Using Semantic Retrieval and Explainable AI Techniques


299. Cultural Fidelity in English-to-Hindi Translation: A Preservation-Fluency Frontier for Gender Recoverability


300. Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression


301. Trinity: Unifying Class-Agnostic Terrain and Semantic Segmentation for Unstructured Outdoor Environments by Leveraging Synthetic Data


302. Supervised Distributional Reduction via Optimal Transport and Dependence Maximization


303. Not All NVFP4 QAT Recipes Are Equal: How Architecture and Scale Shape Model Quality for Anomaly Segmentation



305. The Energy Blind Spot: NVIDIA’s Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution


306. Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks


307. The Future of Facts: Tracing the Factual Generation-Verification Gap


308. On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note


309. Clinical Validation of the Melanoscope AI Mobile Dermoscopy Clinical Decision Support System


310. Detection Without Correction: A Two-Parameter Decomposition of Multi-Stage LLM Pipelines


311. Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?


312. Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems


313. HARP: Measuring Harm Amplification in Multi-Agent LLM Systems


314. Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels


315. Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer


316. Debate Helps Weak Judges Reward Stronger Models


317. Energy-Structured Low-Rank Adaptation for Continual Learning


318. BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving


319. Resource-Constrained Affect Modelling via Variance Regularisation Pruning


320. Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective


321. HEAL: Resilient and Self-* Hub-based Learning


322. AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications


323. Detect by Yourself: Self-Designing Agentic Workflows for Few-Shot Graph Anomaly Detection



325. Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility


326. AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems


327. AdaMerge: Salience-Aware Adaptive Token Merging for Training-Free Acceleration of Vision Transformers


328. Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU


329. When prompt perturbations break your A/B test: A valid statistical test for generative surveying


330. Generic Interpretation Approach for Transformer Models Incorporating Heterogenous Attention Structures


331. Checking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval


332. RAGe: A Retrieval-Augmented Generation Evaluation Framework


333. A Systematic Evaluation of Retrieval-Augmented Generation and Language Models for Space Operations


334. Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline


335. Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit


336. MGRetrieval: Memory-Guided Reflective Retrieval for Long-Term Dialogue Agents


337. RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?


338. When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference


339. Heterogeneous Multi-Agent Modeling for Measurement and Network Analysis of the Data Service Market


340. FD-RAG: Federated Dual-System Retrieval-Augmented Generation


341. Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey


342. Ocean4Rec: Offline LLM-Derived OCEAN Profiles for Request-Time VOD Reranking


343. Quantum Machine Learning-based 6G edge Network: Enabling Adaptive Communication and Model Aggregation


344. Can Quantum Federated Learning Withstand Circuit-Level Backdoors?


345. Ligand-Conditioned Discrete Diffusion for Protein Sequence-Structure Co-Design


346. Advancing Direct Training for Spiking Neural Networks with Circulate-Firing Neurons and Learnable Gradients


347. STARS: Spike Tail-Aware Relational Synthesis for ANN-to-SNN Data-Free Knowledge Distillation


348. Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects


349. Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams


350. LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students’ written reflection assignments


351. REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading


352. Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis


353. Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framework for Future-Facing Learning


354. Short-Term Gain, Long-Term Fragility: AI Labor Substitution and the Erosion of Sustainable Capability


355. Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named


356. Informing AI Policy Assessment using Large-Scale Simulation of Interventions


357. Human-AI Collaboration for Estimating Scientific Replicability


358. StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation


359. Learning after COVID-19 and the ICT career aspirations: Are students entering the AI era with weaker skills?


360. EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter AdaptationTarget


361. Memory-Based vs. Context-Only Conditioning Produces Distinct Behavioral Patterns in Stateful Personalization


362. Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities


363. From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons


364. Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity


365. From Instructor to Collaborator: What a 90-Participant Study Reveals about Human-Agent Collaboration in a Mobile Serious Game


366. Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models


367. The Alignment Floor: When Persona Customization Is Safe


368. The Computational Boundary of Inference: Capability Internalization, Training, and the Turing Jump


369. BioELX: Cross-lingual Biomedical Entity Linking via Alias-based Retrieval and LLM Ranking


370. RAG-Coding: Enhancing LLM Medical Coding with Structured External Knowledge


371. Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models


372. LNN-PINN: A Unified Physics-Only Training Framework with Liquid Residual Blocks