전체 AI 논문 - 2026-09-29
1. FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
- Authors: Hoyoung Lee , Suyeol Yun , Jack Haverty , Yunju Cho , Meesong Kim , Daekyung Park , Sumin Kim , Jihoon Kwon , Jasmine Jia Geng , Andrew Chin , Yin Luo , Edward Tong , Yu Yu , Zach Golkhou , Minkyu Kim , Igor Halperin , Young Cha , Alejandro Lopez-Lira , Chanyeol Choi , Yongjae Lee
- URL: https://arxiv.org/abs/2609.35744
- Abstract:
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution’s own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric’s expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts’ key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.
2. Shockingly Simple Self-retrospection Improves Agentic Models Without RL
- Authors: Jonathan Light , Christopher Zhang Cui , Jeonghye Kim , Roger Creus Castanyer , Emiliano Penaloza , Zhengyan Shi , Alessandro Sordoni , Marc-Alexandre Côté , Xingdi Yuan , Minseon Kim
- URL: https://arxiv.org/abs/2609.35741
- Abstract:
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO’s 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
3. Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
- Authors: Junru Zhu , Shiming Xie , Aime Lu Fan Chen , Xiaoqing Ding , Chunxin Tang , Ruoyu Qi , Yulang Fei
- URL: https://arxiv.org/abs/2609.35732
- Abstract:
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.
4. Reinforcing Agentic Creativity in Scientific Ideation with Night Science
- Authors: Priyanka Kargupta , Silviu Cucerzan , Shweti Mahajan , Allen Herring , Jiawei Han , Ryen W. White , Sujay Kumar Jauhar
- URL: https://arxiv.org/abs/2609.35706
- Abstract:
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
5. Reasoning with Continuous Latent Diffusion
- Authors: Xiang Cheng
- URL: https://arxiv.org/abs/2609.35694
- Abstract:
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: this https URL
6. Report: Progressive Disclosure of Agent Skills
- Authors: Guilin Zhang , Kai Zhao , Priyanka Mudgal , Waleed Ammar , Xiquan Cui , Xu Chu , Alet Blanken
- URL: https://arxiv.org/abs/2609.35692
- Abstract:
Users of Workday’s deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents’ capabilities. However, as an agent’s skills library grows in size, so does the agent’s operational cost. Progressive disclosure (lazy-loading) of skills as needed may reduce operational costs, but its impact on overall latency and skill-retrieval quality remains unclear. In this report, we investigate the impact empirically and find that progressive disclosure improves skill-retrieval quality but marginally degrades overall latency.
7. Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
- Authors: Christian Moya , Elliott Thornley , Guang Lin
- URL: https://arxiv.org/abs/2609.35677
- Abstract:
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.
8. PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
- Authors: Yangqin Jiang , Lingrui Xu , Chao Huang
- URL: https://arxiv.org/abs/2609.35671
- Abstract:
Mobile GUI agents operate through a perception–action loop: at each step they screenshot the device, invoke a vision–language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation—and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app’s GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld’s official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app’s navigation rather than one run, so it serves new tasks, not only repeated ones.
9. Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization
- Authors: Huzi Cheng , Zhewei Zhang
- URL: https://arxiv.org/abs/2609.35643
- Abstract:
Large Language Models can perform multi-step reasoning and improve task performance through different forms of intermediate computation, from token-based traces to computation carried out in latent space. However, a question remains open: do these different forms of thinking rely on the same underlying mechanism? To address this, we train and compare five variants of the same GPTNeoX backbone from scratch on an extended multi-hop reasoning task (ProsQA-Ext): a vanilla model, a Chain-of-Thought (CoT) model, a Pause Token model, and two latent-reasoning models that are optimized end-to-end without intermediate reasoning traces. We find that, strong in-distribution (ID) performance does not guarantee depth generalization. Vanilla, CoT, and Pause Token models solve ID problems well, but rely largely on local graph features and generalize poorly to out-of-distribution (OOD) problems with longer hops. In contrast, latent variants generalize better and show internal dynamics consistent with forward reachability propagation on the graph. Causal interventions and circuit analysis localize this computation to a sparse recurrent search circuit in the bottleneck latent model: an attention head retrieves graph relations, an MLP and the residual stream update the reachability state across recurrent steps, while multiple attention heads together then do the candidate matching. Together, these results show that different thinking mechanisms can learn distinct computational solutions, even at similar ID performance. In this setting, latent recurrence supports a reusable forward-search algorithm that generalizes beyond the training depth.
10. Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
- Authors: Shuyue Stella Li , Xiaochuang Han , Yulia Tsvetkov , Luke Zettlemoyer
- URL: https://arxiv.org/abs/2609.35641
- Abstract:
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate—GPT-Image-2.5—solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.
11. RIDE: Reference-Anchored Inference-Time Diffusion Editing for Scaffold Hopping
- Authors: Ruoxi Gao , Frazier N. Baker , Trieu Nguyen , Xia Ning
- URL: https://arxiv.org/abs/2609.35623
- Abstract:
Scaffold hopping is a critical task in drug discovery, which seeks to discover new, structurally distinct molecules that share key functional groups and similar 3D shape with a reference binding ligand. Existing diffusion-based scaffold hopping methods formulate the problem as conditional generation of scaffolds given the functional groups. However, they lack a principled mechanism to jointly enforce 2D structural novelty and preserve the 3D shape of the reference ligand. Here, we introduce RIDE, a Reference-anchored Inference-time Diffusion Editing framework for scaffold hopping. RIDE recovers the reference diffusion noise trajectory conditioned on the binding pocket and functional groups, selects an optimal trajectory segment for editing via noise perturbation, and conducts a value-guided scaffold sampling to generate new scaffolds. Extensive experimental results demonstrate that, compared to baselines, RIDE consistently generates scaffolds with lower 2D similarity and higher 3D similarity to the reference, with an average improvements of 11.7% and 7.3%, respectively. Further analysis reveals that RIDE can accommodate various reward functions, and can preserve 3D similarity even when this is not explicitly included in the reward. Two case studies illustrate RIDE’s ability to generate distinct scaffolds with different structures and properties, and its ability to introduce substantial 2D variation while maintaining very high 3D similarity. RIDE is publicly available at this https URL .
12. From cacophony to hierarchy: a principled framework for assessing AI consciousness
- Authors: Shamil Chandaria , Arvo Muñoz Morán , Fernando Rosas , Anil Seth , Henry Shevlin , Marcus Hutter , Thore Graepel , Adam Bales , Iulia Comsa , Murray Shanahan , Ruben Laukkonen , Morten Kringelbach , Chris Frith , Shane Legg
- URL: https://arxiv.org/abs/2609.35618
- Abstract:
The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a cacophony of competing theories that often talk past each other. Separating the hard problem from the mapping problem allows the deepest metaphysical disagreements to be set aside: granting that experience supervenes on a system’s organisation, the tractable question becomes at which grain of description that supervenience base sits. We extend Marr’s three levels of analysis into a five-level hierarchy of functional descriptions (behavioural, computational, intrinsic causal-structural, organismic, and organism-environment) grounded in supervenience, coarse-graining, and multiple realisability. The major theories of consciousness are positioned within this hierarchy according to which level they take to be critical, and for each level we develop operationalisable indicators and assess current AI systems against them. A Bayesian model then combines theoretical credences with indicator evidence into an overall credence in a system’s capacity for consciousness. In illustrative assessments, the verdict for current LLMs is driven as much by where theoretical credence is placed as by how the evidence is read: under different stipulated readings and credence distributions, assessments range from below 0.01 to roughly 0.8, showing sensitivity to assumptions. Finally, the consciousness indicators at each level closely overlap with the architectural features needed for general intelligence, suggesting that increasingly capable AI may become a stronger candidate for consciousness. The framework supports a structured agnosticism, in which theoretical commitments are made explicit, credences are updated as evidence accumulates, and assessments take the form of aggregated probabilities rather than verdicts.
13. TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science
- Authors: Chutong Yang , Xiyuan Zhang , Yu Huang , Boran Han , Soonho Kong , Shuai Zhang , Vihang Prakash Patil , Zhen Han , Michael Bohlke-Schneider , Bernie Wang
- URL: https://arxiv.org/abs/2609.35606
- Abstract:
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery, comprising 398 theorem-level challenges from 138 STOC and COLT 2026 papers. Expert-designed rules complete paper-specific context, preserve computational assumptions and quantitative guarantees, and withhold constructions when discovering an algorithm is part of the task. For each task, prover systems receive theorem statements and access to cited prior work. The pipeline supports fresh, versioned challenge batches from newly released papers. We evaluate ten model configurations from four families under direct inference and prover-verifier discussion, and compare four agent workflows under matched model-call opportunities. All evaluations use the full benchmark. In the model comparison, GPT-5.6 Sol max achieves the highest five-run verifier-accepted coverage at 23.6% after 10-round discussion. Discussion and repeated sampling improve coverage. In the separate agent comparison using GPT-5.5 xhigh, decomposition improves coverage over discussion, and agentic planning achieves the highest five-run verifier-accepted coverage at 25.4%. TCSAlgBench provides a refreshable testbed for measuring progress in model reasoning and studying how agent workflows support research-level proof discovery.
14. Signatures of semantic search in the activations of large language models
- Authors: Luke Leckie , Peter M. Todd , Jacob G. Foster
- URL: https://arxiv.org/abs/2609.35599
- Abstract:
When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production (“exploit”) and between-cluster switching (“explore”). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., “water”) increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.
15. Source-preserving alignment for robust evidence localization in scientific PDFS
- Authors: Zihao Liu , Wei Yang , Zixiao Dong , Chenshu Li , Longzhang Liu , Tao Tan , Hong Xie
- URL: https://arxiv.org/abs/2609.35588
- Abstract:
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, and fragmented items alter text sequences and geometry. We present a source-preserving alignment framework: normalize text for robust matching while preserving provenance for accurate localization. It aligns evidence with normalized page text, maps matches back to source-character spans, and renders only their geometry. When exact alignment fails, line-break-aware token alignment recovers supported spans while excluding unmatched noise. Experiments on 1,020 chemistry papers show that the framework achieves a 92.6\% quote-level automatic localization rate, compared with 43.6\% for text search and 19.1\% for a precomputed bounding-box baseline. Component ablation confirms distinct contributions from normalization and approximate token alignment, while human verification assesses the visual correctness of returned highlights. Overall, these results demonstrate that reliable evidence verification requires robust matching and precise localization within a shared source-preserving alignment representation.
16. IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
- Authors: Yung-Chin Chen , Chia-Yu Chen , Naveen Verma
- URL: https://arxiv.org/abs/2609.35586
- Abstract:
Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize these coupled sources of IMC error and often rely on costly search-based calibration. We introduce IMC-CLINIC (Coupled Loss-Informed Newton Iterations for Clipping), a clipping calibration framework based on an analytical surrogate for IMC MatMul output error. The surrogate jointly models operand quantization, accumulated clipping-induced bias, and ADC quantization, enabling efficient evaluation of its gradient and approximate curvature from a small calibration set. IMC-CLINIC jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5-11.5 percentage points over the grid search baseline while reducing calibration time by factors of 10.0-12.1. Its analytical surrogate closely tracks empirical IMC output error, and its optimizer is certified within 1% of the global optimum under the loss objective across all projections on two representative models.
17. Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
- Authors: Sidharth Pulipaka , Ansh Sharma , Stanislau Hlebik , Leonidas Raghav , Vyas Raina , Ivaxi Sheth , Mario Fritz
- URL: https://arxiv.org/abs/2609.35576
- Abstract:
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant’s persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.
18. Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
- Authors: Hyunwoo Yoo , Cassie Huang , Haebin Shin , Li Zhang , Gail L. Rosen
- URL: https://arxiv.org/abs/2609.35571
- Abstract:
While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ‘‘SMILES-to-PDDL’’ attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL generation, we achieve high success rates where end-to-end approaches fail. This provides evidence that a primary bottleneck lies in representation alignment rather than raw model capacity. Our structural analysis demonstrates that intermediate representations are essential in retrosynthesis planning, highlighting the importance of representation-centric design in future systems.
19. RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement
- Authors: Yaxin Du , Xiyuan Yang , Zhifan Zhou , Yujie Ge , Cheng Wang , Jiajun Wang , Sijie Chen , Zehui Liu , Yuxin Zhang , Weicheng Gu , Julian Zhang , Zixing Lei , Siheng Chen
- URL: https://arxiv.org/abs/2609.35561
- Abstract:
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0\% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.
20. From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
- Authors: Kangcheng Deng , Hui Cai , Jiacheng Lu , Chester Zhongshu Qian , Rui Sun , Beidi Luan , Jing Li , Daxin Jiang , Zuo Bai
- URL: https://arxiv.org/abs/2609.35559
- Abstract:
Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as search scaling. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.
21. BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
- Authors: Peilin Feng , Zhengyang Huang , Soujanya Poria
- URL: https://arxiv.org/abs/2609.35551
- Abstract:
In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model’s internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.
22. RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
- Authors: Bo Zhang , Yuchen Wang , Dongbai Li , Matthew Yu Heng Wong , Qingkai Zeng , Lijun Wang , Tien-Yin Wong , Peng Cui , Tianyu Liu
- URL: https://arxiv.org/abs/2609.35549
- Abstract:
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.
23. Continuous Context Management
- Authors: William Hoy , Jingxuan Fan , Nurcin Celik , Xu Pan
- URL: https://arxiv.org/abs/2609.35540
- Abstract:
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student’s initial model scores each sampled student action under the complete history reconstructed from that student’s rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
24. ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
- Authors: Kun Feng , Yuchen Fang , Yiyang Tan , Shuqi Gu , Yongxiang Zhao , Yu Liu , Xingyu Lu , Lintao Ma , Kan Ren
- URL: https://arxiv.org/abs/2609.35532
- Abstract:
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at this https URL .
25. MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
- Authors: Zihan Yu , Jiadong Zhang , Jialin Cheng , Jingtao Ding , Yong Li
- URL: https://arxiv.org/abs/2609.35515
- Abstract:
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal–mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.
26. SRHarness: A Harness for Agentic Symbolic Regression
- Authors: Zihan Yu , Shixuan Zhou , Hao Huang , Jingtao Ding , Yong Li
- URL: https://arxiv.org/abs/2609.35501
- Abstract:
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.
27. Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
- Authors: Yan Zhan , Shaobo Liu , Zhijun Gao
- URL: https://arxiv.org/abs/2609.35472
- Abstract:
Discrete diffusion language models (dLLMs) expose a denoised solution at every step, which makes process reward model (PRM) guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged in the same budget of forward passes, its deterministic form loses to a much simpler baseline. Our PRMs score intermediate denoising states and are trained on the correctness of the final answer. On Dream-v0-Instruct-7B with 8 candidates per GSM8K problem, keeping the candidate with the highest PRM score at every scoring step reaches 65.18%, while independent sampling plus an ORM reranker trained for the task reaches 75.13%. The gap grows to 12.69 percentage points (pp) with 32 candidates, and is 9.85 pp on MATH and 12.16 pp on MBPP. We trace it to two separable failures. First, guidance prunes on a weak signal: on GSM8K, PRM ROC-AUC falls from 0.77 to 0.54 as the mask ratio rises, a decay that persists when states are relabeled with fresh rollouts, and pruning lowers the best accuracy reachable from the candidate pool from 81.05% for independent samples to 67.30%. Second, on GSM8K and MATH, the PRM is a poor final judge: a sequential Monte Carlo sampler at the same budget restores that ceiling to 77.89%, yet selecting with the PRM gives 65.48%, on par with deterministic guidance, while a PRM retrained on final states matches the ORM on identical candidates. MBPP separates the two: there the PRM reaches 65.47% when reranking finished programs, on par with the ORM, but 50.88% when it guides denoising. The results point to two targets for dLLM guidance: keep correct partial solutions alive through early denoising, and leave the final choice to a verifier trained on final states. We release the corpus of denoising states with outcome labels and evaluation toolkit for reproducible comparisons at matched compute.
28. A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans
- Authors: Alessandro Emmanuel Pecora , Stefano Calzolari , Francesco Strada , Andrea Bottino
- URL: https://arxiv.org/abs/2609.35463
- Abstract:
Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify perception, reasoning, and action within a single control loop. A.D.A.M.O. maintains a dual visual-symbolic world model that combines egocentric visual input and synchronized symbolic state to support grounded task-oriented behavior from natural language prompts. To support diagnostic evaluation, we introduce a controlled task suite organized by a cd taxonomy that breaks down spatial tasks into procedural and linguistic complexity. Experiments in controlled scenes show that semantic labeling strongly influences task completion and failure modes, reducing perceptual ambiguity while shifting failures toward downstream execution, whereas reasoning errors remain comparatively rare.
29. AutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain–Computer Interfaces
- Authors: Muyun Jiang , Yi Ding , Wei Zhang , Jinbo Chen , Chenyu Liu , Zhenjie Yang , Yuxin Li , Jingyuan Chen , Yuhao Lu , Yong Li , Shuailei Zhang , Cuntai Guan
- URL: https://arxiv.org/abs/2609.35456
- Abstract:
EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection of EEG decoding architectures across tasks. The Designer Agent performs Pool-Guided Architecture Discovery (PGAD), generating and refining architectures through training and validation across multiple EEG tasks, such as emotion recognition, motor imagery, and sleep staging. The Forecaster Agent performs Performance Estimation from Early Knowledge (PEEK), using architecture code, the training protocol, and early learning curves to predict full-budget validation performance and select promising candidates for continued training. Across 14 EEG datasets spanning motor imagery, emotion recognition, and sleep staging, we evaluate AutoBCI with six LLMs, including Opus 5.5 and GPT 5.6 Sol, and compare the architectures selected by the search procedure against ten baselines: six conventional EEG models and four foundation models. The architecture discovered by AutoBCI with Claude Opus 5.5 achieves 64.16% average test balanced accuracy (bAcc), compared with 63.87% for REVE, the strongest baseline on this metric. Using ten observed epochs, PEEK reduces mean absolute error in predicting average validation bAcc from 2.20 to 1.36 percentage points, a 38.1% reduction relative to the best-observed-score baseline.
30. Just Initialize: A Training-Free Initialization Component for Large-Scale Routing Optimization
- Authors: Jiale Zhao , Sirui Mao , Zimu Chen , Wentao Yang , Zihan Wang , Xuefeng Huang , Junji Cheng , Liyuanjun Lai
- URL: https://arxiv.org/abs/2609.35443
- Abstract:
Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational cost. We instead shift the focus to a useful initialization that can be refined into a high-quality solution with limited downstream refinement. We propose Just Initialize, a training-free and solver-agnostic initialization component for large-scale routing optimization. Just Initialize compresses a large routing instance into a compact surrogate space, optimizes its global routing structure, and recovers the resulting solution as an optimization-friendly starting point in the original space. Extensive experiments on Traveling Salesman Problems (TSPs), Capacitated Vehicle Routing Problems (CVRPs), Vehicle Routing Problems with Time Windows (VRPTWs), and Prize-Collecting Traveling Salesman Problems (PCTSPs) demonstrate that Just Initialize achieves high-quality solutions comparable to or better than state-of-the-art methods while substantially reducing computational cost across instances ranging from 1K to 100K nodes, including an average speedup of approximately 70$\times$, sub-second runtimes on 10K-node instances, and runtimes within tens of seconds on 100K-node instances.
31. Building Transformation Layers for Riemannian Neural Networks
- Authors: Ziheng Chen
- URL: https://arxiv.org/abs/2609.35436
- Abstract:
Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a framework for constructing FC and convolutional layers over computationally tractable Riemannian spaces. This framework incorporates several previous FC layers across different geometries as special cases and is instantiated on ten representative manifolds, including three hyperbolic models, five geometries of the symmetric positive definite (SPD) manifold, and two Grassmannian perspectives. Experiments on different manifolds demonstrate the effectiveness and applicability of our approach. Code can be found at this https URL .
32. Self-Adapting Group of Experts for Multi-Agent Reasoning
- Authors: Mohammad Atif Quamar , Nurbek Tastan , Karthik Nandakumar , Junpei Komiyama
- URL: https://arxiv.org/abs/2609.35412
- Abstract:
Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent’s response depends on its model’s capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often adapt communication by changing this context while leaving individual prompts fixed, even when a problem calls for different skills. We study whether agents’ initial responses can identify a strategy better suited to the current problem and guide its transfer to other agents. To address this, we introduce SAGE (Self-Adapting Group of Experts), a training-free framework that uses answer agreement, prefix consistency, and reciprocal peer review to select a strategy donor. SAGE transfers the selected donor’s reasoning strategy to the other agents while preserving their original roles. This transfer uses only the agents’ original system prompts, without access to the problem or generated solutions. After strategy adaptation, agents exchange responses through a dynamic, sparse directed acyclic graph that routes information from higher-scoring agents to lower-scoring agents. Experiments across multiple agent backbones and reasoning benchmarks show that SAGE achieves higher average accuracy than the evaluated baselines. Our code is available at this https URL .
33. Structural Alignment for Reliable Industrial AI: Bridging Physical Reality, Data, Models, and Human Intent
- Authors: Lizhi Xiao , Sihong Wu , Victoria Xiao , Yiqiao Song , Chen Gu , Jianwei Ma , Xinming Wu , Aimé Fournier
- URL: https://arxiv.org/abs/2609.35400
- Abstract:
Artificial intelligence is increasingly deployed in critical industrial domains, including healthcare, energy grids, subsurface exploration, where failures can have severe consequences for human safety, system stability, and economic outcomes. Yet AI is still evaluated primarily through benchmark accuracy, a model-centric metric that fails to capture the structural complexity and risks of real-world deployment. We propose a framework that views industrial AI reliability as a problem of structural alignment across four interacting worlds: physical, representational, machine, and human cognitive. These worlds are connected through two interfaces: digitalization, linking physical reality to computational representations, and goal encoding, translating human cognition to the machine objectives. Together, they define the space of admissible solutions. We characterize the solution space through four attributes: existence, non-uniqueness, robustness, and interpretability and show how mismatches arise at interfaces and propagate across worlds to produce reliability failures. Applications to healthcare, energy grids, and subsurface exploration illustrate that although dominant failure modes differ across domains, for example, interpretability in healthcare, robustness in energy grids, and non-uniqueness in subsurface exploration, all originate from a shared structural mechanism. By shifting the focus from model-centric evaluation to system-level alignment, this framework offers a principled foundation for assessing and governing reliability in industrial AI systems.
34. A decision-support system applied to Law: Reasoning and explainability of the decision
- Authors: Jeremy Bouche-Pillon (IRIT, IRIT-MELODI, IRIT-ADRIA, IRIT-LILaC), Pascale Zarat{é} (IRIT, UT Capitole, IRIT-ADRIA), Yannick Chevalier , Nathalie Aussenac-Gilles (IRIT-MELODI, IRIT, CNRS)
- URL: https://arxiv.org/abs/2609.35370
- Abstract:
The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ‘‘Law Enforcement Directive (LED)’’, appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal representation of these regulations can be part of decision systems that support LEAs in processing data in compliance with the regulations. Although many new formalisms have emerged to represent legal norms and rules, few are provided with a reasoning mechanism. Furthermore, systems used in decision-making processes in critical contexts such as medical diagnoses or legal decisions cannot be fully automated, and the explainability of their results is essential to ensure user confidence in decisions. This explainability aspect, while crucial, is lacking in most modern approaches that rely on machine learning. This paper describes a framework to operate formal rules from regulations, by focusing on explainability of the decision. After describing the general architecture of the proposed decision support framework, the paper showcases how symbolic AI and the SPARQL query language can support legal reasoning. It then describes an algorithm to generate a justification for the reasoning results, and outlines the procedure to be followed when the reasoning does not lead to a satisfactory conclusion. We notably focus on a method based on decision trees to determine what additional information to request from the user.
35. Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
- Authors: Kajetan Dymkiewicz , Tim Farrelly , Adam Práda , Ishaan Panigrahi , Srishti Gureja , Helen Yannakoudakis , Robert Mullins , Victor Gillioz , Daniel Tan , Maxime Riché
- URL: https://arxiv.org/abs/2609.35356
- Abstract:
Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours co-occur in most training examples, so filtering out examples with undesired behaviour leaves only a small clean subset. We introduce stratified inoculation prompting (SIP). SIP leverages a small clean subset to demonstrate that desired behaviour should persist without the undesired one across different contexts. SIP oversamples these clean examples under diverse non-eliciting prompts while inoculating the rest. SIP substantially reduces expression of undesired behaviour while preserving more of the desired behaviour than IP. These gains persist even when we extend IP to oversample the same clean subset at the same rate as SIP. Moreover, SIP yields lower emergent misalignment rates in all harmful-advice setups we tested. SIP can be further extended to limit the undesired behaviour even under prompts that explicitly request it. We introduce backdoor dilution, which weakens expression under the inoculation prompt, and password-locked inoculation, which concentrates elicitation on a designated password. Taken together, our findings show that changing the training contexts for a small clean subset can significantly improve selective generalisation.
36. Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
- Authors: Lucas Biechy , Cédric Eichler , Adrien Boiret , Nicolas Anciaux
- URL: https://arxiv.org/abs/2609.35350
- Abstract:
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model’s effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U’s improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
37. Jev thinks “I don’t know’’, but doesn’t say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
- Authors: Riccardo Porcedda
- URL: https://arxiv.org/abs/2609.35342
- Abstract:
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev’s central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don’t know’’.
38. TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving
- Authors: Shengqin Wang , Jie Jin , Yu Cheng , Yihang Chen , Weilin Luo , Yuan Xie , Zhizhong Zhang
- URL: https://arxiv.org/abs/2609.35336
- Abstract:
Despite the promise of Large Language Models (LLMs) in computational chemistry, rigorous combinatorial chemistry problems remain difficult because they require quantitatively constrained molecular modification, candidate validation, and systematic revision after failed attempts. Existing tool-augmented chemical agents demonstrate useful planning and tool use, but they rarely provide a unified loop for property-driven molecular optimization and workflow-level composition. To bridge this gap, we propose Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving (TMCS), a step-by-step multi-agent framework that formalizes chemical problem solving as an interpretable, tool-augmented workflow. At the task level, specialized agents leverage external tools, few-shot trajectory memory, and structured reflection to iteratively refine solutions. At the workflow level, TMCS chains generation, understanding, editing, description, and optimization into a closed-loop pipeline. Evaluations across multiple chemical tasks demonstrate that TMCS consistently enhances chemical reasoning across both open- and closed-source base models, achieving state-of-the-art performance.
39. Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
- Authors: Zipei Yu , Yue-Jiao Gong , Zeyuan Ma , Yuncheng Jiang , Zhiguang Cao
- URL: https://arxiv.org/abs/2609.35328
- Abstract:
Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm’s bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While MetaBBO helps advance the performance lower bound of the resulted optimization system, it is currently handcrafted and customized case by case to adapt different optimization problems, which inevitably introduces inherent subjectivity and hence restricts the performance upper bound and usability in practice. In this paper, we address this issue by regarding MetaBBO’s design loop as coding task, where we could introduce openendedness into MetaBBO with recursive self-improvement capability of advanced coding agents. Specifically, we propose a dual-agent framework: i) a task agent continuously refines the codebase of a target MetaBBO approach through code evolution; ii) a hyper agent progressively modifies the task agent and itself to provide open-ended design behavior; iii) the evolved MetaBBO codebase is evaluated and all in-execution information is fed back to the agents for recursive self-referential improvement. As a result, given a naive MetaBBO template, our framework automates a design evolution and finds novel variants superior to up-to-date human-made MetaBBO baselines. Surprisingly, the experimental results also demonstrate that our framework supports fast adaption across different optimization domains. Solid interpretation analysis further reveals interesting design principles emerge in such open-ended process. This work serves as the first exploration on automating design of complex learning-assisted optimization algorithms.
40. Reliability Engineering for AI Systems: Challenges, Methods, and Directions
- Authors: Rong Pan , Yili Hong , Min Xie
- URL: https://arxiv.org/abs/2609.35316
- Abstract:
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.
41. Narrowing the Horizon: Quantifying Topic Saliency Shifts in Generative Monoculture
- Authors: Oriane Peter , Elena Simperl , Kate Devlin
- URL: https://arxiv.org/abs/2609.35302
- Abstract:
As Large Language Models (LLMs) become central to how we access and share information, they play an increasingly powerful role in shaping global knowledge. However, as these models evolve, their outputs risk converging into a \textit{generative monoculture}, where the diversity of perspectives they represent narrows over time. Studies at the model level often fail to pinpoint which specific topics or viewpoints are being marginalised or amplified in this process. In this paper, we introduce a method to measure shifts in topic saliency across model families, tracking what gains or loses prominence during post-training. Applying this approach to a case study of climate change discourse, we demonstrate how homogenisation affects the representation of diverse solutions across different models. We also test interventions to counter this trend, showing that specialised models can help preserve a broader range of perspectives. This underscores the importance of monitoring topic saliency to diagnose the risks of monoculture and to ensure AI systems reflect a pluralism of ideas. Data and Code are accessible \href{ this https URL }{here}.
42. Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework
- Authors: Surajit Das
- URL: https://arxiv.org/abs/2609.35298
- Abstract:
Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Information Gate, knowledge-weighted evidence accumulation, disease similarity, and decisive clinical rules. Missing-aware normalization and coverage auditing distinguish absent from unavailable evidence. Candidate ranking is separate from outcome-label-independent K-means clustering, which uses four derived evidence coordinates (evidence strength, relative magnitude, directional similarity, and evidence completeness), not raw predictors or targets, to derive cohort-level assignments. Across six retrospective cohorts - four dengue (N = 1000, 1523, 989, 1018), malaria (N = 2190), and influenza (N = 4569) - a uniform, label-free, cohort-fitted K = 2 protocol yielded positive-class F1 scores of 0.996, 0.634, 0.936, 0.917, 0.695, and 0.842, and all-record accuracies of 0.996, 0.558, 0.914, 0.893, 0.707, and 0.906, respectively, with full partition-decision coverage using the frozen package and disease-specific knowledge representations. Neither scoring nor clustering uses outcome labels. Logistic regression provides a supervised baseline. Influenza incorporates confirmatory molecular PCR and is not independent pre-test prediction. Results characterize knowledge-grounded evidence separation, auditability, and sensitivity, not prospective clinical validity or comparative superiority. FOL/LLM-based clinical explanation remains unevaluated.
43. AbGaze: Attentive Geometric Representation Learning for End-to-End Antibody Design
- Authors: Jiashuo Wang , Siqi Fan , Yizhen Luo , Zaiqing Nie
- URL: https://arxiv.org/abs/2609.35296
- Abstract:
Computational antibody design requires representations that capture the geometric patterns underlying antigen–antibody interactions, yet existing approaches often rely on scalar distances or surface-intrinsic features, leaving cross-molecular geometry largely implicit. We present AbGaze, an end-to-end antibody design framework based on attentive geometric representation learning, which encodes distance, spatial direction, and surface-normal orientation of antigen surfaces relative to antibody-residue local frames, and adaptively aggregates these geometric interactions according to their interfacial context. The learned interaction representation is shared across multi-CDR co-design, complex structure prediction, and affinity optimization, with local-frame geometric supervision further constraining the representation. AbGaze outperforms prior methods across all three tasks: relative to the second-best method, it improves amino-acid recovery by 7.1% and reduces structural error by 14.9% on average over the six CDRs, improves interface docking quality (DockQ) by 6.6%, and raises the affinity improvement rate (IMP) by 32.5%.
44. EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
- Authors: Shihan Dou , Shaofan Liu , Zhonghang Lu , Jiahang Lin , Shichun Liu , Binghai Wang , Jiajie Jin , Guanting Dong , Tao Gui , Qi Zhang , Xuanjing Huang
- URL: https://arxiv.org/abs/2609.35290
- Abstract:
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ‘‘potpourri’’ approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model’s own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document’s length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.
45. The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
- Authors: Michele Loi
- URL: https://arxiv.org/abs/2609.35286
- Abstract:
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany’s debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.
46. Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale
- Authors: Ghazal Fazelnia , Paul Gigioli , Eliza Klyce , Sharon Zheng , Katie Zelvin , Ye Myat Thein , Anurag Deshpande , Seda Davtyan , Kate Remeika , Maya Hristakeva , Erik Franco , Karen Banzon , Peng Ge , Jacqueline Wood , Nandini Singh , David Murgatroyd , Mounia Lalmas , Yves Raimond , Andreas Damianou
- URL: https://arxiv.org/abs/2609.35285
- Abstract:
Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embedding vectors remain highly effective for retrieval and ranking, but they are opaque to users and not natively expressed for language model workflows. We present Textual User Taste, a system that generates structured natural-language taste profiles from listening behavior, interaction signals, content metadata, and optional user feedback, and deploys them to millions of Spotify users. We describe the end-to-end production lifecycle required to generate, evaluate, optimize, and maintain these representations at industrial scale, including prompt development and compression, user steering, and integration with downstream personalization systems. Because no unique ground-truth taste profile exists, we introduce a multi-faceted evaluation framework to evaluate taste profiles as a production representation: they carry user-specific predictive signal independently, and when integrated with behavioral embeddings, improve MRR by 0.6% for future-track prediction and NDCG@7 by 2.2% for search ranking. Our evaluation also reveals that taste profiles support positive natural-language steering, while exposing important limitations, including challenges with negation and short-term temporal adaptation. These findings position taste profiles not as replacements for behavioral embeddings, but as an interpretable and steerable interface between evolving user context and foundation-model recommender systems.
47. Imprint Reader: From Weight-Update Readout to Behavioral Intervention
- Authors: Guanxu Chen , Qihao Lin , Jing Shao
- URL: https://arxiv.org/abs/2609.35261
- Abstract:
As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter-level traces that can, in principle, be inspected directly. However, current models cannot decode these traces into an explicit account of what they have learned. To this end, we introduce the \textit{Imprint Reader}, a model trained with \textit{Semantic Mount-and-Read Tuning} (SaRT) to describe frozen weight updates. SMaRT mounts each update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, while no-change and random-perturbation controls discourage unsupported claims. On held-out updates, the joint Reader reaches judge-based Pass@100 of $2\%$ for knowledge and $16\%$ for behavior. These results demonstrate the feasibility of natural-language readout while pointing to reliability across updates as the next step. Beyond free-form generation, the Reader provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update. Its coordinate-aligned gradients support intervention through MetaEdit. At a $0.5\%$ pruning rate, Reader-guided selection raises measured harmful-prompt refusal from $57.9\%$ to $64.1\%$ under a safety-maintenance target. Using behavior descriptions without target-task training data, MetaEdit increases the frequency of backtracking and sub-goal expressions in mathematical reasoning traces and raises BFCL Overall from $41.69\%$ to $44.60\%$.
48. Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
- Authors: Huachi Zhou , Yujing Zhang , Jiahe Du , Jiacheng Cai , Zijin Hong , Chuang Zhou , Zheng Yuan , Qinggang Zhang , Qing Li , Xiao Huang
- URL: https://arxiv.org/abs/2609.35255
- Abstract:
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at this https URL .
49. EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents
- Authors: Fengzhou Sun , Yuan Zhang , Xintong Yu , Jinyao Yan
- URL: https://arxiv.org/abs/2609.35233
- Abstract:
Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users’ social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus on instantaneous interactions, leaving the long-term relational disclosure problem unexplored. In this paper, we propose EP-Mem, an Elastic Privacy Memory architecture that reframes privacy as user-owned boundary control across social roles. EP-Mem introduces (1) token-level memory driven by user-configurable a privacy policy that stratifies persons and events, combining domain-level default circulation rules with fact-level whitelist/blacklist exceptions; and (2) a pluggable sidecar with a privacy engine that aligns disclosure controls with memory across summary, detail, and boundary granularities, enforced throughout generation, storage, and retrieval. We construct EP-Bench, to our knowledge the first long-term multi-party benchmark with cross-session correlated events for policy-conditioned relational disclosure. Experiments show that EP-Mem achieves 94.0% privacy classification accuracy, improves disclosure-permission judgment from 22% to 68%, and reduces privacy leakage by 75.6%, while maintaining retrieval performance and cross-benchmark generalization.
50. ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
- Authors: Yang Li , Jinhan Yang , hai liu , Di Wan , Xiyu Chen , Zongsi Xu , Tuo Zhou , Sheng Zhong , Sergey Volkov , Ye Luo , Hao Sun
- URL: https://arxiv.org/abs/2609.35215
- Abstract:
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor’s probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.
51. Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
- Authors: Suwesh Prasad Sah
- URL: https://arxiv.org/abs/2609.35188
- Abstract:
Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by (1.91\times) to (2.19\times) across all prompts and reduced time to first output by 10.0–14.2\%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4–78.1\% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.
52. 5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding
- Authors: Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University)
- URL: https://arxiv.org/abs/2609.35184
- Abstract:
Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later tasks may be unable to use information that was omitted. If the index retains only open-ended text, however, rule-based reasoning lacks checkable premises. We propose 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding. The 5W1H questions organize source-grounded content units; Which points to versioned ontology elements and records mapping relations, scope, and validation status. Time, location, system environment, and participant roles are not merely retrieval labels: together, they constrain the contexts in which facts, bindings, and rules apply. Unbound content remains searchable, while bound content enters a formal reasoning path only after premise checks. The method further distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions. A worked example of migration from an on-premises server to a cloud environment illustrates the different treatment of world-state changes, ontology-version changes, and changes in rule applicability. We formulate three groups of falsifiable hypotheses concerning cross-task evidence coverage, control of contextual misuse, and incremental update cost. The planned evaluation includes a strong typed fact-graph baseline with the same evidence, temporal information, and budget, to test whether benefits arise from 5W1H organization, deferred binding, or additional information and engineering effort. The contribution is a testable indexing mechanism, not a claim to a new universal ontology or a demonstrated performance advantage.
53. FONDANT: Strong and Best-Effort Planning via Antichains
- Authors: Benjamin Aminof , Tuan Khai Nguyen , Sasha Rubin
- URL: https://arxiv.org/abs/2609.35160
- Abstract:
A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available or there is no evidence that the environment is adversarial, one can resort to best-effort policies, which always exist, and which follow the classic decision-theoretic principle that an agent should not use a dominated strategy. A typical positional best-effort policy works as follows: from every state, it follows a strong policy if one exists from that state (such states are called
strong-winning''), else a weak policy if one exists from that state (weak-winning’’), and else is unconstrained (``losing’’). In this work, we introduce a sound and complete planner for both best-effort planning and strong planning. The algorithm that underpins the planner is quite simple: it represents certain sets of states, such as the winning regions, by their $\subseteq$-minimal elements. The algorithm returns uniform policies, i.e., it returns a policy $\pi_t$ that is a strong solution starting in every strong-winning state, and it returns a policy $\pi_w$ that is a weak solution starting in every weak-winning state, and it provides a certificate for the set of losing states. We implemented the algorithm with some simple optimizations (calling it FONDANT), and evaluated it on a benchmark set consisting of the instances that were used in the evaluation of leading strong planners PR2 and FOND-SAT, and the best-effort planner BeSyftP. On coverage, our implementation is at least as good on all domains, and outperforms on some domains; and on wall time, it is slower on small and medium-sized instances, and outperforms on larger instances.
54. PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
- Authors: Jiaan Zhu , Wei Gao , Youhui Bai , Zewen Jin , Ju Huang , Siran Yang , Jiamang Wang , Lin Qu , Cheng Li
- URL: https://arxiv.org/abs/2609.35158
- Abstract:
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill–decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU–worker–role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$–$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.
55. From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale
- Authors: Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University)
- URL: https://arxiv.org/abs/2609.35149
- Abstract:
Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retention or target-contract satisfaction. We formulate agent calibration as constrained behavioral adaptation across three interacting layers: information preservation, harness adaptation, and user acceptance; the layers apply to every scenario, not one-to-one to the three. The basic objective is non-degradation on prespecified capability measures while satisfying target requirements; aggregate improvement is stronger. Information calibration preserves independently validated source content still applicable to the target task. Harness calibration aligns observable artifacts at semantic checkpoints and repairs them through iteration, tool substitution, or local replanning within explicit budgets. User calibration enforces recipient-specific output contracts: templates, schemas, and section-level preferences. A global e-commerce example shows how shared standards coexist with site- and market-specific adapters and validation. We distinguish trainable policies from frozen-backbone configuration or controller optimization, and evidence verification from relative judgment and DPO/GRPO optimization. Recent harness-transfer and judge-validity studies motivate target-native execution records, separate audits of task validity and near-tie ranking, and matched target-native optimization controls. We propose held-out evaluations for model changes, cross-border adaptation, and scale, including a factorial test of source evidence and checkpoint repair and group-level reporting to prevent aggregate gains from masking local failures. This is a methodological proposal; implementation and empirical validation remain future work.
56. Tool Mediation Alters Refusal Mechanisms in Large Language Models
- Authors: Abel Rodríguez , Giuseppe Garofalo , Lieven Desmet , Vera Rimmer
- URL: https://arxiv.org/abs/2609.35117
- Abstract:
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms across a diverse set of open-weight language models. We find that information about the harmfulness of a request remains strongly encoded in the model’s representations and transfers across conversational and tool-mediated inputs. Evidence from representation geometry and neuron-level analysis further indicates that the two interaction modes systematically distribute harm-related computation differently. Crucially, while conversational inputs can be refused at relatively low levels of perceived harmfulness, tool-mediated inputs remain permissive until harmfulness crosses a substantially higher effective refusal threshold. Moreover, tool-mediated refusal is also more brittle: progressively weakening the refusal computation disrupts tool-mediated refusal at lower intervention strengths than conversational refusal, even when benign capabilities remain intact. Together, our findings indicate that tool mediation does not simply reduce the internal perception of harm, but instead impacts its conversion into refusal. Overall, this suggests tool-mediated environments may intrinsically reduce robustness of models to harmful requests, and that conventional safety evaluations may not fully transfer to LLM agents.
57. DuplexCadence: Exact State and Execution from a Speech Model’s Declared Timelines
- Authors: Haixiao Gao , Yimin Zheng , Linyou Xiao , Zeke Xie
- URL: https://arxiv.org/abs/2609.35115
- Abstract:
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model’s native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime’s speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-
58. Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
- Authors: Yitong Li , Jincheng Yu , Junsong Chen , Haopeng Li , Shuchen Xue , Haozhe Liu , Ping Luo , Song Han , Enze Xie
- URL: https://arxiv.org/abs/2609.35110
- Abstract:
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.
59. Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling
- Authors: Bohao Wang , Xiaoyan Zhao , Yang Zhang , Jinghang Guo , Chun Chen , Can Wang , Jiawei Chen
- URL: https://arxiv.org/abs/2609.35109
- Abstract:
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user’s historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.
60. DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery
- Authors: Yulong Li , Rong Xia , Yuxuan Zhang , Jianxu Chen , Xiwei Liu , Haochen Xue , Maosheng Li , Yuhang Liu , Yibo Yuan , Yutong Xie , Chong Li , Jionglong Su , Hagai Rossman , Eran Segal , Imran Razzak
- URL: https://arxiv.org/abs/2609.35107
- Abstract:
We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, from longitudinal clinical phenotypes, medical imaging, and continuous physiological signals to eight molecular layers, together with an evidence network of approximately 4.7 million literature-derived records over 93,566 concepts and 149,383 candidate causal relations. DoAtlas-2 autonomously formulates research questions from evidence gaps and unresolved mechanisms, prespecifies their causal designs, and generates validated analyses. Supporting, challenging, and unresolved results continuously revise mechanistic interpretations, the causal evidence state, and the discovery frontier, so that DoAtlas-2 self-evolves within a closed loop of hypothesis generation, empirical testing, and renewed discovery. DoAtlas-2 has systematically evaluated 2,031 research questions. In the Human Phenotype Project (HPP), it formulated 4,014 candidate pathway questions across vascular, early-glycemic, and hepatic-metabolic systems, and screening of the first 1,079 yielded statistical support for 756. Representative studies identify blood pressure as a convergence node linking adiposity, hepatic, and lipid phenotypes to vascular outcomes, and show that an adiposity-inflammation-blood-pressure pathway is largely attenuated by joint adjustment for body mass index (BMI) and smoking. The discovered vascular network constitutes a completely interpretable predictive foundation, admitting exact attribution of every prediction and closed-form mediation effects. DoAtlas-2 thereby unifies causal mechanism discovery, population-evidence testing, and interpretable prediction within one continuously evolving foundation.
61. Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
- Authors: Zehao Lu , Xingguo Xiong , Wopke van der Werf , Thijs L. van der Plas , Ioannis N. Athanasiadis
- URL: https://arxiv.org/abs/2609.35089
- Abstract:
Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches—direct zero-shot prompting, a staged workflow, and a multi-agent system—with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model–approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.
62. When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents
- Authors: Geonwoo Kim (1), Brent ByungHoon Kang (1) ((1) Korea Advanced Institute of Science and Technology (KAIST))
- URL: https://arxiv.org/abs/2609.35088
- Abstract:
Tool-enabled agents form calls from model-visible interfaces, while hosts later select their implementation. Standard dispatch omits the descriptor-handler relation. An unchanged and schema-valid call can therefore acquire a different security effect during rollout, reconnect, or delayed approval. We call this failure schema-epoch drift. We present formation-consistent dispatch (FCD), which connects implementation analysis to execution authority. Reviewed profiles produce provenance-bound over-approximations of declared in-scope effects from official source. Under a closed-target approval policy, a verifier applies each formed call to a summary and captures a successor only when its effects fit the call’s security contract. Atomic admission and a final-hop fence preserve this decision to the effect. The exact source retains priority, and the captured successor becomes eligible only after source retirement. Stock releases and deployment changes reproduced the failure. Four profiles covered 32 official releases: 29 required no release-specific change and three escalated. A frozen 16-release expansion matched a separate source oracle. In a preregistered stock comparison, FCD completed all three pending calls whose effect remained private and blocked all three whose omission became public. Exact pinning and release-wide denial stopped all six calls, while release-wide approval completed all six but produced three public effects. A separate lifecycle experiment carried a formation-captured certificate across source retirement. The same safe certificate installed later governed new formations without expanding the pending call’s authority.
63. What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages
- Authors: Ben Moore , Liam Dunne
- URL: https://arxiv.org/abs/2609.35077
- Abstract:
Production large language models retrieve and cite web pages alongside generated answers, yet the page-level features that predict citation frequency remain poorly characterised. We present an observational study of approximately 2 million LLM citations from four commercial engines (ChatGPT, Claude, Google AI, Gemini) over six months, joined to 10,000 crawled pages from nineteen B2B SaaS workspaces. Sixty-plus features are tested using a nine-method consensus framework combining mixed-effects regression with domain fixed effects, FDR correction, stability-selection Lasso, double machine learning, generalised additive models, and temporal hold-out replication. Four findings survive all checks. First, prompt-content alignment (Jaccard overlap between page tokens and the full workspace prompt corpus, including non-citing prompts) is the dominant page-level predictor (beta = +0.37, 95% CI [+0.33, +0.41], q ~ 10^-73). Second, the standard AEO checklist (FAQ blocks, structured data, Core Web Vitals) shows positive effects in pooled data that reverse or collapse to zero once domain fixed effects are applied: Simpson’s paradox with practical consequences for the AEO literature. Third, domain-level AI authority exceeds the strongest non-alignment page-level feature by a factor of six in mean absolute SHAP value. We release the analytic pipeline as a methodological contribution.
64. Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation
- Authors: Jiashen Ren , Wenlin Zhang , Bohan Zhang , Xiaopeng Li , Zichuan Fu , Wanyu Wang , Junyi Li , Xiangyu Zhao
- URL: https://arxiv.org/abs/2609.35036
- Abstract:
Persona prompting is widely used to construct user simulations with large language models (LLMs), yet it relies on a largely untested assumption: specifying one user attribute should change that attribute alone. We test this assumption and identify a systematic failure of selective control: across all eight black-box LLMs we audit, changing a target attribute also shifts responses on unspecified, non-target attributes. For example, describing a user as more risk-seeking shifts color choices, even though the prompt never mentions color; we term this cross-attribute influence. Semantic, contextual, and internal analyses collectively suggest that models treat a persona prompt as evidence about the user and extend the inferred profile to unspecified preferences, a process we call trait-conditioned completion. We next ask whether explicitly specifying non-target attributes restores selective control. When a non-target attribute is assigned a clear direction, models generally follow the declaration and suppress the target attribute’s influence. However, when the same attribute is declared neutral, the target continues to affect choices across all five open-weight checkpoints, even when the model correctly reports the declared state. This disparity, the neutrality gap, demonstrates that successful persona following does not imply selective persona control, which additionally requires keeping non-target attributes stable. We operationalize this distinction with a three-state diagnostic that leaves the non-target attribute unspecified or declares it directional or neutral; because directional tests can be passed by simply following the stated persona, the neutral state reveals failures they miss. In a post hoc analysis of independent items, neutral declarations leave 51-81% of items target-sensitive, against at most 1 of 320 item-pole comparisons under directional ones.
65. JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
- Authors: Zhixi Cai , Fucai Ke , Sukai Huang , Maria Garcia de la Banda , Peter J. Stuckey , Gholamreza Haffari , Hamid Rezatofighi
- URL: https://arxiv.org/abs/2609.35032
- Abstract:
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at this https URL .
66. WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
- Authors: Anton Emelyanov , Maria Tikhonova , Zaven Martirosian , Sergei Averkiev , Alena Fenogenova
- URL: https://arxiv.org/abs/2609.35026
- Abstract:
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface’s own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with parameters as a user or an agent acts. A task declares the events it requires, and success is decided by matching them, with no judge model and no scraping of rendered pages. The same instrumentation supports controlled UI variation: one configuration switch re-renders a task through a different implementation of a single control while the prompt and the success conditions stay completely identical, so sensitivity to interface form can be measured under a fixed task specification. The WebPageBench release consists of three components: 152 tasks, divided into 65 canonical scenarios and 87 control variants across light/dark UI-modes; a common runner evaluated with six browser/DOM harness configurations and five screenshot-only GUI-agent families; and a public leaderboard of 24 model-harness pairs. On the public 152-task leaderboard the gap between what agents declare finished and what the log confirms reaches 41 points (one configuration declares every task finished and satisfies the conditions on 59%).
67. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
- Authors: Haotian Luo , Haoyu Wang , Zeyu Qin , Huanjin Yao , Yibo Wang , Zhuotao Tian , Shuai Wang , Jiaya Jia
- URL: https://arxiv.org/abs/2609.35025
- Abstract:
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into this http URL production line still rests on human labour and on human-in-the-loop collaboration. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent’s ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where data is delivered sample by sample and each sample is accepted against a set of criteria rather than put straight into training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task for the same suite that meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. AutoDataBench provides a direct measure of an agent’s capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run. Code and data are available at this https URL .
68. Environmental requirements for the use of social information by artificial life agents using evolved plastic artificial neural networks
- Authors: Hugh Charterton , James M. Borg , Aniko Ekart
- URL: https://arxiv.org/abs/2609.35018
- Abstract:
Evolved Plastic Artificial Neural Networks (EPANNs) consist of two principal processes, the first, evolution, and the second, development and in-life learning. In the context of the origins of social = learning, very few studies have been carried out using ALIFE models based on EPANN requirements. Studies in this field have usually involved an imitative teacher/pupil relationship. This, however, ignores the possibility that the observed behaviour is a consequence of social information cues rather than direct imitation or teaching. Starting with the first of the EPANN processes (evolution), a series of experiments was undertaken using artificial neural network (ANN) based agents in a variety of foraging environments to examine under what minimal environmental conditions the use of social information might have evolved, as measured by the number of generations taken to meet a specified fitness criterion. NEAT (Neuroevolution of Augmenting Topologies) was the ANN used as its evolutionary algorithm would evolve a network’s topology as well its weights. Unintentionally, in the experiment there was a simple network topology based on the location of the nearest food item which enabled agents to swiftly meet the fitness criterion. With this topology, additional information, social or otherwise, was not required and could have proved to be a hindrance. However, this does indicate that for the use of social information to have evolved, it would require a greater degree of complexity in the environment to do so.
69. TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation
- Authors: Nicolas Dahan (ISIR, ALMAnaCH), Fran{\cc}ois Yvon (MLIA, ISIR), Rachel Bawden (ALMAnaCH)
- URL: https://arxiv.org/abs/2609.35017
- Abstract:
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assigns an interpretable verdict to every term occurrence: glossary-conforming occurrences are settled deterministically, while divergences are assessed under a two-step LLM-as-judge procedure using the full document context: the first detects and labels terminology errors; the second sorts valid document-level variations from inconsistencies. Validated against expert error annotations and document-level human MQM scores, TermJudge ranks first in both system- and segment-level meta-evaluation, ahead of glossary-conformity and quality-estimation baselines. When applied to eight systems translating academic documents, under two prompting conditions, we observe that glossary injection improves terminology translation in all paired comparisons, by removing genuine errors rather than valid variation. TermJudge is released as open-source code.
70. Automated feature engineering, AutoML, and decision-focused learning for improved energy consumption forecasting
- Authors: Nasser Alkhulaifi
- URL: https://arxiv.org/abs/2609.35013
- Abstract:
The rising cost and demand for energy, together with environmental sustainability goals, create major challenges for energy management. Energy Consumption Forecasting (ECF) supports planning by predicting future consumption, but Machine Learning (ML) models for ECF often depend on expert-driven Feature Engineering (FE). This thesis addresses that dependence through three contributions. First, it establishes and evaluates a comprehensive FE pipeline for ECF and investigates domain-specific features. Second, it introduces AutoEnergy, a domain-tailored automated FE algorithm that generates interpretable features from timestamps and lagged consumption and integrates with AutoML for end-to-end ECF modelling. Across eighteen real-world energy datasets spanning residential, commercial, industrial, renewable, and grid domains, AutoEnergy reduces forecasting error by 19.52%-84.72% relative to baseline AutoML and established automated FE methods, while running 1.31-4.41 times faster, with gains varying by dataset. Third, AutoEnergy is integrated with Decision-Focused Learning (DFL) for a Battery Energy Storage System problem, jointly forecasting electricity prices and demand while optimising charging and discharging decisions. On a real-world UK property dataset, this approach reduces operating costs by 22.9%-56.5% compared with the same DFL models without automated FE. Overall, the results show that domain-specific automated FE can reduce reliance on manual feature design, improve forecasting accuracy, and translate predictive gains into measurable operational benefits in energy management.
71. From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing
- Authors: André Ricardo Ducca Fernandes , Jean-Pierre Briot , Simone Diniz Junqueira Barbosa1 , Hélio Côrtes Vieira Lopes
- URL: https://arxiv.org/abs/2609.34994
- Abstract:
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations – chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.
72. Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces
- Authors: Ruozhao Yang , Mingfei Cheng , Xiaofei Xie
- URL: https://arxiv.org/abs/2609.34974
- Abstract:
LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users’ interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArena-Single and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.
73. APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
- Authors: Puneet Mathur , Dinesh Manocha
- URL: https://arxiv.org/abs/2609.34973
- Abstract:
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.
74. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias
- Authors: Yinhong Liu , Zhili Tan , Zilin Wang , Zhijiang Guo
- URL: https://arxiv.org/abs/2609.34971
- Abstract:
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.
75. Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks
- Authors: Hangzun Liu , Yuling Fan , Fang Tian , Zhilong Bie , Zaiwen Feng , Yongliang Qiao
- URL: https://arxiv.org/abs/2609.34966
- Abstract:
Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.
76. ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization
- Authors: Feiming Wang , Daibo Li , Kun Yuan
- URL: https://arxiv.org/abs/2609.34960
- Abstract:
Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system for Proof-Obligation-Driven Theory Construction. Given a published algorithm, target theorem, and source proof, ProofLoom autonomously constructs the Lean model and supporting theory. Open proof obligations drive the development of definitions, interfaces, lemmas, and proof plans. Signature contracts record evidence and obligations for model revisions; an independent Judge rejects unsupported assumptions and weakened conclusions. Planner expands the published argument into intermediate claims, and Audit checks whether the Lean proof follows it. Across tasks, SOptLib accumulates verified mathematics and construction experience: reusable results are extracted, generalized, and verified, while modeling decisions and failed proof routes are recorded. Later tasks retrieve these results and records and contribute new developments, forming a cycle of construction, accumulation, and reuse. On fifteen textbook and research-paper tasks, ProofLoom obtains mean human ratings of 6.3/7 and 6.4/7, compared with 4.9/7 and 5.0/7 for the strongest of six baselines. Across 33 developments, it produces 490,693 lines of algorithm-local Lean code with no sorry. The formalizations also expose 28 incorrect formulas, proof gaps, and algorithm-analysis mismatches in published sources across 22 developments, each with checked evidence.
77. AX is the New AEO
- Authors: Ido Finder , Assaf Elovic , Gad Shalev
- URL: https://arxiv.org/abs/2609.34951
- Abstract:
In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models’ training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter breadcrumbs across forum threads, listicles, and off-site citations, so AI engines are likelier to surface and recommend them. But being surfaced is no longer enough: an agent opens the results and reads them before deciding, and one buyer question sends it through several rounds of search and fetch. What decides the outcome at this drill-down step is whether the agent can fetch and read the business’s own site: agent experience (AX). We argue that AX is the new AEO. We run 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses, matched on fame, prior model knowledge, and two AEO proxies, then split based on their AX level. Only 7-10% of the finished answer comes from the model’s training knowledge, whether or not the site is readable. Agent-ready businesses have answers built from their own pages 78% of the time against 56% and are clearly recommended 1.9x more often, while every grounded answer about a not-agent-ready business costs the agent 64% more. Holding business, harness, and question fixed, answers built from the site are 41% more accurate. The dominant failure is not fabrication but omission: web-built answers are 3.7x more likely to contain none of the facts the buyer asked for. Baselines differ sharply across the four harnesses, with clear-recommendation rates varying sevenfold from stack to stack, yet the effect holds in every one. In the agentic web era, being readable beats being talked about, and improving a site’s AX is the strongest lever a business has.
78. VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
- Authors: Mengyang Zhao , Zhuolin He , Haiyang Yu , Yuxuan Liang , Yifang Xu , Yuchuan Wu , Xiaolei Chen , Zhengtao Yao , Fan Shi , Yang Liu , Bin Li , Xiangyang Xue
- URL: https://arxiv.org/abs/2609.34949
- Abstract:
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.
79. Proactive Dialogue Policy Optimization via Cognitive-State Transition
- Authors: Minghui Ma , Mengqi Chen , Bin Guo , Jingqi Liu
- URL: https://arxiv.org/abs/2609.34948
- Abstract:
Proactive dialogue requires agents to continually adapt their policies to user feedback while progressing toward task objectives over multiple turns. To move beyond imitation learning on static datasets, recent approaches use user simulators to collect interactive data for policy optimization. However, many simulators do not explicitly model the evolution of user cognition, limiting the consistency and state dependence of feedback across turns. Moreover, representing each action only by a high-level strategy label overlooks the large utterance space and cannot distinguish alternative realizations of the same strategy. To this end, we jointly design a $\textbf{Cog}$nitive User $\textbf{Sim}$ulator $\textbf{(Cog-Sim)}$ and $\textbf{C}$ognitive-$\textbf{S}$tate $\textbf{T}$ransition–Driven $\textbf{P}$olicy $\textbf{O}$ptimization $\textbf{(CSTPO)}$. Cog-Sim maintains the user’s cognitive and affective states and generates responses through constrained state transitions across turns, so feedback depends on both the realized utterance and the user’s current state. CSTPO organizes each action as a hierarchical strategy–utterance representation: a high-level strategy label constrains utterance sampling, and utterances are optimized within each label. Sparse complete-branch sampling reuses shared dialogue prefixes and estimates separate strategy-level and utterance-level advantages, enabling fine-grained optimization at both levels. Across three tasks, Cog-Sim exhibits monotonic dose–response relationships and is preferred over prompt-based simulators for naturalness. CSTPO improves Qwen3-14B’s performance to a level comparable to that of GPT-5.5-based planning methods.
80. PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents
- Authors: Huayi Lai , Shichao Song , Qingchen Yu , Simin Niu , Mengwei Wang , Hanyu Wang , Xun Liang
- URL: https://arxiv.org/abs/2609.34930
- Abstract:
Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive execution, leaving unclear whether agents can formulate, execute, and revise an explicit plan while preserving user preferences throughout long-term interaction. To address this gap, we introduce \textbf{PDEU-Bench} (\textbf{P}ersonalized plan \textbf{D}efinition, plan \textbf{E}xecution, and plan \textbf{U}pdate \textbf{Bench}mark), a benchmark for evaluating the complete planning lifecycle of personalized tool-using agents. PDEU-Bench comprises 214 long-horizon interaction tasks spanning 12 everyday domains and 94 tools, with stage-specific assessments of preference adherence and plan quality. Extensive evaluations of 15 representative open-source and closed-source LLMs reveal a pronounced gap between local tool execution and dynamic planning: LLMs can often instantiate preferences in individual calls, yet struggle to construct coherent plan definition and plan update. We further evaluate mainstream personalization and memory-augmentation methods. Although these methods improve particular stages, none of the evaluated methods reliably propagates user preferences throughout the complete lifecycle, and their gains frequently fail to transfer to subsequent execution. Fine-grained error analysis further reveals that preference omissions and conflicts persist throughout the planning lifecycle, highlighting the need for future research to parameterize LLMs with preference-aware information retrieval and memory capabilities. We provide the relevant code and data in the appendix to support future research.
81. RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models
- Authors: Dmitrii Kharlapenko , Sergei Bratchikov , Konstantin Korolev , Aleksandr Nikolich
- URL: https://arxiv.org/abs/2609.34920
- Abstract:
On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.
82. DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
- Authors: Jeremy Canale
- URL: https://arxiv.org/abs/2609.34913
- Abstract:
Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for. Dossiers are generated from canonical facts under 61 executable rules, with 42 authoritative records and 32 narrative documents; every gate is certified decidable from those records. Attacks never change an authoritative value, so an attacked dossier keeps the reference decisions of its clean copy. A success is attributable only when the agent receives the injection and takes the exact injected action, which it does not take on the paired clean dossier; the DGF score is the share of applicable fixed attacks a model blocks. Reading documents and records themselves, five of six models were outcome-strict (disposition, findings, actions and authorization all correct) on 82 to 85 of 85 gates. Over 2,622 attacked gate runs, seven direct-order, false-data and false-authority attacks obtained one attributable success against these five, whereas task-aligned attacks imitating the organization’s own process passed against four of them: a record note citing a fake review procedure lowered GPT-6 Luna Pro from 34 to 6 outcome-strict gates and DeepSeek V4 Pro from 33 to 7. DGF scores ranged from 96.2 to 26.9, and a policy-aware adaptive attacker writing in records succeeded against five of six models. The approval tool executed no forged approval, yet deceived agents submitted approvals that the rules forbid. The open-source package dgf-bench computes the DGF score with one command.
83. Dual-Stream Simultaneous Translation via 2D Grid Attention
- Authors: Yu Pu , Wei-Qiang Zhang
- URL: https://arxiv.org/abs/2609.34902
- Abstract:
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations—broadcast and Hadamard—reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.
84. DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models
- Authors: Qirui Liu , Yichen Sun , Yan Wang , Zhixuan Chu , Linbo Jiang , Jianan Lin , Kui Ren
- URL: https://arxiv.org/abs/2609.34896
- Abstract:
Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.
85. Fewer Assumptions by Design: A Reusable Skill for LLM-Assisted Verus Verification
- Authors: Andrada-Livia Antoneac (Alexandru Ioan Cuza University of Iaşi, Bitdefender), Dorel Lucanu (Alexandru Ioan Cuza University of Iaşi), Dragoş Teodor Gavriluţ (Alexandru Ioan Cuza University of Iaşi, Bitdefender)
- URL: https://arxiv.org/abs/2609.34886
- Abstract:
LLM-assisted Verus verification is a less tedious method to verify Rust implementations, but paired with self-referential structures, e.g., Doubly Linked Lists (DLLs)—notoriously difficult to formalise for verification—it becomes a substantially more demanding verification task. Moreover, a specification weakness can arise when verification relies on unproven or invalidated assumptions, such as axiomatic lemmas and assume statements. We investigate whether LLM agents can synthesize strong DLL specifications while minimizing these trusted base. The analysis follows three different approaches: manual verification, property-specific verification, and a defined skill for the specific case of DLLs and certain properties of this type of data structure. The skill encodes domain knowledge and a task-decomposition strategy. We show that an LLM agent equipped with a carefully designed verification skill can generate strong, low-trust specifications for DLLs in Verus.
86. One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair
- Authors: Xiang Xia , Cheng Yan , Fan Xu , Zhijun Fan , Shuyuan Zhang , Wuyang Zhang
- URL: https://arxiv.org/abs/2609.34879
- Abstract:
Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both operation choices and their concrete realizations, making complete-sequence regeneration costly. Moreover, regeneration repeats operation selection even when failure arises from how those operations are realized. The resulting challenge is to reduce this repetition while preserving exploration of alternative operations and realizations. Therefore, we formulate repair as hierarchical search over operation supports, which we introduce as sets of permitted operation types that define reusable search regions for concrete tool-call sequences. We propose ReCommit, a training-free, diffusion-guided framework for improving tool-agent failure recovery while reducing repair computation. ReCommit amortizes operation-level proposal computation across repair trials by reusing operation-type scores from a single parallel readout of a masked diffusion language model. These scores guide search across supports, while realization search explores alternative entity bindings, arguments, and action composition within each support. Experiments on real failures across four enterprise services in the Agent-Diff benchmark show 75.9\% and 63.2\% relative recovery gains with 61.3\% and 51.3\% reductions in mean full-budget repair time at repair budgets $B=3$ and $B=13$, respectively, over the strongest evaluated 8B comparison method. ReCommit achieves a favorable recovery–cost trade-off, including in comparisons with the evaluated 32B models.
87. On the Limits of Metacognitive Monitoring in LLMs
- Authors: Dongqi Han , Yifan Yang , Dongsheng Li
- URL: https://arxiv.org/abs/2609.34864
- Abstract:
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.
88. AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces
- Authors: Zhijie Deng , Ling Li , Junhao Ji , Siwei Lyu , Zhipeng Xu , Zulong Chen , Rongyao Fang , Shuai Bai , Xuming Hu , Jiaheng Wei
- URL: https://arxiv.org/abs/2609.34854
- Abstract:
Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.
89. From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
- Authors: Jiangtao Lin , Bangyang Wei , Siyi Liu , Yihang Ding , Yuhan Dong
- URL: https://arxiv.org/abs/2609.34850
- Abstract:
Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewards produced by different objectives? We introduce assignment geometry to study this interaction. Mean-matched smoothing controls target dispersion, while within-stratum reassignment changes correspondence and preserves the complete target distribution. Across five reward objectives, intact correspondence retains the largest clean preference margins among the compared soft targets within a common accuracy-equivalence budget. Attenuation orderings change with the reward objective, revealing different responses to the same target assignments. Independent reassignments and a related source construction reproduce the retention direction. An attenuation-retention profile compares these combinations through margin magnitude, edit response, and accuracy. Against independently calibrated scaling, APLOT uniform targets deliver additional attenuation on both aggregate and presentation edits. These findings establish a joint design space in which target placement and reward objective shape reward properties beyond preference accuracy.
90. Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
- Authors: Yugu Li , Zehong Cao , Peizhen Li , Yang Zhang , Siyi Hu , Jianglin Qiao
- URL: https://arxiv.org/abs/2609.34848
- Abstract:
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5$\times$ reduction in token credit magnitude.
91. Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
- Authors: Wolfgang Maass
- URL: https://arxiv.org/abs/2609.34840
- Abstract:
An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-\chi)$ of the per-body value with $\chi$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.
92. BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
- Authors: Suyoung Kim , Jahyun Koo , Hyeonjin Kim , Inhyeok Bang , Seunghyun Lee , Hyunjae Oh , Baeseong Park , Dongsoo Lee
- URL: https://arxiv.org/abs/2609.34832
- Abstract:
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0–21.0\% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.
93. Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models
- Authors: Ji Huang , Mengfei Li , Shuai Shao
- URL: https://arxiv.org/abs/2609.34828
- Abstract:
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item’s response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
94. UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning
- Authors: Zenghuang Fu , Zhaoyang Li , Qiuyuan Ai , Xiaofeng Han , Zelong Zheng , Haoyu Wu , Tianyu Fu , Chenxu Zhao , Minghui Wu , Guannan He , Changwei Wang
- URL: https://arxiv.org/abs/2609.34810
- Abstract:
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy’s sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source’s influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of $82.8\%$ and $83.6\%$, WebShop success rates of $75.0\%$ and $82.0\%$, and Search-QA aggregate accuracies of $45.3\%$ and $49.8\%$, respectively. On 3B WebShop, UniOPSD improves over SDAR by $7.0$ percentage points. Our code is available at this https URL
95. SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL
- Authors: Zenghuang Fu , Ningqi Chen , Mingda Jia , Xiaofeng Han , Zhaoyang Li , Qiuyuan Ai , Zelong Zheng , Haoyu Wu , Tianyu Fu , Chenxu Zhao , Minghui Wu , Guannan He , Changwei Wang
- URL: https://arxiv.org/abs/2609.34805
- Abstract:
Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score–outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by $1.31$ and $1.07$ percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected–fresh value gap alongside a near-zero fresh–fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at this https URL
96. STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
- Authors: Wanchun Ni , Tao Qi , Leonel Aguilar , Jiugeng Sun , Marlene Wagner , Verena Zimmermann , Mennatallah El-Assady
- URL: https://arxiv.org/abs/2609.34799
- Abstract:
Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.
97. BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
- Authors: Yuheng Wu , Berk Gokmen , Sujeeth Jinesh , Lauren McLane , Aarav Wattal , Qi Yang Huang , Zhaozhuo Xu , Thierry Tambe
- URL: https://arxiv.org/abs/2609.34785
- Abstract:
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design’s verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B’s RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.
98. Applying Language Models in medical Medicine: Recent Trends and Perspectives
- Authors: Erik Aerts
- URL: https://arxiv.org/abs/2609.34780
- Abstract:
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.
99. Page-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFs
- Authors: Abdelhak kelious
- URL: https://arxiv.org/abs/2609.34776
- Abstract:
We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the complete system scores 0.4450 MRR@10 and 0.4013 Recall@10, compared with 0.3430 and 0.2994 for BM25. Dense retrieval alone and a simple lexical–dense fusion both underperform BM25. Reranking improves the hybrid system, whereas adding query decomposition produces the largest further gain, with higher latency and more detected output artifacts. The complete system slightly exceeds the reported anonymous overall mean on two answer metrics but falls below it on most page-retrieval metrics. These results identify accurate page selection, rather than semantic retrieval in isolation, as the main opportunity for improvement in this setting.
100. Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs
- Authors: Yadong Wang , Siping Yue , Yu Tian , Chuanxing Geng , Xiang Chen
- URL: https://arxiv.org/abs/2609.34772
- Abstract:
Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.
101. When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
- Authors: Tianyi Guan , Jianhui Chen , Liangming Pan
- URL: https://arxiv.org/abs/2609.34771
- Abstract:
Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.
102. Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision
- Authors: Shuxing Zhang , Yongquan Ni , Zhenyu Ding , Yawen Lin
- URL: https://arxiv.org/abs/2609.34768
- Abstract:
Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-modal teacher-student framework that lifts commercial radar to full-body, per-frame, metric 3D mesh reconstruction with per-joint uncertainty. Three innovations: (1) a mesh-foundation-model teacher - SAM 3D Body produces whole-body MHR ground truth (70 joints, 18,439 mesh vertices) from a single RGB frame with zero training, slashing annotation cost by orders of magnitude; (2) StudentPoseFormer - set encoding with masked attention pooling, a temporal Transformer, and a CVAE multi-hypothesis head that outputs both the pose mean and per-joint variance, honestly reporting where the radar cannot see; and (3) a multi-stage ground-truth quality pipeline (confidence gating, depth validation, temporal smoothing, bone-length consistency, bad-frame rejection) plus systematic information-lever ablations. On the public MM-Fi benchmark (same TI IWR6843 sensor, cross-subject), our full configuration reaches 7.45 cm 12-joint MPJPE, with ablations proving the causal value of point accumulation (k = 3, -0.34 cm), Doppler (-0.85 cm; -2 cm at the wrist on fast actions), and velocity loss (-0.27 cm). On our own synchronized radar + RGB-D corpus with block-level held-out splits, the pipeline achieves 21.47 cm end-to-end (per-joint hierarchy from 4.8 cm at the hip to 34.7 cm at the wrist - matching physical information limits), could be improved to 15 cm with ~30k diverse samples, and a scaling law shows sample diversity, not volume, is the binding constraint. Deployment inference is radar-only - no camera, no image.
103. SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing
- Authors: Vasilis Perifanis , Nikolaos Pavlidis , Symeon Symeonidis
- URL: https://arxiv.org/abs/2609.34736
- Abstract:
Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing rarely states what a query actually requires. We introduce SeLMRoute, a routing framework that separates the extraction of candidate-independent semantic evidence from the learning of candidate performance and the application of deployment objectives. A decision model first evaluates a set of interpretable questions about the query, such as its reasoning requirements and use of external knowledge, with each judgment retained as a probability distribution. The resulting probabilistic semantic state is used by a lightweight supervised router to estimate candidate model performance. Routing objectives are applied after performance estimation, which allows the same semantic state to support performance-oriented and cost-aware decisions. On the LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries), SeLMRoute achieves an average accuracy of $72.08\% \pm 0.45$, while grouped five-fold out-of-fold evaluation reaches $72.64\%$, compared with $69.23\%$ for the strongest fixed candidate. The representation achieves the highest mean performance among the evaluated semantic, dense, lexical, and domain-level representations. In a separate 13-model performance-cost setting, SeLMRoute improves performance in all five grouped splits, with a mean PerfGain of $2.66\%$. Our code is available at this https URL .
104. From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram
- Authors: Josef Pavlíček , Petra Pavlíčková , Irena Štrausová
- URL: https://arxiv.org/abs/2609.34735
- Abstract:
Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human composers operate within personal and cultural contexts that influence harmonic decisions and deliberate departures from established patterns. This study investigates six narrative-driven popular songs by Bob Dylan, Johnny Cash, and Ritchie Valens. Original human harmonies are compared with outputs of an explainable computational harmonizer operating on the same melodies without access to the original chord progressions. We examine harmonic vocabulary, functional persistence, repetition, non-diatonic events, and tension-resolution patterns using Chord Wheel Diagrams and BPMN-based representations. Results show that high melody-chord compatibility does not necessarily imply preservation of the original human harmonic decision pattern. Some generated harmonizations retain the economical structure of the reference, while others alter harmonic diversity or suppress distinctive events while remaining compatible with the melody. Rather than quantifying artistic quality, the study introduces narrative-conditioned harmonic structure as a complementary perspective for computational music analysis. The findings suggest that generative systems may benefit from modeling not only harmonic correctness, but also structural identity, context, and human compositional intention.
105. PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs
- Authors: Zhentao Tan , Jianrong Zhang , Ruijie Quan , Yi Yang
- URL: https://arxiv.org/abs/2609.34715
- Abstract:
Physical trajectories contain more than snapshots of a system: they also reveal how its states evolve under governing conditions. However, representation learning for parametric partial differential equations (PDEs) has largely relied on reconstruction-based objectives that emphasize recovering observed physical fields. In this paper, we investigate predictive representation pretraining as an alternative to reconstruction-based learning. We find that predictive representations preserve rich physical information, yet this advantage alone does not ensure accurate field evolution. Based on these observations, we introduce PDE-JEPA for parametric PDE dynamics. Specifically, we first train an encoder using a masked-latent prediction to capture the underlying regularities of PDE dynamics. To explicitly adapt the pretrained representation toward a more dynamics-aligned state space, we then introduce a geometry projector that aligns latent trajectory geometry with the evolution geometry of physical fields. Finally, building on this geometry-aligned latent space, we further develop a physics-structured latent predictor that decomposes the dynamics into parameter-independent evolution and parameter-dependent response components. Extensive experiments on nine widely used PDE benchmarks demonstrate that our framework outperforms existing state-of-the-art methods by an average of 33.4\% in-distribution, while achieving an average improvement of 51.4\% when extrapolating to unseen governing parameters. The project page is available \href{ this https URL }{here}.
106. RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents
- Authors: Hao Li , Hangfan Zhang , Zhiyao Cui , Chunjiang Mu , Yiqun Zhang , Bo Zhang , Danyang Jia , Shuyue Hu
- URL: https://arxiv.org/abs/2609.34712
- Abstract:
Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance-cost trade-off can be improved through within-task large-small model collaboration, as smaller models can handle some stages even when they cannot solve the full task. In this paper, we introduce RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience. Each iteration consists of four stages: Subtask Mining derives subtask definitions and identification rules from training trajectories; Routing Strategy Evolution proposes and evaluates diverse model assignments; Model-Specific Skill Evolution compares routed and large-model-only trajectories to diagnose failures and develop reusable execution skills; and Pareto-Optimal Router Selection updates the Pareto population using historical and newly generated routers while retaining dominated routers as experience for subsequent evolution. Routing between DeepSeek-V4.1-Flash and Qwen3.5-9B, RSI-router consistently surpasses the DeepSeek-only baseline at roughly half the inference cost (48.3%) across five agentic benchmarks. In particular, on ALFWorld, ScienceWorld, and WebShop, it cuts inference cost by 74.7-82.2% while simultaneously improving performance; on Terminal-Bench 2.0, it achieves a 16.7% relative performance gain at 18.0% lower cost. Moreover, RSI-router establishes a stronger performance–cost Pareto frontier than 9 routing methods.
107. FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management
- Authors: Peiyu Zang
- URL: https://arxiv.org/abs/2609.34710
- Abstract:
Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark that studies five configurable factors through controlled comparisons on a single simulator, using paired seeds and a frozen calibration. We evaluate four foundation models and four agent scaffolds. In the Model Track, Coach points Z-scores span 0.19, while Manager points Z-scores span 0.68, with GPT-5.6 showing a sharp rise in passivity under responsibility expansion. Its responsibility ladder rises from 46.1 to 58.1 points with recruitment, then falls to 46.8 under full management, localizing the regression to the final responsibility boundary. Across that boundary, its skipped-decision rate rises from 1.1% to 57.9%. Within the Flash-Pro pair crossed across every scaffold, scaffold choice changes Manager points Z-scores by up to 0.48 relative to the fixed stateless scaffold. The 3Y cohort shows a directional reversal in mean ranking between years one and three, while a selected Claude Code+Pro configuration peaks in year three and remains below that peak, showing that responsibility scope and horizon expose behavior changes that a single headline score conceals.
108. ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems
- Authors: Tarun Chintada , Neelamadhav Gantayat , Ishaan Romil , Renuka Sindhgatta , Soujanya Soni , Sameep Mehta
- URL: https://arxiv.org/abs/2609.34701
- Abstract:
Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.
109. VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience
- Authors: Siqi Zhang , Meng Wei , Chenyang Wan , Shaohao Zhu , Shufan Shen , Xihui Liu , Zhihua Wei , Tai Wang , Jiangmiao Pang
- URL: https://arxiv.org/abs/2609.34687
- Abstract:
Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.
110. Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents
- Authors: Xi Wang , Songlei Jian , Yiming Zhang , Bin Ji , Zhaoye Li , Ma Jun , Baosheng Wang , Jie Yu
- URL: https://arxiv.org/abs/2609.34686
- Abstract:
As large language models increasingly operate as tool-using agents, post-jailbreak safety feedback is often assumed to serve as a reliable safeguard; however, how lingering jailbreak context shapes subsequent agent behavior remains largely unexplored. To systematically examine this dynamic, we introduce a paired continuation framework across 192 parent tasks spanning 42 domains, evaluating 12,148 analyzed continuation pairs (curated from a 12,288-pair initially design) across eight diverse agents. We find that identical safety feedback induces sharply model-dependent behavioral routing rather than uniform protection: redirecting unsafe trajectories toward legitimate completion (\emph{rescue}), sustaining unauthorized execution (\emph{persistent unsafe}), or triggering over-refusal on benign tasks (\emph{collateral loss}). Through layer-wise activation patching, we discover a shared \emph{late-commit pattern} where causal intervention effects surge sharply near the final layers (relative depths of 0.958–0.984) despite an over 30-fold variation in peak magnitude across architectures. Crucially, critical-layer representations correlate with macroscopic routing outcomes, and intervening at these layers causally alters concrete next-step tool actions. Building on this causal foundation, we test whether localized intervention-derived features can serve as predictive proxies for full-trajectory routing outcomes on unseen parent tasks under leave-one-parent-task-out evaluation, finding that they provide viable predictive signals in responsive agents with peak ROC AUCs reaching 0.675 for \emph{rescue}, 0.777 for \emph{collateral loss}, and 0.702 for \emph{persistent unsafe}. These findings establish a mechanistic lens and a predictive baseline for anticipating the safety and utility trade-offs of post-jailbreak feedback in autonomous agents.
111. A General Harness for Protein Foundation Model Fitness Prediction
- Authors: Yang Tan , Qijia Tian , Gangyu Sun , Bozitao Zhong , Mingchen Li , Yuanxi Yu , Nanqing Dong , Liang Hong
- URL: https://arxiv.org/abs/2609.34654
- Abstract:
Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.
112. OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
- Authors: Zongshang Shen , Wangsong Yin , Daliang Xu , Mengwei Xu , Xuanzhe Liu
- URL: https://arxiv.org/abs/2609.34653
- Abstract:
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.
113. Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
- Authors: Weiyuan Li , Jinghan Xu , Aili Chen , Xintao Wang , Shuang Liang , Jiaqing Liang , Deqing Yang
- URL: https://arxiv.org/abs/2609.34649
- Abstract:
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.
114. MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics
- Authors: Danilo Gusicuma , André Freitas
- URL: https://arxiv.org/abs/2609.34636
- Abstract:
This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying structural and causal constraints. The benchmark tests whether answers preserve simulator-licensed ambiguity, quantified claims, episode-graph transition evidence, repairs, and trace-support judgments. Its 1,120 items are generated deterministically from admissible interpretation sets, component states, scenario restrictions, confluence constraints, and derivation steps across 18 catalog mechanisms and six task families. Each mechanism undergoes converter checks of structure and topology and behavioral checks against quantitative simulations. GPT-5.5 accuracy decreases as family-specific mechanistic complexity increases, from 76.1% in the lowest-complexity bucket (B1) to 38.0% in the highest-complexity bucket (B4). The negative association remains after controls for rendered-prompt and expected-answer length. These results show that qualitative simulators can support auditable NLP benchmarks for mechanistic inference.
115. A Persistent State for Auditable Mixture-of-Experts Routing
- Authors: Abdurrahman Javat , Allan Kazakov
- URL: https://arxiv.org/abs/2609.34634
- Abstract:
Mixture-of-Experts (MoE) models repeatedly route tokens to sparse subsets of experts, but conventional routers expose no routing-specific record of how cross-layer influences accumulate. We introduce Scratchpad-Augmented Mixture-of-Experts (SA-MoE), which gives each router access to a low-dimensional persistent state that is not provided to the experts. Learned layerwise writes update this state, and their realized post-update changes exactly decompose the state-mediated contribution to any later routing margin, forming a routing ledger. Across sparsely upcycled SmolLM2- and Gemma-based models and three independent training seeds per architecture, this pathway adds less than 1% analytical forward compute and is strongly used by trained routers: local removal of its router contribution changes the selected Top-2 expert set in 87.6% and 69.9% of decisions, respectively. Relative to a matched latest-write-only control, persistent accumulation increases long-horizon future-routing accessibility by 19.4 and 12.2 percentage points, with positive effects in every seed. More than 90% of absolute ledger contribution comes from non-recent writes in both families, and full-forward suppression of ledger-selected writes changes later routing and output distributions. The ledger is an exact provenance object for the persistent-state pathway, not a complete causal explanation of routing. Sensitivity-aware scores better predict full-forward intervention effects, and post-hoc methods recover related cross-layer attribution without architectural modification. SA-MoE instead makes one routing-specific computational history explicit and directly inspectable within the model’s natural forward computation.
116. LLMs for Executable Multi-Agent System Specification Generation
- Authors: Andreas Kouvaras , Periklis Mantenoglou , Alexander Artikis
- URL: https://arxiv.org/abs/2609.34619
- Abstract:
MAS specifications express the effects of the actions of the agents and their environment, as well as other temporal phenomena, such as the intervals during which an agent may perform an action. The specification of a MAS should also be executable in order to allow for run-time monitoring. Constructing the specification of a MAS requires formal language expertise, while machine learning techniques depend on labelled data which are rarely available. To address these issues, we propose
genRTEC', a method that leverages pre-trained Large Language Models (LLMs) to generate executable MAS specifications, in the language of theRun-Time Event Calculus’ (RTEC), from natural language descriptions. genRTEC constructs MAS specifications with complex hierarchical and cyclic dependencies based only on short natural language descriptions of the concepts involved. We present an extensive empirical evaluation of genRTEC, spanning various MAS specifications, including both a qualitative and a quantitative assessment. Our results demonstrate that genRTEC constructs executable MAS specifications of high predictive accuracy without compromising reasoning efficiency.
117. After the Fix: How Corrected Agent Histories Transfer to Related Tasks
- Authors: Yanfei Zhang , Xu Lin
- URL: https://arxiv.org/abs/2609.34603
- Abstract:
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-state inheritance, under eleven conditions. ThinkingBox’s Full/Skill/Hybrid correction gains are 44/29/32 percentage points, with corrected performance 25/22/18 points above independence; inference weakens at the task-family level. Yet 12 of Full’s 15-point larger correction gap over Skill come from worse uncorrected performance, not better corrected memory. Moreover, 22 of Full’s 46 upward transitions restore observed baseline success. Neither APEX regime establishes comparable aggregate correction benefits. Action evidence connects workflow gains with reusable obligations and convention conflicts with source-local choices. Text APEX’s accepted execution reaches 52% versus its summary’s 40%, without robust global/group-level superiority or an estab- lished advantage over independence. Smaller handoffs reduce input but increase calls. The value of repairing experience is therefore distinct from the value of reusing it: memory updates require both a previous-version reference and a fresh-start reference.
118. TULIP: Targeted LLM Unlearning at Layers Identified Per-Input
- Authors: Yejin Kim , William F. Shen , Seokwon Jung , Daeun Park , Seong Joon Oh
- URL: https://arxiv.org/abs/2609.34591
- Abstract:
Representation-level unlearning intervenes on the intermediate hidden states of LLMs. Although knowledge is distributed across layers, existing methods operate at a single fixed layer for the entire forget set. We ask whether such a fixed layer is sufficient. To answer this, we design a hijacking experiment that grafts hidden states of the target model into an oracle trained only on the retain set. The oracle cannot produce the forget answer on its own, yet it produces the answer from the grafted state. Thus, the answer is formed at an intermediate layer and merely read out afterward, so unlearning should focus on formation, not readout. Moreover, the layer where formation ends varies widely across inputs. Motivated by these findings, we propose Targeted Unlearning at Layers Identified Per-input (TULIP). For each input, TULIP uses the logit lens to locate the formation-readout boundary and removes the hidden state’s alignment with the forget answer’s unembedding vector there. TULIP consistently outperforms output- and representation-level baselines on TOFU, PISTOL, and WMDP across Llama, Qwen, and Zephyr models. It also remains robust to paraphrase and quantization attacks. Beyond standalone use, its per-input layer selection serves as a plug-and-play component that further improves existing methods.
119. SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences
- Authors: Mingyue Huo , Shivam Mehta , Bhavin Jawade , Yinghong Lan , Haoqi Li
- URL: https://arxiv.org/abs/2609.34582
- Abstract:
Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher’s dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.
120. Calibrated Uncertainty for Informative Path Planning in Aquatic Environmental Monitoring
- Authors: Samuel Yanes Luis , Alejandro Casado Pérez , Alejandro Mendoza Barrionuevo , Dame Seck Diop , Sergio Toral Marín , Saniel Gutiérrez Reina
- URL: https://arxiv.org/abs/2609.34577
- Abstract:
Informative Path Planning for scalar field reconstruction uses predictive uncertainty to direct sensing vehicles toward maximally informative locations. Gaussian Processes provide this signal but their stationary isotropic kernels are misspecified for non-homogeneous phenomena such as oil spills, producing miscalibrated estimates that degrade planning. We investigate whether replacing the Gaussian Process with a well-calibrated Deep Ensemble improves path planning outcomes, and whether uncertainty quality interacts with the choice of planning algorithm. Five strategies ($\epsilon$-Greedy, Value Greedy, Uncertainty Greedy, Monte Carlo Tree Search, and Receding Horizon Orienteering) share a common Deep Ensemble backbone trained on physics-based oil spill simulations. On held-out stochastic spill scenarios, the Deep Ensemble reduces normalised reconstruction error by $83\%$ relative to the Gaussian Process baseline. Crucially, well-calibrated uncertainty amplifies the importance of the planning strategy: the performance gap between algorithms is negligible under miscalibrated models but becomes substantial under the ensemble, where multi-step lookahead planners outperform greedy selection by up to $32\%$ in reconstruction error and achieve IoU above $0.85$. Monte Carlo Tree Search is the recommended planner, matching Orienteering in reconstruction quality at an order-of-magnitude lower computational cost.
121. Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
- Authors: Hengrui Zhang , Yuhu Cheng , C. L. Philip Chen , Xuesong Wang
- URL: https://arxiv.org/abs/2609.34575
- Abstract:
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
122. Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence
- Authors: Rohit Saxena , Utkarsh Upadhyay
- URL: https://arxiv.org/abs/2609.34572
- Abstract:
Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model’s subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.
123. PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations
- Authors: Rui Xu , Yinghui Xu , Libo Wu
- URL: https://arxiv.org/abs/2609.34571
- Abstract:
Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model’s activation space and apply Euclidean operations—addition, scaling, and linear interpolation—under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold’s intrinsic geometry—local metric tensors, geodesic distances, and Ollivier-Ricci curvature—and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.
124. FlowState: Execution State as Memory for Long-Horizon LLM Agents
- Authors: Minghao Li , Bangyan Li , Zifan Wang , Yulong Li , Hu Xu , Gan Zhang , Jingtong Wu , Wenqiang Xu
- URL: https://arxiv.org/abs/2609.34565
- Abstract:
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $\tau^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.
125. SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
- Authors: Bingqing Jiang , Guoxi Zhang , Jasper Wang , Auric Wang , Bingning Wang , Tianyi Lin , Zichao Yu , Yujin Han , Ziye Ma , Difan Zou
- URL: https://arxiv.org/abs/2609.34557
- Abstract:
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.
126. SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning
- Authors: Jaeho Jung , Sung Hoon Jung
- URL: https://arxiv.org/abs/2609.34548
- Abstract:
Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLACT, which reflects only on the end-goal at each step without explicitly considering intermediate sub goals. To address this problem, we propose SGG-ReflAct (Sub-Goal Guided Re flAct), a reasoning backbone that integrates sub-goals generated through a single path LLM planner into the reflection process. We further extend this framework to BeamSGG-ReflAct, which replaces the single-path planner with a beam search based LLM planner for structured plan exploration. We run experiments on ALF World, ScienceWorld, and Jericho with multiple LLM models. SGG-ReflAct out performs REFLACT in nearly all settings, achieving best success rate gains of 14.9 percentage points on ALFWorld and 8.0 percentage points on ScienceWorld with Llama-3.1-8B-Instruct. Our experimental analysis shows that SGG-ReflAct re duces hallucinated actions and achieves its largest gains on procedurally ordered tasks. Furthermore, experimental results with BeamSGG-ReflAct show that the backbone’s effectiveness depends on plan quality: explicitly specifying the re quired operations recovers gains that plan searching alone cannot achieve. These results demonstrate that SGG-ReflAct offers a practical and highly effective rea soning backbone, enabling LLM agents to achieve reliable performance in com plex, long-horizon tasks through easy integration.
127. Remember Before You’re Asked: MemDream for Self-Probing Memory Evolution
- Authors: Mingfei Lu , Mengjia Wu , Runsong Jia , Zhe Luo , Yi Zhang
- URL: https://arxiv.org/abs/2609.34545
- Abstract:
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
128. APOLO: Automatic Prompt Optimization for Ontology Learning
- Authors: Huu Tan Mai , Roman Kochnev , Cuong Xuan Chu , Lukas Lange , Heiko Paulheim , Daria Stepanova
- URL: https://arxiv.org/abs/2609.34540
- Abstract:
Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimization for Ontology Learning, by casting OL as an explicit prompt optimization problem over LLM modules. To obtain training data, we employ a multi-agent system that generates text-ontology pairs from existing expert-curated ontologies. We then propose two ontology learner architectures: a greedy and an autoregressive learner, and optimize both using GEPA, a greedy evolutionary prompt optimizer built on DSPy. Experiments on two ontologies - a biomedical (DOID) and a plant ontology (PO) show consistent improvements after optimization across nearly all model and mode combinations, with autoregressive learners achieving the largest gains. Our results demonstrate that prompt optimization is a viable and lightweight alternative to fine-tuning for OL, and that the autoregressive formulation better captures ontological structure than the greedy approach.
129. The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions
- Authors: Xiaoting Lyu , Xinbo Ma , Yufei Han , Hangwei Qian , Ziyang Lin , Bin Wang , Bin Wang , Wei Wang
- URL: https://arxiv.org/abs/2609.34537
- Abstract:
Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations), a benchmark for evaluating scientific agents under scientifically plausible perturbations throughout multi-turn problem solving. \textsc{SciARP} transforms 620 scientific problems into interdependent tasks of 3–13 turns and defines 13 perturbation types spanning problem understanding, evidence processing, reasoning, and conclusion formation. Clean and perturbed versions of each task are independently executed under matched settings, producing paired live trajectories for evaluating both task success and process reliability. Experiments across eight LLMs from four model families reveal three key robustness characteristics. First, different classes of scientific perturbations exhibit distinct robustness profiles and can decouple task progression from scientific reliability: agents may continue advancing through the task even after their information or reasoning has become unreliable. Second, stronger clean-task performance does not necessarily translate into stronger robustness, as models with higher clean-task accuracy can exhibit larger degradation under perturbation. Third, perturbation effects exhibit strong temporal dynamics: they may remain latent for multiple turns before emerging and subsequently propagate through downstream dependencies. Together, these findings show that current scientific agents remain insufficiently robust to scientifically plausible perturbations, with failures often remaining undetected, propagating, and resisting recovery.
130. PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use
- Authors: Mingfei Lu , Mengjia Wu , Yi Zhang
- URL: https://arxiv.org/abs/2609.34526
- Abstract:
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores ($\Delta$) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6\% to 18.3\% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.
131. EOPSA: Efficient On-Policy Self-Distilled Safety Alignment
- Authors: Qirui Liu , Yichen Sun , Yan Wang , Yu Mi , Wei Cao , Yue Shen , Zhixuan Chu , Kui Ren
- URL: https://arxiv.org/abs/2609.34519
- Abstract:
On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher’s corrective efficacy degrades precipitously as the student’s generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by $\sim$50% and backpropagates through merely $\sim$2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.
132. Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
- Authors: Xingtong Yu , Jiarun Zhou , Guanlin Ding , Wenkang Wei , Jiarui Liu , Chang Zhou , Fangzhou Ge , Chenyi Xu , Xikun Zhang , Renqiang Luo , Jie Zhang , Hong Cheng , Xinming Zhang , Hui Zhang , Yuan Fang
- URL: https://arxiv.org/abs/2609.34510
- Abstract:
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at this https URL .
133. Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
- Authors: Manya Singh , Arjun Pakrashi
- URL: https://arxiv.org/abs/2609.34506
- Abstract:
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement ($\rho = 0.24–0.55$), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.
134. PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
- Authors: Xijing Wang , Yinsheng Yao , Jinru Ding , Yidong Jiang , Ziwen Xu , Yiwen Jiang , Jie Xu , Dawei Cheng
- URL: https://arxiv.org/abs/2609.34492
- Abstract:
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs’ ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at this https URL .
135. When Does Structured Knowledge Help Neural Theorem Proving?
- Authors: Sareh Nabi , Roland Vogl , Marzieh Nabi
- URL: https://arxiv.org/abs/2609.34460
- Abstract:
Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analogies, generalizations, cross-domain bridges) remains implicit. We introduce MathAgent, which builds this layer as a knowledge graph, MathKG, and uses it to augment LLM theorem provers. MathKG connects 364 Mathlib theorems and definitions by 9,434 typed semantic edges inferred via LLM-based relation extraction anchored to verified Mathlib declarations. We run a controlled ablation across four augmentation modes (no context, knowledge-graph context, Mathlib retrieval, both) and five models: Qwen3-8B/32B, their Lean-specialized derivatives Goedel-Prover-V2-8B/32B, and Claude Sonnet 4.6, on miniF2F, plus PutnamBench and MathOlympiadBench for Sonnet. Three findings emerge. (i) Specialization dominates augmentation: Lean fine-tuning adds 33-38 percentage points of solve rate in every mode, and a specialized 8B model beats a $4\times$ larger general one by 29-35 points, while no augmentation mode improves solve rate by more than 3 points. (ii) Augmentation is capability-conditioned: knowledge-graph context helps small models but hurts large ones, with the specialized model gaining more relative to its general base at every scale. (iii) Yet the augmentation modes solve different problems: an oracle selecting the best mode per problem solves 6% to 58% more than the unaugmented prover, a complementarity effect that strengthens on harder problems (32% more on PutnamBench). These results motivate adaptive strategies that select augmentation by model capability and problem. Code, data, and artifacts are available at this https URL
136. Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning
- Authors: Yijie Sun , Sanquan Sun , Yanda Zhu , Yuanyang Zhu , Yaohua Hu , Chunlin Chen
- URL: https://arxiv.org/abs/2609.34459
- Abstract:
Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challenges, we propose a novel interpretable framework, called escaping local views (ELV), which introduces semantically structured latent concepts to render policy decisions transparent. Specifically, each agent extracts low-dimensional semantic concepts from its local observation and action-observation trajectory. These concepts are jointly encoded into a contextual latent variable via a variational autoencoder (VAE), which builds a bridge between local views and global semantics. To explicitly model the decision of each agent, we employ a dual-path attention mechanism in which one module estimates the salience of individual concepts relative to the global context, while the other captures higher-order cooperative patterns with pairwise concept interactions. Furthermore, we incorporate a concept prediction module that derives an intrinsic reward from next-concept prediction errors, which incentivizes agents to explore regions of semantic novelty. Experiments in multiple environments verify that ELV not only achieves competitive performance but also explicitly provides how agents reason about their decisions.
137. CORTEX: Learning to Share and Specialize in Dense Language Models
- Authors: Chuiyang Meng , Ming Tang , Vincent W.S. Wong
- URL: https://arxiv.org/abs/2609.34449
- Abstract:
Large language models are trained on heterogeneous data mixtures, where different knowledge domains require both shared knowledge and specialization. Existing modular approaches typically impose explicit components or discover modules through interpretability analysis after training. In this work, we propose CORTEX, a learning dynamics-inspired framework that learns internal modularization within dense language models. CORTEX partitions trainable matrices into parameter groups and learns module assignments from domain-conditioned gradient and cross-domain gradient similarity. We introduce the selective lesion score and module-domain mutual information to characterize the target-domain lesion effects and alignment, and analyze how module assignment affects the trade-off between assignment bias and update magnitude. Experiments with 160M, Qwen3-8B, and Qwen3-32B backbone models show that CORTEX achieves the highest synthetic-domain exact match and largest average perplexity reduction, while remaining competitive on real-domain evaluations and forming identifiable modules.
138. Social Circuits behind Multi-agent Echo Chambers
- Authors: Chuiyang Meng , Wenlu Yu , Ming Tang , Cheng Li
- URL: https://arxiv.org/abs/2609.34444
- Abstract:
Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent’s internal activations and affects its decision. In this work, we introduce Social Circuits, a framework for tracing message effects through receiver activations. We compare the receiver’s answers before and after changing a message. Then, we restore selected activations recorded under the original message to determine how much of the message effect these activations reproduce. Based on Social Circuits, we propose Circuit-Guided Deliberation (CGD), which learns to select useful messages using receiver activation changes. We establish when activation replacement preserves receiver decisions and bound the gap between CGD’s task performance and the best achievable through message selection. Experiments show that receiver activation changes explain the message effects and guide message selection that improves the task performance. Across three models and four datasets, CGD achieves the highest or joint-highest average accuracy in our main comparisons while generating fewer tokens than multi-agent baselines.
139. Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation
- Authors: Hanzhong Zhang , Jindong Wang , Siyang Song
- URL: https://arxiv.org/abs/2609.34419
- Abstract:
Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG methods rely on end-to-end architectures that directly map speaker behaviours to listener expressions without an explicit intermediate internal state. We reformulate FRG as generation mediated by a structured internal-state process and propose the \textbf{Intentional Agent}, which shifts FRG from direct stimulus-response mapping to stimulus-grounded generation through explicit intermediate states. To represent temporal internal-state evolution, we propose an internal dynamics model that integrates emotional drives with an iterative Inner Thought Flow (ITF) within a structured intermediate state used for subsequent generation. This state can continue to update during conversational silences. Furthermore, to bridge abstract internal states with physiological actions, we formulate FRG as a downstream affective mapping from this latent thought flow to facial expressions. Experiments on the REACT 2025 dataset show an FRDist of 72.39 and an FRDiv of 0.5057; perceptual plausibility is evaluated separately through blinded human ratings. A blinded human evaluation of 96 reactions found no significant difference in mean score between Full and ground truth ($5.527$ vs.\ $5.195$, $p_{\mathrm{Holm} }=.076$), while Full significantly outperformed Event-Triggered and Heuristic-Only (both $p_{\mathrm{Holm} }<.001$). The Reaction Quality Scorer (RQS) correlated strongly with human judgements (Pearson $r=.855$; Spearman $\rho=.821$, both $p<.05$), supporting its use as an automatic metric. These results underscore the immense potential of endogenous dynamics in building highly autonomous, human-like agents.
140. OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue
- Authors: Rui Xu , Yikai Zhang , Aili Chen , Zicheng Zhao , Xu Yinghui , Libo Wu
- URL: https://arxiv.org/abs/2609.34418
- Abstract:
Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure—sharply peaked at character-critical tokens yet diffuse at generic utterances—and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.
141. Mathematics for and by human cognition: A resource-rational search for bottlenecks in problem-solving
- Authors: Sneha Aenugu
- URL: https://arxiv.org/abs/2609.34410
- Abstract:
Human cognitive constraints are generally viewed as limiting factors in problem-solving. We argue that these constraints can instead play a critical role in driving advances in mathematics and beyond. We propose a theory of mathematical abstraction as a resource-rational search for bottlenecks in problem-solving. Bottlenecks arising from cognitive constraints create pressure to restructure existing knowledge, potentially giving rise to novel formalisms with applications beyond the problems that originally motivated them. Drawing on episodes from the history of mathematics, we illustrate how such bottlenecks can drive the development of novel abstractions and examine how cognitive constraints and affective responses shape this process. Finally, we discuss the implications of this account for machine mathematical discovery and argue that incorporating human-like constraints may facilitate the discovery of useful mathematical abstractions.
142. SkillFocus: Evolving Agent Skills via Capability Decomposition
- Authors: Ning Wang , Zhiren Gong , Bingdong Li , Peng Yang , Aimin Zhou
- URL: https://arxiv.org/abs/2609.34397
- Abstract:
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24\% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task–capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.
143. Org-Agent: Beyond Personal Assistants Towards Organizational Agents
- Authors: Luyao Zhuang , Yujing Zhang , Zijin Hong , Yilin Xiao , Xiao Huang
- URL: https://arxiv.org/abs/2609.34392
- Abstract:
Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-user memory and knowledge use. Both capabilities are governed by organizational constraints across three aspects: user identity, authority, and access permissions; the attribution and temporal validity of information; and rules for resolving conflicting requirements across users and completion requirements for joint decisions. These constraints shape what information or decisions must be obtained before an action can proceed and what conditions must be satisfied during its execution. Motivated by this, we introduce Org-Agent, a unified constraint-centric reasoning framework that organizes task execution in three stages. Specifically, Org-Agent decomposes a task into atomic subtasks and constructs a task dependency graph whose edges encode the dependencies among them. Building on this graph, it schedules the subtasks in dependency order through topological sorting. It then executes each subtask while accounting for the task’s constraints, supported by evidence-acquisition and memory-management tools. Experiments on MUSES-Bench and GroupMemBench demonstrate the effectiveness of Org-Agent on both capabilities, and ablations further support the contributions of dependency modeling and tool use.
144. PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents
- Authors: Hanzhong Zhang , Ziwei Xiang , Weicheng Xie , Shizhe Liu , Siyang Song
- URL: https://arxiv.org/abs/2609.34372
- Abstract:
The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent’s memory processing to be inconsistent with the pre-defined personality, and makes it difficult to validate whether agent behaviours follow this personality. In this paper, we propose Personality-Integrated Memory (PersMem), which integrates personality into the agent’s memory processing pipeline, making it consistently personality-dependent. PersMem processes memory using four steps, where the personality is mapped to operation-specific parameters controlling: (i) affective appraisal annotating emotion states of the user input; (ii) retention of previously stored memories along with the current input; (iii) passive affect-driven memory retrieval exploring memories similar to user input in semantics and personality-guided emotions; and (iv) active goal-driven memory retrieval that refines and selects passively retrieved memories for the reply. Consequently, consistency with the pre-defined personality can be examined by inspecting memory-processing traces during human-agent interactions. We evaluate these personality-dependent differences in attachment and Big Five settings. PersMem exceeds the chance baseline for four-way attachment classification by 23.1 percentage points. In Big Five dialogue comparisons, PersMem achieves 67.5% accuracy, 6.7 percentage points above a baseline using uniformly sampled memories. On CoSER, PersMem achieves an average score of 66.13, with scores of 69.33 for Character Fidelity and 84.33 for Storyline Quality. Together, these results show that PersMem produces distinguishable personality-related memory-processing patterns.
145. CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency
- Authors: Junwoo Bae , Jin-Hyun Ahn
- URL: https://arxiv.org/abs/2609.34360
- Abstract:
Split Federated Learning (SFL) enables resource-constrained clients to participate in collaborative training, but vanilla SFL exchanges smashed data and gradients at every batch, which incurs significant communication overhead. Recent methods reduce this overhead with an auxiliary network at the client-side cut layer. However, we identify that this approach makes the client optimize a local objective that differs from the end-to-end objective, which fundamentally limits the collaborative training between the client and the server. We propose Compensated Feedback based SFL (CoeF-SFL), a communication-efficient framework that retains the end-to-end objective without any auxiliary network. In CoeF-SFL, the client and the server exchange the smashed data and the gradients once per round and reuse them during local training. Since this reuse makes the gradients stale on the client side, we compensate them with a curvature-based correction in the activation space and develop two variants. CoeF-D approximates the Hessian with a diagonal gradient outer product, while CoeF-J exploits the tractable Jacobian-based Hessian of a surrogate loss that upper-bounds the true loss. We provide the theoretical background of each method, characterizing its compensation. Across vision and language tasks, model capacities, cut layers, and data distributions, CoeF-SFL significantly outperforms auxiliary-network-based methods under the same communication frequency, and the improvement is most substantial on vision tasks. Code is available at this https URL
146. Improving Large Language Models for Code through Runtime Program-State Reasoning
- Authors: Hongwei Li , Spandan Garg , Yufan Huang
- URL: https://arxiv.org/abs/2609.34359
- Abstract:
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
147. SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception
- Authors: Hu Xu , Chun Li , Siyuan Qiu , Zeyan Li , Jianfeng Xu
- URL: https://arxiv.org/abs/2609.34353
- Abstract:
Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open: which remote evidence is indispensable given the receiver’s own observation? We introduce a closure-fidelity perspective on ego conditioned remote perception. Under a finite deductive abstraction and explicit conditions, its rate–distortion function decomposes over an irredundant core, and the exact zero-distortion rate becomes $P_A H(\pi_A)$. This analysis suggests a concrete design principle: transmit compact evidence and recover derivable context with bounded receiver-side inference. Guided by this principle, SemRD-V2X is an operational neural proxy that combines exact-budget BEV support selection, pointwise channel compression, and masked shared-weight reconstruction before standard fusion. Experiments on simulated V2XSet and real-world DAIR-V2X validate the resulting design. In a controlled five-run V2XSet comparison against a locally reproduced V2X-ViT-v1 baseline on one Tesla V100, SemRD-V2X reduces the analytical feature payload by $26.6\times$ while improving AP@0.5/AP@0.7 by 4.13/8.57 points, with 3.81\% additional mean compute latency. These results position closure fidelity as both an analytical lens and an actionable design principle for communication-efficient cooperative perception.
148. Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation
- Authors: Michael Vasilakakis (1), Dimitris K. Iakovidis (1) ((1) Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
- URL: https://arxiv.org/abs/2609.34349
- Abstract:
Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic and deep generative models, where their interpretability remains limited. This paper proposes a novel fuzzy distribution modeling methodology for synthetic tabular data generation based on fuzzy sets theory. Feature distributions are represented using fuzzy sets and feature dependencies are modeled through Fuzzy Cognitive Maps, resulting in a low-parameter, and an interpretable data representation. Synthetic samples are generated by sampling fuzzy concepts rather than raw values, enabling native support for mixed data types, missing values, and domain constraints. The methodology further supports linguistic queries and IF-THEN reasoning, facilitating transparent simulation of decision-making processes. Experimental results on benchmark datasets demonstrate competitive performance with respect to utility, fidelity and privacy compared to state-of-the-art methods, while offering substantially improved interpretability. These results establish fuzzy distribution modeling as a principled and effective approach for synthetic tabular data generation in fuzzy systems and decision support applications.
149. SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing
- Authors: Zhiwei Chen , Tianchun Wang , Zhongtao Rao , Haiming Zhu , Ding Cao , Tianxiang Zhao
- URL: https://arxiv.org/abs/2609.34342
- Abstract:
A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents’ behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupported strategic assumptions, inconsistent opponent estimates, and harmful interference from irrelevant historical interactions. To address these issues, we propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate. SAGE first anchors reasoning to an equilibrium policy that provides a strategically valid prior. It then conditions deviations from this anchor on a soft belief over opponent behavioral tendencies, enabling opponent-specific exploitation. Finally, SAGE distills strategically related interactions into counterfactual hypotheses about previously missing considerations, allowing past experience to recalibrate the model’s reasoning. We evaluate SAGE on three repeated imperfect-information games: Leduc Hold’em, Liar’s Dice, and Goofspiel, against various opponent types in each game. Compared with reasoning-intensive LLM agents, including Suspicion-Agent, ReTA, Agent-Pro, EMO, and Hypothetical Minds, SAGE achieves up to a 127.6% payoff improvement in Liar’s Dice while reducing input and output token usage by up to 80% and 90%, respectively. In direct match-up play, it attains non-negative mean payoff against 5/10, 8/10, and 8/10 evaluated opponents in Leduc Hold’em, Liar’s Dice, and Goofspiel, respectively, while using relatively fewer tokens. Code is available at this https URL .
150. Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
- Authors: Chanuk Lee , Minki Kang , Sangwoo Park , Woongyeong Yeo , Jinheon Baek , Sung Ju Hwang
- URL: https://arxiv.org/abs/2609.34327
- Abstract:
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
151. Test-Time Scaling via Budgeted Multi-Attribute Verification
- Authors: Bo Xue , Ji Cheng , Shen-Huan Lyu , Yuanyu Wan , Shuang Qiu
- URL: https://arxiv.org/abs/2609.34322
- Abstract:
Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated along several costly attributes, and the goal is to certify as many candidates as possible whose mean scores exceed the prescribed thresholds on all attributes. We propose \textsc{BMA-GAI}, an algorithm that combines cost-aware arm selection with adaptive sampling of attributes. Every observation serves both to guide adaptive allocation and to support anytime-valid certification, which removes the need for a separate confirmation stage. We establish an asymptotic coverage guarantee for \textsc{BMA-GAI} and derive a matching information-theoretic converse that characterizes the intrinsic complexity of the problem, thereby proving that \textsc{BMA-GAI} is first-order optimal away from critical budget levels. Experiments on synthetic benchmarks and an LLM answer-verification task show that \textsc{BMA-GAI} allocates the verification budget more efficiently and certifies more high-quality candidates than competing methods.
152. Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models
- Authors: Kang Yang , Gaofeng Dong , Liying Han , Mani Srivastava
- URL: https://arxiv.org/abs/2609.34316
- Abstract:
This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.
153. ControlScope: Workflow Revision and Reliability in LLM Agents
- Authors: Jingjie Ning , Xueqi Li , Yibo Kong , Dongting Li
- URL: https://arxiv.org/abs/2609.34313
- Abstract:
How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call’s data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the actions an agent selects. We evaluate one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld. Across two source programs per task and three reasoning-reviewer draws on 20 filesystem tasks, FULL completes 15-16 tasks versus 13 for KEEP; across four fast draws it completes 10-13 versus 13. Fresh student-record confirmation reproduces a batch-read repair. ALFWorld fast panels yield KEEP/ARG/FULL scores of 85/86/87 on 87 tasks across 52 scenes and 134/134/127 on 134 tasks across four scenes; reasoning on the 87-task cohort also yields 85/86/87 with substantial review cost. An AppWorld V1 official-test panel of 585 task instances from 195 scenario templates shows small net differences. Frozen replays expose viable agent-written replacements interrupted by later revision in two failed file-organization runs. An offline source-trajectory midpoint comparison shows later reviews completing an insufficient repair. Five-call protection saves 19.4% of logged model output and loses one success across 20 fresh source runs. An argument-only shortcut shows that the broader sampled policy can overlook a cheaper successful edit available in both operation sets. These outcomes tie repair access to actual choices and subsequent execution.
154. MoSPR: Histology-to-Gene Expression Prediction with Morpho-Spatial Macrostates and Low-Rank Molecular Programs
- Authors: Dongmyung Shin , Geongyu Lee , Yesung Cho , Park Jong Bae
- URL: https://arxiv.org/abs/2609.34280
- Abstract:
Predicting molecular profiles from histopathology remains challenging because whole-slide images contain spatially organized, heterogeneous tissue patterns, while gene expression comprises thousands of correlated targets. We introduce MoSPR (Morpho-Spatial Program Regression), a linear framework that couples an adjacency-informed histology representation with a low-rank molecular basis. MoSPR clusters frozen patch embeddings into morphology microstates, aggregates their spatial adjacencies across the training cohort, and groups microstates with similar adjacency patterns into shared macrostates. Each slide is then represented by global morphology and macrostate-specific deviations, which are linearly mapped to coefficients of a training-derived low-rank gene-expression basis. Across three cancer cohorts from The Cancer Genome Atlas, MoSPR achieves the highest mean gene-expression prediction scores among all evaluated methods. Without pathway-level supervision, pathway scores derived from its predicted expression profiles rank first in eight of nine comparisons across three pathway collections. Ablation studies on the breast cancer cohort show complementary gains from adjacency-derived macrostate representation and low-rank molecular prediction. Moreover, with half of the training data on this cohort, MoSPR exceeds the full-data gene-prediction score of the strongest competing baseline. Finally, its linear formulation enables exact decomposition of each predicted expression profile into global and macrostate-specific molecular contributions, providing an interpretable link between spatially coherent macrostate regions and their associated molecular programs. Our code is available at this https URL .
155. BIABench: Evaluating AI agents on real-world bioimage analysis tasks
- Authors: Zixuan Pan , Davide Panzeri , Lukas Johanns , Marilin Moor , Yu Zhou , Hedi Peterson , Yiyu Shi , Jianxu Chen
- URL: https://arxiv.org/abs/2609.34274
- Abstract:
Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth. The tasks span eleven analysis subtasks and modalities from H&E histology to single-molecule localization microscopy. Each submission receives an outcome score, which compares the output files with the ground truth using field-standard metrics, and a process score, in which a vision-language model judges method choice and quality control against an expert-written rubric. We evaluated general-purpose and biology-specific agents across several language models, with repeated runs of every task. Routine two-dimensional tasks were solved well, but on some tasks that added a third dimension or a time axis no agent scored above 0.19. Neither biological specialization, stronger models nor detailed expert instructions closed this gap. The agents were also unreliable, with scores varying more between repeated runs of one agent than between different agents, and without ground truth a correct run could not be told from a wrong one by its process score or by the time spent. Released openly with its data and code, BIABench provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
156. Query Expansion and Key Specialization in Transformer Attention Geometry
- Authors: Vidit Gupta , Siddhesh Nadkarni , Mihik Chaudhari , Vinaya Sawant , Prachi Tawde
- URL: https://arxiv.org/abs/2609.34273
- Abstract:
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that $PR_Q - PR_K$ is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of $QK^\top$ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
157. Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
- Authors: Weijun Luo , Kelvin Luu , Xinyi Liu , Guangze Luo , Miguel Romero Calvo , Soham Dan , Daniel Yue Zhang , Ying Liu , Mohamed Elfeki
- URL: https://arxiv.org/abs/2609.34262
- Abstract:
Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless can become exploitable, making benchmark validity an ongoing maintenance problem. We introduce a process-verification framework that audits passing trajectories, distinguishes evidenced reward hacking from verifier weakness, and localizes exploitable surfaces for repair. Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations often increase with model generation but not monotonically. On SWEBench Pro V1.0, confirmed violation rates rise from 24% to 73% between Opus 4.7 and Fable 5 on matched tasks; later cohorts fall to 11% for Fable 5.1 and 0% for GPT-6 Astra. These comparisons are descriptive: configurations were not normalized, and the latest models also pass fewer exploitable tasks. Violations concentrate around a small set of recurring surfaces, especially unintended access to reference solutions through git history. Three repair case studies across two benchmarks show why blocking a recorded exploit is insufficient: the same protected information can remain accessible through another route. Therefore, we combine minimal patches with exploit replay and fresh agent evaluation, auditing new passes under the original standard. No evaluated attempt against the final patches reached the protected channel, and every post-patch pass was judged legitimate. Benchmark integrity requires ongoing maintenance: audit passing behavior, repair the enabling surface, and re-evaluate both exploit access and legitimate solvability.
158. QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models
- Authors: Bang Hu , Guowei Zhu , Changze Lv , Xiaoqing Zheng , Fengzhe Zhang , Fan Zhang , Wei Cao
- URL: https://arxiv.org/abs/2609.34259
- Abstract:
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
159. Evolving Support Priorities in Empathetic Reinforcement Learning
- Authors: Pengyu Huang , Zhiyuan Han , Wenwen Tong , Hewei Guo , Jiangnan Chen , Sirui Chen , Lewei Lu , Beier Zhu , Xun Yang
- URL: https://arxiv.org/abs/2609.34249
- Abstract:
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
160. Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
- Authors: Chidera Biringa , Lucas Yannul , Xiaowen Wang , Marco Ayala , Nicholas Yi , Alex Moyse , Nishant Manchanda , Vivek Gupta
- URL: https://arxiv.org/abs/2609.34242
- Abstract:
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
161. AdaGuard: An Adaptive Guard Model with User-defined Policies
- Authors: Yunhao Feng , Yifan Ding , Yuxiang Xie , Zheng Li , Mingrui Lao , Zeyuan Wang , Yanming Guo
- URL: https://arxiv.org/abs/2609.34241
- Abstract:
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent’s behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1–100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at this https URL
162. When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
- Authors: Rishabh Sharma , Rishika Lall
- URL: https://arxiv.org/abs/2609.34227
- Abstract:
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking’s gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
163. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures
- Authors: Dong Xu , Zhangfan Yang , Jiantao Wu , Shipeng Zhang , Zexuan Zhu , Jiangqiang Li , Jun Zhang , Junkai Ji
- URL: https://arxiv.org/abs/2609.34215
- Abstract:
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure that records which action wins without reflecting the absolute level of performance. When all actions fail, they tie at zero reward, and independent runs produce the same tied set with high probability, creating an illusion of stability that masks near-zero recovery success. We formalize this limitation through a set-path symmetry result, proving that for equal-cost Bernoulli actions the success probabilities (0.9, 0.8) and (0.2, 0.1) yield identical best-action-set distributions at every sample size. No procedure based solely on which action wins can distinguish these two regimes. We further prove that certifying exact population ties is impossible in finite time, and that the assignment of outcomes to checkpoints carries information beyond marginal outcome distributions. The pooled success probability is the missing scalar that resolves the ordinal ambiguity. Experiments on 864 frozen RecoveryBench episodes and two planning cohorts totaling 3,456 responses confirm the theoretical predictions. Agreement and held-out quality can move in opposite directions, and permuting checkpoint-to-action bindings changes 8 to 13 percent of cell-level conclusions. Based on these findings, we propose reporting four diagnostic quantities (agreement, all-zero fraction, held-out success, and pooled success) that expose this failure mode with no additional data collection.
164. GlyphBench: A Playground for Language-Model Reinforcement Learning
- Authors: Roger Creus Castanyer , Marc-Alexandre Côté , Matthew James Sargent , Augustine N. Mavor-Parker , Glen Berseth , Pablo Samuel Castro
- URL: https://arxiv.org/abs/2609.34214
- Abstract:
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench’s value as a testbed for systematic research on how language-model agents learn, interact, and generalize.
165. Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning
- Authors: Linbo Shao , Huilin He , Yating Lou , Dawei Cheng
- URL: https://arxiv.org/abs/2609.34211
- Abstract:
In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semantics, but at the cost of behavioral realism; textual descriptions for contextual reasoning remain scarce. We address this gap through a semantic enrichment framework grounded in original transaction behavior to simulate multimodal financial data. We (1) propose a multi-agent semantic enrichment framework that generates interpretable financial semantics grounded in transaction behavior through role-specialized agents and consistency refinement, and (2) newly contribute a valuable multimodal financial fraud dataset, MS-FFSD, enriched with structured semantics and textual semantics while preserving real-data-grounded transaction behavior. Furthermore, we systematically analyze the quality and utility of semantic enrichment. Results demonstrate statistical fidelity and framework generalizability, while showing that richer semantics benefit fraud modeling and context-aware LLM reasoning. Overall, this work advances multimodal financial fraud research and bridges emerging LLM and multi-agent capabilities with operational anti-fraud practice. The framework and dataset are released at this https URL .
166. PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
- Authors: Shane K.A. Dalumura Hettige , Jonas Oppenlaender
- URL: https://arxiv.org/abs/2609.34195
- Abstract:
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas through tool calls and observes the result after every turn. The canvas includes a starting shape which cannot be erased, and the agent’s goal is to incorporate this shape into the most original drawing it can produce. The task is open-ended, and the agent itself decides when the drawing is finished. The benchmark tests incremental visual planning over a short horizon and the transfer of creative ability from pretraining to multi-turn tool use. We evaluate 14 multimodal language models from small to frontier scale. Across the primary study and six sensitivity analyses, we collect 2,700 drawings and crowdsource creativity and recognizability ratings for every drawing and for 300 human reference drawings. We also present ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings (r = 0.85 on random held-out test split). Figural divergent thinking varies widely across the 14 models, and GPT-6 Astra produces the most creative drawings. Relative to the human drawings, the agent drawings score higher in creativity but lower in recognizability. We release the final drawings, per-round canvas snapshots, tool call traces, stimulus bank, benchmark harness, crowdsourced ratings (N = 72,000), and ViDrA checkpoint.
167. CASS: Contribution-Aware Structured Sparsity for Model Merging
- Authors: Yan Li , Guiping Cao , Meng Xu , Tao Jiang , Yaguang Song , Ming Tao , Yaowei Wang , Dongmei Jiang
- URL: https://arxiv.org/abs/2609.34184
- Abstract:
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters’’ and overlooking their inherent modularity. In this paper, we propose \textbf{C}ontribution-\textbf{A}ware \textbf{S}tructured \textbf{S}parsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.
168. Efficient Reasoning via Constrained Optimization in Latent Space
- Authors: Zhinan Hou , XingChen Li , Keyou You
- URL: https://arxiv.org/abs/2609.34181
- Abstract:
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Codes are available at \href{ this https URL }{ this https URL }.
169. Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen
- Authors: Xukui Qin , Youting Wang , Xinjie He , Ziyang Luo , Runxiong Wu , Yan-Syuan Chen , Zhongyao Chu
- URL: https://arxiv.org/abs/2609.34180
- Abstract:
How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study’s strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.
170. RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints
- Authors: Richard Krueger , Lucas Krause , Zach Pocquette
- URL: https://arxiv.org/abs/2609.34179
- Abstract:
Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a constrained evidence decision rather than a leaderboard choice. RAGWarrant normalizes evaluator outputs and operational telemetry, applies predeclared quality and hard-risk gates, assigns evidence-class claim ceilings, preserves negative outcomes, and emits auditable PROMOTE, BLOCK, REJECT, or INCONCLUSIVE decisions. We evaluate the framework across T2-RAGBench, MultiHop-RAG, CRAG, HotpotQA, synthetic reproduction, and bounded local generative experiments. On HotpotQA, operational savings were blocked because answer quality fell beyond the declared margin. A bounded CRAG study selected a lower-cost quality-tied policy, but related generative gains were unstable and a held-out guardrail failed closed. We claim an auditable promotion-control abstraction, not optimizer superiority, human validation, or production readiness. The tagged artifact reproduces from a fresh clone, runs as a hardened Docker job, accepts external evaluator exports, and verifies artifact integrity.
171. ReplayLens: Auditing Agents’ Use of Outcomes
- Authors: Dong Xu , Zhangfan Yang , Jiantao Wu , Shipeng Zhang , Zexuan Zhu , Jiangqiang Li , Jun Zhang , Junkai Ji
- URL: https://arxiv.org/abs/2609.34177
- Abstract:
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action’s name, or the record’s position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.
172. RoutePrism: Tracing Construction Order Effects in Agent Memory
- Authors: Dong Xu , Zhangfan Yang , Jiantao Wu , Shipeng Zhang , Zexuan Zhu , Jiangqiang Li , Jun Zhang , Junkai Ji
- URL: https://arxiv.org/abs/2609.34160
- Abstract:
Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces which sources, compiled contexts, and answers differ. Because record content, timestamps, policy, and the answer model all stay fixed, any observed difference is localized to the memory construction step. A matched four-condition intervention tests whether a record displaced by reordering actually carried task-relevant evidence: restoring that single record recovers over 60 percentage points of lost accuracy, while substituting a non-supporting record of equal length does not. We evaluate the protocol on PersonaMem-32K (63 primary queries, 29 users) and 470 LongMemEval-S questions with histories spanning 38 to 62 sessions, replicating the core intervention across five answer models. Survivor selection, defined as the choice of which record a cluster retains, drives most source-level changes, while different memory policies (compaction, bounded recency, MemoChat-style summarization, A-MEM) produce distinct failure signatures at the source, context, and metadata layers.
173. TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora
- Authors: Jiaming Tian , Liyao Li , Wentao Ye , Haobo Wang , Lihua Yu , Zujie Ren , Gang Chen , Junbo Zhao
- URL: https://arxiv.org/abs/2609.34157
- Abstract:
Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depends on surrounding schema and table context. Heterogeneous schemas, value formats, and serializations further weaken one-shot matching. We present TableSeek, a structure-preserving agentic search framework for heterogeneous table corpora. Instead of ranking tables once, an LLM agent iteratively follows sparse clues, inspects schema-preserving previews, identifies schema- and value-level mismatches, and refines its investigation. TableSeek uses cells and schemas as evidence anchors while retaining complete tables as evidence units, enabling fine-grained localization without losing the context required for interpretation and answerability checking. Without relying on retriever training or a precomputed semantic index, TableSeek produces transparent evidence-seeking trajectories and achieves competitive end-to-end performance against strong retrieval-and-reranking pipelines on heterogeneous table benchmarks. These results suggest that active, structure-preserving evidence seeking is a promising paradigm for open-domain table retrieval.
174. Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization
- Authors: Xuanqi Zhang , Ruinan Jin , Running Yang , Yuxuan Zhang , Minghui Chen , Wenlong Deng , Xiaoxiao Li
- URL: https://arxiv.org/abs/2609.34151
- Abstract:
Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution, where each recurring task type maintains a persistent tool space which is constructed from accumulated output experience. We identify three limitations of existing methods: (1) output-unaware selection: they rely primarily on tool descriptions or model priors rather than observed tool outputs; (2) statelessness across request: they select tools independently for each request without consolidating prior output experience into persistent task-specific state; (3) inference cost: they repeatedly search, rank, or reason over candidate tools for subsequent requests of the same task. We address these limitations through output-aware tool scoring, persistent task-specific tool spaces, amortized tool selection, and reusable configurations across models. We introduce LOTS (Likelihood-Only Tool Scoring), which evolves an agent’s tool space from accumulated output experience while keeping model parameters fixed. After each request, LOTS holds the model’s generated answer and estimates each tool’s contribution by measuring how much the answer likelihood changes when its observed output is removed. These contributions are aggregated within each recurring task to rank tools and update its persistent space. Across three benchmarks, LOTS improves task performance while substantially reducing tool context. More importantly, sequential experiments demonstrate that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.
175. Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?
- Authors: Tien Tran , Namho Koh , Daiki E. Matsunaga , Ayush Jain , Kee Eung Kim
- URL: https://arxiv.org/abs/2609.34139
- Abstract:
Mobile GUI agents deployed in real settings must work across different applications that support the same functionality. Most existing benchmarks test each task in only one app, so a high score can mean the agent understands the task, or only that it knows that particular app. We introduce AnyAppBench, a category-controlled live Android benchmark that evaluates cross-application generalization while keeping the user goal fixed. It spans 10 functional categories, 100 task templates, and 520 task–application pairs over 52 applications. Agents run from raw instructions and with app-independent sub-goals, and a VLM judge labels every failed run under a fixed failure taxonomy whose reliability is measured by human annotation. We find that, across 13 agents, success on the original application does not transfer reliably to new applications with the same goal. Furthermore, providing high-level sub-goal decomposition produces only small, category-dependent changes that do not close the gap, and the mix of failure types changes with the target interface. Based on those insights, we believe the AnyAppBench benchmark provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Our code, data and the leaderboard can be found at the project website this https URL .
176. You Can’t Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning
- Authors: Yian Wang , Ali Ebrahimpour-Boroojeny , Hari Sundaram , Varun Chandrasekaran
- URL: https://arxiv.org/abs/2609.34137
- Abstract:
Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75\%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the $\kappa$-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.
177. Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms
- Authors: Mingxi Zou , Wei Zhu , Zhuo Wang , Langzhang Liang , Zhiwen Tang , Yinghui Xu , Zenglin Xu
- URL: https://arxiv.org/abs/2609.34136
- Abstract:
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable local law that can be shared across interchangeable agents and adapt coordination as populations or interaction conditions change, without redefining a global organization. We introduce Waggle, a shared anonymous policy over bounded local views that jointly selects task actions, semantic communication, and local commitment updates. Repeated execution of the same law allows coordination to form, persist, and reorganize online without explicit roles or global topology. To learn this law across interchangeable agents and evolving coordination, we develop Swarm-Consistent Distillation (SCD), combining anonymous-orbit consistency with rollout-grounded prediction of the next local coordination field, with no added inference-time components. Across diverse coordination settings, the same learned law remains effective as populations and interaction budgets change, retains over 96% of substrate-specific oracle quality, and transfers without retraining; SCD further improves reorganization after counterevidence. Together, these results show that LLM-agent organization can emerge and adapt through repeated execution of a learned local law.
178. Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment
- Authors: Renxiang Wang , Jiaming Cui
- URL: https://arxiv.org/abs/2609.34135
- Abstract:
A skill bank that helps one multi-agent system may leave another’s behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4–32 agents and GPT and Qwen model ladders. Source evolution meets a joint quality, cost, model-tier, and confirmation goal in 14 of 16 settings. We then evaluate Evo2Team, which selects, adapts, and confirms source skills for the target team, alongside six frozen selectors across 28 transfer directions. Evo2Team’s target-side exploration cost is below that of evolving a new target bank in every direction, even when reused reference evaluations are charged once. Twenty of 28 held-out outcomes meet the positive-transfer criterion, including three saved diagnostic tests. Selection alone does not explain these outcomes: KNN and CORAL choose different banks in two AgentsNet directions but produce identical recorded executions. When Evo2Team changes execution, gains can reach many tasks, as in a Count-Frequency direction that improves 28 of 32 tasks over KNN. Seven positive AgentsNet outcomes save 6.1–14.6\% in deployment cost while using transferred skills on only three to six of fifteen tasks. In five earlier accepted directions, all 22 task records using transferred skills pass three fixed-graph confirmations, but four fail in recorded executions on new graphs. Graphs and model responses change together in this comparison. These results show that skill transfer must be assessed through the actions agents take, the tasks those actions reach, and the quality and cost of the final deployment.
179. StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents
- Authors: Wenle Liao , Zhao Wang , Jingchao Zhang , Jiajie Jin , Yimeng Xu , Zhicheng Dou
- URL: https://arxiv.org/abs/2609.34134
- Abstract:
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, making it difficult for agents to track which analytical artifacts remain valid over increasingly long horizons and changing dependencies. Consequently, stale artifacts may be silently inherited, propagating errors to downstream stages. To address this challenge, we propose StateGuard, an analytical-state validity management framework for long-horizon data agents. StateGuard externalizes evolving analytical progress into a state graph containing constraints, versioned variables, intermediate conclusions, and cross-state relations, treating each state as an executable, verifiable, and traceable object rather than textual memory alone. StateGuard maintains state validity through evidence-grounded verification and hierarchical intervention. To equip StateGuard with these capabilities, we first introduce Manager-Oriented Counterfactual Supervision, which constructs 3K state-centric trajectories through counterfactual runtime synthesis to fine-tune StateGuard for state maintenance, verification, and repair. We then apply Validity-Guided Policy Optimization, using runtime validity evidence to provide fine-grained learning signals for protocol correctness, state grounding, and intervention quality. Experiments on three diverse long-horizon data-analysis benchmarks show that StateGuard consistently improves data-agent performance while reducing dependency-induced downstream error propagation, demonstrating the advantages of explicit analytical-state management for reliable long-horizon data analysis.
180. From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
- Authors: Mingxi Zou , Langzhang Liang , Zhuo Wang , Yiyang Zhao , Lizhen Qu , Zenglin Xu
- URL: https://arxiv.org/abs/2609.34132
- Abstract:
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
181. GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
- Authors: Shaoqing Zhang , Kehai Chen , Xuefeng Bai , Zhuosheng Zhang , Pengfei Zhang , Yang Xiang , Min Zhang
- URL: https://arxiv.org/abs/2609.34113
- Abstract:
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at this https URL
182. SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction
- Authors: Sejin Sim , Jinsoo Bae , Seoung Bum Kim
- URL: https://arxiv.org/abs/2609.34111
- Abstract:
The rapid advancement of artificial intelligence has observed increased application in predicting vehicle interior noise levels within the automotive industry. However, the collection of labeled data for training models in this context involves significant costs. Previous studies in semi-supervised regression (SSR) have effectively mitigated the reliance on labeled data by incorporating unlabeled data. Nonetheless, these approaches often introduce a high computational cost due to the training of multiple models and data sampling. This study introduces SpecRegMatch, a novel SSR method aimed at addressing the computational cost associated with training by leveraging a single model, thus eliminating the need for multiple data samplings. SpecRegMatch integrates consistency regularization and information maximization to robustly train the model, achieved through various augmentations applied to both the embedding vectors and predicted values. Experimental results demonstrate that SpecRegMatch achieves state-of-the-art performance across various scenarios, even when using a single model. It attains a remarkable performance, as indicated by an R^2 score of 0.434. This is especially noteworthy in scenarios where labeled data is scarce. You can access the code for our proposed method at this https URL .
183. A Differentiable Optimization Framework for Registering Sequential Bounding Boxes with Point Cloud Stream
- Authors: Xuesong Li , Jinguang Tong , Jie Hong
- URL: https://arxiv.org/abs/2609.34103
- Abstract:
Refining a sequence of coarse 3D bounding boxes against a LiDAR point-cloud stream demands tracks that are geometrically accurate (high IoU) and temporally coherent (low roughness), preferably without training data. The usual recipe keeps the two concerns apart: register each frame independently, then smooth the trajectory afterwards with a Kalman~RTS or Savitzky–Golay filter. Smoothing displaces boxes from a geometric optimum and never re-optimises, so it trades accuracy for smoothness. We instead fold the temporal smoothness constraint into a training-free registration objective and solve for all poses jointly with L-BFGS. The payoff depends on how well the object is seen. On well-observed tracks it is large: within the low-roughness budget, the joint objective beats both post-hoc smoothers on paired multi-seed statistics and cuts roughness several-fold relative to frame-wise registration at matched accuracy. Treating visibility as an experimental variable exposes the limit. The advantage decays monotonically as views become one-sided, until it is indistinguishable from zero for near-edge-on objects and slightly negative under a ray-cast simulator with range-dependent density and ego motion, where the decoupled pipeline is in fact ahead at tight roughness budgets. We locate that boundary and trace it to one term: orientation alignment ties yaw to the estimated velocity and fails once that estimate is noisy. A ground-truth-free rule can choose the temporal scale and keep every track inside the roughness budget.
184. K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings
- Authors: Yunfei Bai , Enrico Chionna , Akash Amol , Kawaljit Singh KC , Joern Tinnemeyer
- URL: https://arxiv.org/abs/2609.34082
- Abstract:
Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillation (OPSD) with verifiable supervision, we construct a teacher from the model’s own best-of-N generations, certified by a process-level verifier, and rescue failed prompts by resampling under a hint that exposes the verified answer. We then perform an on-policy model update by training on verified completions with a cross-entropy inner-loss, outperforming the bounded token-wise generalized Jensen-Shannon divergence (JSD) used by on-policy distillation. Using K-OPSD, we fine-tune Qwen3-VL models on the AECV-Bench dataset. The resulting models attain the top average judge score (0.819) and combined accuracy (0.738), achieving competitive results against open-source baseline models. The recipe transfers to the out-of-domain ArchCAD dataset, where the 8B model gains most. We present the verifier suite and the continual learning and self-improving pipeline, our results provide preliminary evidence that verifier-guided self-distillation is a promising route toward more reliable machine reading of architecture drawings.
185. GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation
- Authors: Tanmoy Kanti Halder , Akash Ghosh , Arijit Roy , Sriparna Saha
- URL: https://arxiv.org/abs/2609.34079
- Abstract:
Large language models (LLMs) have demonstrated strong capabilities in biological reasoning; however, genomic disease inference remains largely dependent on memorized gene-disease associations rather than understanding biological pathways. This shortcut learning undermines robustness and generalization, and breaks down when molecular identifiers are unavailable. We present GenoMorph, a multimodal genomic reasoning framework that shifts disease prediction from associative gene-disease mapping toward pathway-grounded reasoning. GenoMorph couples a frozen DNA foundation model with question-conditioned cross-attention fusion, self-adaptive latent reasoning (LatentSp), a residual reasoning gate for iterative genomic evidence reinjection, and rejection sampling fine-tuning regularized by hierarchical optimal transport (OT). Rather than learning direct gene-disease mappings, GenoMorph aligns genomic sequence representations with latent pathway dynamics, enabling reasoning trajectories that follow molecular interactions before producing disease predictions. LatentSp dynamically allocates computation according to reasoning confidence, reducing unnecessary reasoning steps and improving inference efficiency. We further construct an anonymized benchmark from the Kyoto Encyclopedia of Genes and Genomes (KEGG), replacing every gene and molecular identifier with anonymous symbols while preserving sequences and pathway topology, thereby removing memorization shortcuts. GenoMorph raises the weighted F1 from 0.7863 (BioReason) to 0.9412, and rejection sampling fine-tuning with self-adaptive latent reasoning pushes it to 0.9725 while cutting latency nearly 60%. On the anonymized benchmark it reaches 0.9465 F1, substantially outperforming prior systems and confirming that accurate disease prediction can arise from pathway reasoning rather than memorized gene-disease associations.
186. PhysFieldBench: Can Multimodal Models Understand Physical Fields?
- Authors: Yuezhou Ma , Huikun Weng , Jialong Wu , Chenyi Zhao , Hang Zhou , Haonan Shangguan , Jianmin Wang , Mingsheng Long
- URL: https://arxiv.org/abs/2609.34072
- Abstract:
Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.
187. Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- Authors: Piyush Jha , Aishik Ghosh , Vijay Ganesh
- URL: https://arxiv.org/abs/2609.34069
- Abstract:
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.
188. Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
- Authors: Timothy DeLise , Seth Cromelin
- URL: https://arxiv.org/abs/2609.34049
- Abstract:
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry information from beyond the current window and transmit it to later states. This study presents a series of experiments using five open-weight models spanning Qwen, Llama, Mistral, and Muse Glimmer. We investigate whether information originating outside the immediate context window can persist through a rolling KV cache and remain useful for retrieval. Initial results show that retaining previously computed states improves retrieval across the models tested compared with recomputing the final fixed window from raw tokens. We then measure how far this effect extends and find that Muse Glimmer and Mistral 7B show the strongest \emph{latent information relay}: they can recover information even after the relevant source tokens have left the cache. Both models incorporate sliding-window attention in their published architectures, an association that motivates testing whether training with sliding windows promotes more reliable information retention.
189. Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
- Authors: Erfan D. Dehkalani , Seetha Shankaran , Abbot R. Laptook , C. Michael Cotten , P. Ellen Grant , Yangming Ou
- URL: https://arxiv.org/abs/2609.34039
- Abstract:
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med’s configuration and scalability, and the Invocation Agent’s accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med’s configuration and scalability and the Invocation Agent’s accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
190. Jev in Medicine: A Benchmark Evaluation. Preliminary Results
- Authors: Alfredo Madrid-García , Beatriz Merino-Barbancho
- URL: https://arxiv.org/abs/2609.34024
- Abstract:
Jev is a non-generative “System One” model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev’s accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev’s probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev’s probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was “I don’t know or cannot answer”, Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
191. A Computer Vision Approach to Visual Fraud Detection in Phishing Websites Using YOLOv8
- Authors: Basil Sajid Shaikh , Hajar Homayouni
- URL: https://arxiv.org/abs/2609.34015
- Abstract:
Phishing remains one of the most common vectors for financial and identity fraud, and most detection systems still rely on inspecting a page’s URL, HTML markup, or domain registration history. These signals are easy for an attacker to rotate or obfuscate, and they say very little about what actually convinces a victim to hand over a password or a card number: the way the page looks. This paper describes a visual, image-based approach to phishing detection that treats a rendered webpage the same way a human eye would, as a picture that either matches a trusted brand or doesn’t. A YOLOv8 convolutional neural network was trained to classify full-page website screenshots as phishing or legitimate based on layout, logo placement, color scheme, and login-form structure, rather than on text extracted from the page. The system reached 92% classification accuracy on a held-out test set, processed a single screenshot in roughly 100 milliseconds, and, after a round of data augmentation aimed specifically at lighting, compression, and scaling variation, cut the false-positive rate by 11% relative to the pre-augmentation baseline. The paper walks through the dataset construction, the augmentation strategy, the model architecture and training setup, and the resulting performance, and closes with a discussion of where this kind of visual detector fits alongside, rather than instead of, existing URL- and content-based defenses.
192. EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
- Authors: Andre R Goncalves , Vincent Liu , Priyadip Ray
- URL: https://arxiv.org/abs/2609.34007
- Abstract:
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model’s embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event’s clinical description from a biomedical language model trained on clinical ontologies, mapped into the model’s input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients’ records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1–0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.
193. Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents
- Authors: Kaiwen Luo , Ming Gao
- URL: https://arxiv.org/abs/2609.33955
- Abstract:
Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an integrated methodology spanning evaluation specification, modular LLM judges, intent-preserving user simulation, and human-in-the-loop governance. The specification defines conversation-level end states and actionable failure ownership. Atomic judges share versioned evidence and feed an explicit aggregation graph. The simulator is released only after task-preservation and stability checks. Independent human audits estimate measurement fidelity, renew tiered reference sets, and route disagreements to label correction, guideline revision, or judge improvement. Production studies show that system-level fidelity improved across repeated audits, that human reviewers and automated judges improved together under the shared feedback loop, and that their combined workflow had the strongest descriptive performance in both reported task-completion settings. Because the studies are observational and the human reference itself required revision, these findings demonstrate operational usefulness rather than causal or universal superiority. The contribution is a practical framework for making multi-turn agent measurement reliable, actionable, and maintainable as the evaluated system and its evidence evolve.
194. HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents
- Authors: Tingsong Xiao , Nithish Balachandar Moudhgalya , Chandrayee Basu , Lichao Wang , Luyang Kong , Benjamin Z. Yao , Zhe Jiang , Jie Hao
- URL: https://arxiv.org/abs/2609.33920
- Abstract:
Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment interaction make search costly. Efficient search therefore requires effective reuse of trajectory feedback. Standard MCTS maintains prefix-specific statistics, without explicitly accumulating outcomes for decision groups that recur across different paths. To fill this gap, we propose HyperMCTS, a training-free method that augments an ordered MCTS tree with a cross-trajectory hypergraph. Hyperedges represent groups of canonical decisions and accumulate their observed returns within the current task. Our hypergraph-guided HyperUCT selection rule aggregates evidence from overlapping hyperedges into an action prior, allowing outcomes collected under one prefix to inform selection under another while preserving execution histories in the tree. On DeepPlanning, HyperMCTS improves average planning accuracy by 2.3–7.3 percentage points over the strongest baseline for each of three backbone models. It enables Qwen3.6-27B to outperform Claude Opus 4.6 (max) on Shopping Planning, while achieving higher accuracy with fewer LLM calls and output tokens than the evaluated MCTS-based baselines. SealQA experiments further demonstrate improvements in question answering.
195. When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
- Authors: Zhihao Zhang , Chao Wang , Rujia Li , Qingze Wang , Xiaoyan Sun , Jun Dai
- URL: https://arxiv.org/abs/2609.33910
- Abstract:
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can outlive the context that originally justified the approval, creating residual authority reusable without renewed consent. We expose this failure mode through a longitudinal attack that starts from a target security-sensitive action, identifies the authority required to execute it, induces benign interactions that legitimately obtain that authority, and later replays the residual authority during adversarial execution. Across controlled and live settings, we demonstrate that residual-authority replay arises in practice and substantially increases the success of prompt-injection and context-rebinding attacks. We evaluate 508 AgentDojo attack cases across six LLM families using production-derived authorization semantics. With residual authority, attack success rate (ASR) increases by up to 35.1 percentage points compared with a fresh authorization state. In live context-rebinding attacks on 55 Terminal-Bench cases across three real-world production coding agents, residual-authority replay increases ASR by 24.9 percentage points on average. These findings expose a fundamental mismatch between persistent authorization and the contextual nature of user consent in long-lived LLM agents.
196. Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
- Authors: Donghao Huang , Jinling Pei , Zhaoxia Wang
- URL: https://arxiv.org/abs/2609.33878
- Abstract:
Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.
197. When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
- Authors: Janvijay Singh , Vaishnavi Shrivastava , Dilek Hakkani-Tur , Ece Kamar , Asli Celikyilmaz
- URL: https://arxiv.org/abs/2609.33870
- Abstract:
LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.
198. R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution
- Authors: Mingda Zhang , Qiang Huang , Yanjin Li , Zijia Wang , Qika Lin , Xiaoying Tang , Tiesunlong Shen
- URL: https://arxiv.org/abs/2609.33867
- Abstract:
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at this https URL .
199. How code helps different tasks? A decompositional lens on LLM post-training
- Authors: Zheng Yu , Yiwei Li , Yishen Chen , Xiang Li , Jiale Han , Benyou Wang , Jingbang Chen
- URL: https://arxiv.org/abs/2609.33845
- Abstract:
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decompositional lens for studying these effects in LLM post-training. We first decompose an execution-verified code corpus into interpretable categories based on the computational patterns of its solutions. Through controlled fine-tuning experiments, we compare individual categories with a balanced mixture across instruction-tuned models on question answering, mathematics, and code generation. The resulting response maps reveal recurring gains in average question-answering performance, while the same category can improve one model or task and degrade another. The best-performing category also varies with the starting model and target task. We then compose compact mixtures guided by these results and examine whether benefits observed in individual categories persist under joint training. On selected model–task pairs, mixtures whose constituents each improve the target task outperform both their best constituent and full-corpus training while using roughly 10–15\% of the full corpus. These exploratory findings illustrate a \emph{less is more} pattern and highlight how the value of code data in post training depends on which categories are combined for which model and task.
200. Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation
- Authors: Gowthamkumar Nandakishore
- URL: https://arxiv.org/abs/2609.33843
- Abstract:
The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability bin’s accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, $0.214$. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card’s headline accuracy ($0.767$ vs. $0.766$). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at $q=0.05$ (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model’s publisher, the dataset’s publisher, or TypeSafe.
201. Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory
- Authors: Jayant Parashar , Eugene F. Douglass , William C. Bastian , Suchendra M. Bhandarkar
- URL: https://arxiv.org/abs/2609.33822
- Abstract:
An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model. Its organizing overhypothesis is that tasks of a shared kind may exhibit recurring failures whose remedies transfer within that kind. Vestrum expresses failures as recognizable classes, proposes changes across verification, retrieval, decomposition, and knowledge synthesis, and screens their scope before evaluating them as a bundle. A persistent lessons file informs subsequent proposals. Across five settings and two baseline harnesses, the frozen harnesses improve held-out performance: UltraHorizon rises from 47.6 to 59.8 over GAM, Terminal-Bench 4 Hard from 63.7% to 70.3% of checks passed over Claude Code on eight held-out tasks at 1.03x test cost, and cell-type annotation agreement from 67.5% to 77.8% on held-out sections of one slide, alongside gains on LoCoMo and AMA-Bench. Across our searches, verification grounded in evidence helped both intermediate steps and final answers, at lower cost at intermediate steps, while critics asked to rebuild finished answers broke more than they repaired. On the three memory benchmarks, Vestrum also scores above the evaluated GEPA configurations in every paired evaluation.
202. Dual-Vocabulary Language Model for Cross-Tokenizer Distillation
- Authors: Kedi Chen , Chen Lin , Yutao Sun , Wei Zhang
- URL: https://arxiv.org/abs/2609.33816
- Abstract:
On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher’s LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student’s input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
203. Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative
- Authors: Vicent Ribas , Anna Oliveras Tous
- URL: https://arxiv.org/abs/2609.33786
- Abstract:
A diffusion model can predict a follow-up medical scan from a baseline, but a clinician needs a per-voxel map of where that prediction can be trusted. Many such maps approximate the diagonal of the Tweedie posterior covariance, and are evaluated against another approximation of it, so whether the estimator or the target limits them is unclear. We compute the exact diagonal on six checkpoints across fourteen model-corpus conditions. Hutchinson at M=200 tracks it at rank agreement of at least 0.92 everywhere, yet in four of the fourteen the exact diagonal is anti-correlated with the denoising error, reaching -0.13, so a faithful estimator reproduces that reversal. All four are real-image conditions; on the models’ own samples the reversal does not appear, so evaluating on generated samples flatters this family. What limits these maps is the target, not the estimator. We then introduce Tweedie Probe-Tangent (T-PT), a gradient-free residual probe that corrupts one model-supported prediction repeatedly and measures the voxel-wise variance of the denoiser’s response. T-PT reads a different functional of the same Jacobian, and its exact second-order form ranks with the diagonal wherever the diagonal reverses; at thirty probes it returns a map too unstable to reproduce that ranking, while Hutchinson at M=5 already reproduces it, so T-PT there is not evidence against the reversal. We offer it as an instrument, not a better approximation. On brain MRI at full resolution, where every Jacobian-based estimator we test runs out of memory, T-PT leads a twenty-chain Monte-Carlo ensemble on five of eight endpoints inside tissue and trails it on none, at 16x fewer network evaluations; over the whole volume the ensemble leads, and fifty chains close the tissue gap. On lung CT the ensemble is ahead throughout. Both lose most of their discrimination where the change is, which remains open.
204. Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong
- Authors: Megan Diehl , Ser-Nam Lim
- URL: https://arxiv.org/abs/2609.33778
- Abstract:
Modern multi-hop LLM agents are equipped with built-in mechanisms to detect errors in intermediate reasoning steps. Such errors trigger corrective actions from these agents, which mostly follow the paradigm of retrying the steps or the reasoning trajectories. Not only are these retries expensive, we present in this paper that they are also potentially unnecessary. To this end, we introduce Evidence-Inference Reconstruction (EIR), which uses structured state to guide one retrieval trajectory, accumulating source evidence in the process. We show that as long as the relevant evidence has been collected, EIR is capable of generating the correct answer in a single final model call even if erroneous evidence has been mixed in due to incorrect intermediate reasoning steps. In one evaluation, using Haiku 4.5 and GPT-4.1 Mini, we evaluate EIR on matched 1,000-question subsets of HotpotQA, 2WikiMultiHopQA, and MuSiQue, showing that EIR improves Answer F1, the overlap between the model’s and the correct answer, over the baseline by 8.3–32.8 points, Agentic SSR by 10.6–29.1 points, and Reflexion by 1.1–15.9 points. Additionally, we show that EIR averages 4.85 total model calls per question, compared with 35.29 for Agentic SSR and 12.41 for Reflexion. Together, these results corroborate EIR’s central premise: separating evidence retrieval from the final answer model call can improve answer accuracy while utilizing substantially less computation.
205. Learning Strategies to Break Judges
- Authors: Guruprerana Shabadi , Aaditya Naik , Rajeev Alur , Mayur Naik
- URL: https://arxiv.org/abs/2609.33773
- Abstract:
As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important question: how can we trust the judge? In this work, we propose an agent-guided method to find weaknesses of agentic judges that expose interpretable failure mechanisms. Our method focuses on mathematical reasoning and proceeds in two stages: first, we deploy adversarial agents to mutate a set of sound proofs by introducing errors, attempting to misguide judges—in other words, injecting errors that judges are unable to catch. Then, we distill these attempts into a small set of mutation strategies which allow us to analyze the failure modes of the judges. To ensure that these strategies are not overfit to the initial set of proofs, we evaluate them by applying the mutation strategies to a held-out set of proofs and querying the same judge. We deploy our method on GPT-5.6-sol and Claude Opus 5, paired with their agent orchestrators, Codex and Claude Code, respectively. These are used both as mutators to introduce errors and as judges to evaluate correctness of mathematical reasoning. We find that across all the agentic judges, we are able to distill mutation strategies that consistently bypass their evaluations, thereby enabling us to ascertain actionable failure modes. Our analysis also reveals that judge reliability degrades at the frontier: errors in Olympiad-level proofs or graduate-level mathematical texts are detected more consistently, whereas flaws in research-level manuscripts are more likely to escape detection.
206. Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents
- Authors: Weiyi Xu , Xiaowen Yang , Wen Da , Hang Xu , Canwei Li , Hongjie You , Pusen Dong , Yucheng Zeng , Zhaokai Luo , Mu Chuan
- URL: https://arxiv.org/abs/2609.33772
- Abstract:
Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.
207. HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration
- Authors: Eliott Jacopin , Éric Jacopin , Koichi Takahashi
- URL: https://arxiv.org/abs/2609.33731
- Abstract:
The Model Context Protocol (MCP) isolates servers by design: only the host can orchestrate cross-server workflows. When the host is a large language model, the resulting orchestrations are non-deterministic, non-reproducible, and pay one inference round-trip per tool call. We present a coordination architecture in which a Hierarchical Task Network (HTN) planner generates a verifiable cross-server plan once, and a runtime middleware executes it deterministically across multiple MCP servers, binding cross-action data dependencies via a template mechanism (\verb ${context.X} ) substituted at execution time. The architecture mirrors MCP’s isolation constraint: each compound task decomposes into server-local primitive actions, and inter-server data flow is bound at execution time via JSON-path output extractors. We instantiate the architecture on five HTN domains spanning laboratory robotics, bioinformatics and multiscale modelling, and demonstrate end-to-end execution from a browser-based plan controller against eight live third-party MCP servers querying real biological databases.
208. Robust Biomolecular Complex Design Across Protein Conformational Landscapes
- Authors: Qingyuan Zeng , Zongqi Xu , Anglin Liu , Ziqi Gong , Pengxiang Cai , Zixin Guan , Yunan Chen , Sen Gao , Min Zhou , Jintai Chen
- URL: https://arxiv.org/abs/2609.33726
- Abstract:
Proteins populate conformational ensembles, yet structure-based biomolecular design typically optimizes candidates against a single target conformation. Consequently, a candidate that fits one state can lose favorable interactions or develop steric clashes when the target adopts another. We introduce FlexEvo, a model-agnostic evolutionary framework that adapts candidates once at inference time from a single target conformation to improve compatibility with alternative natural conformations unseen during adaptation, without retraining the source model or requiring a conformational ensemble. FlexEvo casts cross-state adaptation as geometry-constrained bi-objective optimization, balancing preservation of input-state interactions against robustness to plausible conformational perturbations. To limit the search space and reduce invalid structural edits, geometry-derived FlexBoxes define protected anchor regions, adaptable regions for local exploration, and forbidden regions for clash avoidance. A unified all-atom representation supports topology-preserving adaptation across diverse binder categories, while Pareto selection preserves nondominated candidates across the two objectives. We evaluate FlexEvo across multiple generation baselines and nine representative binder categories spanning diverse molecular sizes and structural topologies. FlexEvo reduces the category-balanced mean relative performance degradation from 47.8% to 4.4%, while adding only 1.4–3.1 minutes of adaptation per sample. These results establish single-state inference-time adaptation as a practical route toward robust biomolecular complex design across protein conformational landscapes.
209. Self-Designed Evaluators and Warm Memory for Long-Horizon Agents
- Authors: Saeid Asgari , Emre Kiciman , Leonardo de Oliveira Nunes , Ranveer Chandra
- URL: https://arxiv.org/abs/2609.33717
- Abstract:
A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent’s own base model, given only the world’s public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite’s evaluator gives a small, consistent gain, and using them to warm up ACE’s memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.
210. BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation
- Authors: Anglin Liu , Yanlin Wu , Ruichao Chen , Yuting Zhang , Qingyuan Zeng , Pengxiang Cai , Ziqi Gong , Muchen Li , Jintai Chen
- URL: https://arxiv.org/abs/2609.33713
- Abstract:
Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model’s relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM’s own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model’s preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.
211. Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse
- Authors: Futa Waseda , Shuhei Kurita , Isao Echizen
- URL: https://arxiv.org/abs/2609.33707
- Abstract:
Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.
212. SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications
- Authors: Feilian Huang (Independent Researcher)
- URL: https://arxiv.org/abs/2609.33699
- Abstract:
Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification and failed to write the code. We present SpecRead, a benchmark that isolates specification comprehension from generation ability. SpecRead v2.1 contains 385 questions over 10 open-source OpenTitan IP blocks: exact retrieval, cross-section reasoning, contradiction detection in mutated specifications, and spec-RTL consistency checking, plus 82 controls (41 distractor, 41 consistent-RTL). Type-4 items are built from real RTL mutations; we retain only mutations that Icarus Verilog simulation shows to change observable behavior. A with-spec vs. without-spec ablation suggests the questions require the excerpt, not training recall alone (without-spec accuracy 3/20 on the t1/t2 subset), though memorization of the source text may still help spot mutations. As an initial characterization with a small model, Ministral-3B scores 33.2% overall (128/385; macro average 39.0%): 55.2% on retrieval, 51.7% on cross-section reasoning. On the two contradiction-focused types, the verdict-plus-location measure gives 48.0% (t3) and 63.3% (t4), with a 51.2% false-positive rate on distractors and 100% on consistent-RTL controls. Layered scoring shows the model locates contradictions well (78.9-81.6% location accuracy) but scores lower on their category (43.9-49.7%). A structured “rule-table” prompting intervention lowers accuracy on every question type except t2 (tied). SpecRead is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring. The benchmark is regenerable for type-3 items via mutation injection, and built exclusively from public sources.
213. One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion
- Authors: Yulin Yuan , Ying Zhang , Xiangming Meng
- URL: https://arxiv.org/abs/2609.33698
- Abstract:
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF’s throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.
214. TopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology Optimization
- Authors: Bin Lou , Yuxuan Cheng , Huaizhi Zong , Junhui Zhang , Bing Xu
- URL: https://arxiv.org/abs/2609.33688
- Abstract:
Deep learning has emerged as an efficient alternative for predicting high-performance material distributions in topology optimization. Existing methods struggle to accurately capture load-transfer information, limiting out-of-distribution generalization, while their model architectures often incur high computational costs. To address these challenges, this paper proposes TopoMamba, a topology prediction framework incorporating a load-support relation-guided multi-directional state-space model. Coupling physical fields with load-support relations enables more effective modeling of mechanical dependencies. A load-support relation-guided spatially adaptive fusion mechanism dynamically adjusts multi-directional scan features according to spatial conditions. Mamba is coupled with the solid isotropic material with penalty method to enhance structural mechanical performance while maintaining computational efficiency. Results on two-dimensional topology optimization benchmarks demonstrate that TopoMamba achieves superior topology prediction accuracy, out-of-distribution generalization, and computational efficiency over state-of-the-art models. The proposed load-support physics-guided framework enables efficient optimization of more complex structural systems.
215. SWE-Game: Can Coding Agents Build the Games We Want?
- Authors: Xiaoyu Chen , Lai Wei , Jin Wang , Xiangyu Zou , Ruochen Fan , Enze Luo , Mingzhe Yao , Jiahui Zhu , Yuhua Wen , Linghe Kong , Weiran Huang
- URL: https://arxiv.org/abs/2609.33678
- Abstract:
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.
216. Auditing Agent Actions through Query-Conditioned Attribution
- Authors: Yifan Liu , Praveen Venkateswaran , Abdulhamid Adebayo , Dong Wang
- URL: https://arxiv.org/abs/2609.33676
- Abstract:
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate $\textit{query-conditioned agent action attribution}, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with $A^3Bench$, a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9\% and evidence MAP by 42.1\% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5\% vs.\ 60.4\%) while reducing empirical deployment latency by 29.9\% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.
217. RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games
- Authors: Miaobo Hu , Shuhao Hu , Xiaobo Guo , Xin Wang , Bokun Wang , Peng Zhang , Daren Zha , Jun Xiao
- URL: https://arxiv.org/abs/2609.33669
- Abstract:
Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group’s law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least $1-\zeta_{risk}-\zeta_{util}$. The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects $\alpha=0.08$, raising the weak-response proxy from 4.2082 to 4.2889 with $0/12$ held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility $4.3659\pm0.0177$ and violation rate $0.0178\pm0.0057$.
218. CompoWorld: Compositional Environment Scaling for General Agents
- Authors: Xiao-Wen Yang , Weiyi Xu , Wen Da , Hang Xu , Canwei Li , Hong-Jie You , Pusen Dong , Yucheng Zeng , Zhaokai Luo , Yu-Feng Li , Yao Hu , Mu Chuan
- URL: https://arxiv.org/abs/2609.33665
- Abstract:
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
219. Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification
- Authors: Miaobo Hu , Shuhao Hu , Xiaobo Guo , Xin Wang , Bokun Wang , Rui Chen , Daren Zha , Jun Xiao
- URL: https://arxiv.org/abs/2609.33662
- Abstract:
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk $\rho=0.08$, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost $1.16\times$. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval $[-0.0364,-0.0157]$. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.
220. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents
- Authors: Tianzhuo Yang , Zirui Mi , Yantao Huang , Guoxi Zhang , Jiawei Chen , Yaodong Yang , Jingwei Yi
- URL: https://arxiv.org/abs/2609.33658
- Abstract:
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5\% of routine-looking unauthorized actions yet completes only 28.7\% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2\% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4\% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
221. Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
- Authors: Keliang Li , Heng Wang , Chen Hu , Daxin Jiang , Hong Chang , Shiguang Shan
- URL: https://arxiv.org/abs/2609.33646
- Abstract:
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation. We introduce Probe to Act (P2A), an active probing framework for the browser-agent loop that moves this alignment into decision time. P2A addresses an asymmetric bridge between symbolic DOM hypotheses and screenshot layout by rendering on-demand symbolic DOM structure back into pixels. Before committing a state-changing browser operation, the agent can issue lightweight probes to translate DOM handles into pixel evidence, map screen regions back to DOM candidates, register visual-only targets, and commit verified notes. These interleaved processes naturally produce evidence-based memory: only probed, acted-on, or explicitly committed observations are kept across steps, preserving only decision-critical evidence in long-horizon contexts. P2A can be used as a prompting strategy for proprietary models under the standard DOM+SoM interface, and can be distilled into open-weight models through cold-start synthesis and self-bootstrapped SFT. Across three browser-use benchmarks, P2A shows clear gains on task success rate for both proprietary and fine-tuned models; on VisualWebArena, for example, it improves Gemini-3-Pro from 54.1% to 61.2% and Qwen3-VL-8B from 24.6% to 32.9%, while matching the costly full-observation history ($\sim$3$\times$) at only $\sim$1.2$\times$ the peak retained input context of action-only history.
222. Scalable and Data-Driven Decision Support in the Maintenance, Repair, and Overhaul Process
- Authors: Houkun Zhu , Helena Ebel , Dominik Scheinert , Florian Schmidt , Jens Altenkirch , Odej Kao
- URL: https://arxiv.org/abs/2609.33641
- Abstract:
Several businesses apply maintenance, repair, and overhaul (MRO) principles to the life-cycle of their existing products. In cases like casted gas turbine component Product Lifecycle Management (PLM), repairing components in frequent intervals can extend the lifetime expectation of the product, provide higher cost efficiency compared to newly produced components, and even improve the part design during the repair cycle. Another aspect of repair concerns sustainability, as products often contain rare materials. The emissions produced by the repair process are usually smaller than mining materials and casting new components. To optimize the repair process further, we propose the Smart Expert System (SES), which assists engineering experts with machine learning-based decision support throughout the repair process. We elaborate on its IT architecture and present machine learning models employed for representative MRO use cases. The SES is evaluated using actual industry data from a leading gas turbine company and demonstrably fulfills formulated requirements concerning the suitability of the overall decision support and the stability of the enclosing IT architecture.
223. Trajectory Unlearning on LLM-based Agents
- Authors: Yingdan Shi , Ren Wang
- URL: https://arxiv.org/abs/2609.33639
- Abstract:
Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an agent knows, an agent should not reproduce undesired behaviors through its action trajectories. In this work, we introduce trajectory-level unlearning, a new problem formulation that targets the removal of specific action trajectories in long-horizon agentic tasks, rather than factual knowledge. We identify two fundamental challenges that distinguish trajectory unlearning from knowledge unlearning: (1) our unlearning target is what the agent \emph{does}, not what it \emph{says}; and (2) trajectories are sequentially dependent action sequences that cannot be decomposed into isolated prompt-response pairs without losing inter-step structure. To address these challenges, we propose Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolates the normalization statistics, yielding a stable and bounded unlearning signal that does not corrupt gradient updates for normal task trajectories. We construct trajectory unlearning benchmarks from two application scenarios, household tasks (ALFWorld) and online shopping (WebShop), and design three complementary metrics for evaluating forgetting quality and model utility. Experiments on ALFWorld and WebShop demonstrate that GiRPO effectively unlearns target trajectories while preserving task success rates, outperforming existing knowledge-unlearning baselines on both forgetting quality and task utility.
224. ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
- Authors: Shengbin Yue , Hongru Wang , Siyuan Wang , Xiaoxin Chen , Wei Chen , Zhongyu Wei
- URL: https://arxiv.org/abs/2609.33618
- Abstract:
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration $\rightleftharpoons$ Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.
225. EAT: Expert Account Tracker for Efficient MoE Inference
- Authors: Yuexian Li , Yifei Yang , Zouying Cao , Hai Zhao
- URL: https://arxiv.org/abs/2609.33614
- Abstract:
Mixture-of-Experts (MoE) models have emerged as a revolutionary method to scale Transformer models. However, traditional MoE architecture still suffers from inefficiency since a large number of experts are unnecessarily activated. Existing approaches for reducing the number of activated experts often overlook the historical performance of each expert. In this paper, we propose EAT, a novel method called Expert Account Tracker (EAT), which utilizes history-awareness metrics and adaptive thresholding to dynamically select the most important experts, thereby reducing the activated expert number while effectively maintaining the model performance. Experiments show that EAT outperforms the existing baseline Top-P method across multiple models and datasets, achieving over 25% an average reduction compared to the vanilla method in the number of activated experts and performing better token generation speed compared to the baseline. Furthermore, the performance of pruned models can be efficiently recovered via OPD using only 9K data. Additionally, through ablation studies, we find that excessively reducing the number of activated experts can significantly harm model performance, and the importance of experts varies across layers, with higher-level experts being generally more critical.
226. Supervision Recovery for Time Series Anomaly Detection via Context-Anchored Pairing
- Authors: Yifei Gao , Tian Lan , Yimeng Lu , Xuming An , Meng Wang , Wenjun He , Yijie Li , Chen Zhang
- URL: https://arxiv.org/abs/2609.33610
- Abstract:
Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised objectives or surrogate abnormal patterns, providing limited supervision for context-dependent normal–anomalous distinctions. We propose Context-Anchored Pair Supervision (CAPS), a supervision-recovery framework for TSAD. CAPS views ideal anomaly supervision as a matched comparison between normal and anomalous outcomes under the same temporal context, and seeks to recover such supervision without target-domain anomaly labels. Using simulated normal–anomalous pairs, CAPS learns structure and anomaly-semantic representations through reconstruction, background consistency, and within-pair counterfactual recombination. The resulting anomaly representations form a continuous semantic space with coarse modes and induce a sampleable multimodal prior. CAPS conditionally realizes sampled semantics as residual-form effects on target reference trajectories. The resulting context-anchored normal–anomalous counterparts provide temporal supervision for discriminative detector learning. Experiments on nine datasets show that CAPS achieves the strongest aggregate performance across all four evaluation metrics among the compared methods, while complementary ablations and transfer analyses support the roles of context anchoring, semantic disentanglement, and conditional realization.
227. JustQuant: You Don’t Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
- Authors: Kaicheng Yang , Kaisen Yang , Chunyu Liu , Xianglong Yan , Haotong Qin , Junyi Wu , Tianao Zhang , Xun Zhang , Shaoqiu Zhang , Youbang Sun , Yulun Zhang
- URL: https://arxiv.org/abs/2609.33601
- Abstract:
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-training quantization (PTQ) and quantization-aware training (QAT) methods have made progress in 4-bit activation quantization by introducing smoothing, SVD branches, rotations, mixed precision, or advanced formats such as NVFP4. These additional operators and data types impose demanding requirements on inference engines and hardware, limiting the broad adoption of low-precision models. Can quantization be achieved using only plain low-bit operators? To answer this question, we propose JustQuant, a simple yet effective framework that moves the complexity of low-bit quantization from deployment-time operators into the training process. We first revisit model quantization from the perspective of knowledge distillation and show that a key reason existing PTQ and QAT methods fail is that they typically exploit supervision at only a single level. We then introduce Theseus QAD, a quantization-aware distillation method that progressively applies multi-level supervision, analogous to the gradual replacement process in the Ship of Theseus. Extensive experiments on DiT and diffusion large language models show two distinct regimes. For smaller models, Theseus QAD can serve as a lightweight warm-up stage that substantially improves subsequent QAT with plain operators, while naive QAD may collapse in the same setting. For larger models, Theseus QAD provides a stronger distillation training path than ordinary QAD. Across both regimes, JustQuant improves low-bit quantization quality while avoiding the complex operators required by many existing PTQ methods.
228. OpenFC: Learning Verification Policies towards Open-Search Fact Checking
- Authors: Xinming Wang , Kaixiang Qiu , Yansong Lin , Chunji Lv , Yi Chen , Boran Wang , Hong-Ming Yang , Xu-Yao Zhang
- URL: https://arxiv.org/abs/2609.33579
- Abstract:
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39\% average accuracy and 63.30\% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.
229. Dr. Free: You Don’t Need Difficulty Rewards for Self-Evolving Search Agents
- Authors: Zhipeng Qian , Zihan Liang , Yufei Ma , Jie Ma , Ben Chen , Huangyu Dai , Lingtao Mao , Xinyu Sun , Tong zhao , Xuxin Zhang , Qingpeng Cai , Peng Jiang , Qibin Hou
- URL: https://arxiv.org/abs/2609.33565
- Abstract:
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over $7\times$. Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.
230. OSCC: Certified Observation-Safe Coupling Optimization for Gradient-Noise Control in Imperfect-Information Learning
- Authors: Miaobo Hu , Shuhao Hu , Xiaobo Guo , Xin Wang , Bokun Wang , Rui Chen , Daren Zha , Jun Xiao
- URL: https://arxiv.org/abs/2609.33543
- Abstract:
Coupled rollouts can reduce the noise of counterfactual action comparisons, but two issues prevent standard common-random-number constructions from serving as a general learning primitive in imperfect-information environments. First, an invalid coupling may expose hidden state, synchronize endogenous policy randomness, or misalign chance events after counterfactual histories diverge. Second, in multi-action policy optimization, lower return-contrast variance is not by itself the relevant objective: the optimizer depends on the return covariance matrix after projection through the local policy-gradient geometry. We introduce observation-safe counterfactual coupling (OSCC), a framework that defines an admissible class through marginal preservation, information-state safety, branch-local policy randomness, semantic event alignment, and trace-before-oracle replay. We derive a gradient-aware coupling criterion showing that, for marginal-preserving couplings, policy-gradient noise changes are determined by policy-Jacobian-weighted off-diagonal return covariance. This motivates OSCC-Select, a calibration-only selector that chooses among independent, root-only, continuation-only, and fully coupled rollouts using separate safety and gain certificates. Its gain target combines projected gradient noise with measured physical sampling cost and falls back to independent sampling whenever a simultaneous lower confidence bound does not certify improvement. On 100,000 fixed-root Leduc comparisons, the fully coupled CP-GRPO instantiation reduces return-contrast variance from 41.1158 to 18.1441, a 55.87% reduction, while preserving the declared branch marginals. With three actions, OSCC-Select chooses continuation coupling and attains gradient-noise trace 0.0783 versus 0.0917 for return-variance selection. Increasing calibration from 64 to 2,048 groups raises certification from 0.327 to 0.995.
231. Reasoning on the Simplex: Geometric Fixed-Point Models
- Authors: Talgat Daulbaev , Ilya Glazkov , Maxim Rakhuba , Ivan Oseledets
- URL: https://arxiv.org/abs/2609.33540
- Abstract:
Looped reasoners spend test-time compute by iterating a weight-tied map, but a small residual does not mean the state is a fixed point when that map lives in unconstrained latent space. We propose Geometric Fixed-Point Reasoning (GFPR), in which the iterated state is the prediction itself: a field of categorical beliefs on a product of simplices, whose argmax is the answer at every step. Because the state is a belief, task structure can be imposed through compact convex relaxations, either as structured readouts or directly in the recurrent state; in the latter case the update remains a continuous self-map, so a fixed point exists for any parameters. At about 7M parameters, GFPR reaches 95.1% exact match on Sudoku-Extreme, 92.0% on Maze-Hard, and 100% sequence accuracy on S_5 length 128, above the published FPRM numbers at the same scale. The same update also trains a 201M language model on FineWeb-Edu in which each site is a distribution over the vocabulary; with 24 Picard steps it is above GPT-2 small on four zero-shot multiple-choice tasks and above GPT-2 medium on ARC-Easy.
232. EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research
- Authors: Siyuan Li , Jiangfeng Zhang , Rui Yao , Weihua Qiu , Mingyang Xu , Zixuan Yuan
- URL: https://arxiv.org/abs/2609.33524
- Abstract:
Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
233. PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment
- Authors: Xiaoda Wang , Minxiao Wang , Maxwell A Xu , Patrick Langer , Kaiqiao Han , Defu Cao , Xiao Luo , Yuzhe Yang , Yan Liu , Xiao Hu , Yizhou Sun , Wei Wang , Carl Yang
- URL: https://arxiv.org/abs/2609.33516
- Abstract:
Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.
234. What Happens During Autonomous Deep Research After the User Steps Away?
- Authors: Yimin Liu , Yijia Zhang , Yanmin Li , Tangwen Luo , Yuze Li , Ziling Yao , Zhi Yang
- URL: https://arxiv.org/abs/2609.33509
- Abstract:
In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommendations. We introduce DRaligned, a counterfactual behavioral evaluation framework built on PDR-Bench. By varying one task-relevant user factor while keeping the remaining context fixed, we compare acquisition requests, working drafts, and final reports. Source-grounded extraction, blinded local judgments, and deterministic aggregation yield coarse directional measurements while leaving ambiguous cases unresolved. Our experiments show that strong user-specific delivery can emerge from a largely shared research process: agents investigate similar broad questions but allocate requests differently, and final recommendations distinguish user conditions more clearly than explicit requests do. Reports can also integrate user factors that were not jointly visible during acquisition. In readable draft-to-report comparisons, recommendations often retain their coarse user-specific direction despite substantial rewriting. Final directional differences recur across tested agent models, execution harnesses, and evaluator models, even as execution paths vary. These findings describe how initial user information shapes autonomous research and clarify the relationship between the process an agent follows and the recommendations it delivers.
235. When Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned Routing
- Authors: Jian Chen , Zixuan Yuan
- URL: https://arxiv.org/abs/2609.33505
- Abstract:
Modern learned systems increasingly combine learned components with search, repair, or external solvers. Benchmarks often measure the resulting end-to-end system, while scientific claims may concern only one component, creating an attribution problem: evidence can fail to support the requested component-level claim while still supporting a positive conclusion about the larger system. Existing evidence-to-claim methods primarily calibrate claim strength. We argue that composite systems require a second dimension: scientific subject. We address this problem with subject-typed claim licensing, which separates weaker conclusions about the requested subject from positive but non-substitutive credit about another subject. We instantiate this idea in SCOPE-Routing for preference-conditioned multigraph routing. Non-authors reproducibly apply the declared semantics; held-out review yields fewer reference-relative upward deviations than unstructured review, while the difference from a strong evidence checklist remains unresolved; and a controlled routing study shows that score-optimal and claim-eligible methods can differ while valid hybrid-system credit is preserved. These results motivate treating claim strength and scientific subject as distinct dimensions of evidence-based evaluation.
236. RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse
- Authors: Ruoling Qi , Yirui Liu , Xuaner Wu , Yuxin Jin , Jian Chen , Jiayu Qin , Yin Chen , Jiawei Shao
- URL: https://arxiv.org/abs/2609.33503
- Abstract:
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.
237. Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning
- Authors: Debasmita Dey , Tanmay Sen , Himel Mallick
- URL: https://arxiv.org/abs/2609.33492
- Abstract:
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.
238. Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
- Authors: Yirui Liu , Ruoling Qi , Xuaner Wu , Yuxin Jin , Jian Chen , Penghang Liu , Yafei Huang , Jiawei Shao , Xuelong Li
- URL: https://arxiv.org/abs/2609.33477
- Abstract:
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing systems materialize recurrent-state checkpoints, restricting prefix reuse to checkpoint-aligned positions. We present SuffixReplay, the first prefix caching system that lets hybrid LLMs reuse cached prefixes at every cache-supported page boundary without materializing recurrent-state checkpoints. Our key insight is to just let linear states forget the distant past. Modern linear-attention mechanisms use recurrent decay and gating to attenuate the influence of old inputs. Therefore, instead of checkpointing every prefix boundary, SuffixReplay approximates the state at a matched boundary by replaying only a recent suffix of the layer’s input hidden states, which we retain as anchors. At the algorithmic level, SuffixReplay combines layer-wise and token-wise anchor sparsity with a bounded replay budget to control storage, computation, and quality. At the system level, it uses an independently managed anchor sidecar and a pipelined replay path to overlap anchor movement and state reconstruction with the native serving pipeline. We evaluate SuffixReplay on three hybrid LLMs: OLMo-Hybrid-7B, Qwen3.5-4B, and Qwen3.6-27B-FP8. Across these models, SuffixReplay retains 91.4-100% of full-prefill quality on average across LongBench and RULER, while using only 0.36-0.51x the amortized per-token storage of SGLang’s default 8192-token checkpoint cache. Integrated into SGLang, SuffixReplay reduces median TTFT by 15-70% on branching workloads, sustains 2.3-4.3x SGLang’s throughput when the working set exceeds HBM, and matches SGLang on high-hit continuation traffic.
239. LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs
- Authors: Haochen Luo , Yifan Li , Binh Minh An , Xiaolong Luo , Zhengzhao Lai , Yuan Zhang , Chen Liu
- URL: https://arxiv.org/abs/2609.33470
- Abstract:
Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challenges, including nonlinear payoffs and multi-leg strategy construction, requiring structured decisions rather than simple directional bets. We introduce LiveOption, an evaluation framework for LLM-based agents in option trading. LiveOption formulates the problem as structured sequential decision-making under realistic execution and capital constraints, and provides a reproducible environment with standardized interaction protocols. The framework includes three task suites covering portfolio overlays, event-driven earnings trading, and 0DTE intraday trading. We further propose a hierarchical metric suite that evaluates action validity, decision quality, risk characteristics, and outcome-level performance. Experiments show that current agents often fail to achieve competitive returns in most scenarios. LiveOption offers a principled testbed for evaluating structured decision-making beyond outcome-based metrics.
240. When Does the Concept of “Dog” Emerge in an Audio LLM?
- Authors: Zhe Wang , Shiqi Liu , Ruiyun Zhong , Tiechong Zhu , Yihua Tan
- URL: https://arxiv.org/abs/2609.33458
- Abstract:
Multimodal large language models answer audio questions, but how they represent auditory semantics and use them in decisions remains unclear, limiting our understanding of response formation. We study dog barking in Qwen2.5-Omni-7B using Jacobian lens (J-lens) readout and directional interventions. We define the dog direction as a J-lens-derived hidden-state vector associated with dog; adding or removing its component modulates dog-related information. We find this information decodable without dog/bark prompt cues or animal-identification requirements. Directional interventions change response tendencies and some final answers, with effects concentrated in late-layer states immediately before generation across species classification, vocalization classification, and sound description. The dog direction shows no comparable advantage over controls in animal/other classification. These results provide causal-intervention evidence that the dog direction affects output scores in a task-dependent manner, most consistently at L22 and L24 immediately before generation.
241. What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation
- Authors: Zzizhuo Lin , Quanling Liu , Yi Yang , Yawei Luo
- URL: https://arxiv.org/abs/2609.33455
- Abstract:
On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student’s reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher–student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.
242. MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI
- Authors: Md. Tanvir Rahman , Nabil Anan Orka , Asaduzzaman Khan , Mohammad Ali Moni
- URL: https://arxiv.org/abs/2609.33440
- Abstract:
Objective cognitive assessment from neural signals supports neurorehabilitation, but individual-level prediction from task-based fMRI (tfMRI) remains difficult because neural features coexist with substantial demographic and scanner-related variation. We present the Multi-task Activation and Contrast Network (MAC-Net), a covariate-aware deep learning framework for modeling individual cognitive function from regional tfMRI. By isolating tfMRI features into a dedicated neural pathway and restricting participant variables to a terminal late-fusion pathway, MAC-Net prevents dominant covariates from suppressing high-dimensional clinical representations during feature learning. Evaluating baseline data from 6,500 Adolescent Brain Cognitive Development Study participants under family-aware cross-validation, MAC-Net was benchmarked against linear models, random forests, and alternative deep architectures. The N-back plus Monetary Incentive Delay configuration achieved $R^{2}$ values of 0.174, 0.238, and 0.277 for fluid, crystallized, and total cognition, outperforming covariate-only baselines (0.178) and alternative deep models (0.217). N-back was the most informative paradigm, whereas incorporating the Stop Signal Task marginally degraded performance. Feature attributions via Integrated Gradients, DeepLIFT, and Input Gradient were highly concordant, localizing working-memory-related frontal, parietal, and cingulate regions. These findings demonstrate that covariate-aware multi-task modeling yields reproducible cognitive-function estimations, establishing a robust neural engineering framework for clinical translation.
243. Raven: The Harness of Harnesses for Composable Agentic Intelligence
- Authors: EverMind AI
- URL: https://arxiv.org/abs/2609.33439
- Abstract:
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model–harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
244. APEX: An Extensible Model for Agent-Assisted Production Scheduling
- Authors: Felix J. Grumbach , Stefan Görlitz
- URL: https://arxiv.org/abs/2609.33430
- Abstract:
Production scheduling requires realistic models that reflect operational constraints and efficient methods that balance competing goals. Putting these methods into use also requires data integration, model adaptation and specialist expertise. We present APEX, an extensible production scheduling framework built around a general model and hybrid multiobjective search. Agent assistance supports both scheduling and model refinement: agents prepare data and explore scenarios in natural language, while coding agents help implement and test new constraints and objectives. Shared construction and checking procedures connect these adaptations to the scheduling core. We benchmark eight APEX configurations against NSGA-II, SPEA2, MOEA/D and SMS-EMOA on 69 public job-shop, flexible job-shop and permutation flow-shop instances, assessing workload completion time (makespan), total job flowtime and computation time. Hybrid configurations achieve the best aggregate solution quality, although the leading method depends on the problem class and objective. A separate synthetic workflow study uses OpenAI’s GPT-6-astra as an interaction layer between the human planner and the algorithmic core, testing rule additions, plan and objective changes, and what-if comparisons. All 24 sessions completed the requested changes and passed independent checks of saved models and schedules. A separate coding evaluation produced six native implementations of an additional objective or hard constraint through predefined extension hooks. All passed independent checks without modifying the core.
245. Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence
- Authors: Md. Tanvir Rahman , Nabil Anan Orka , Asaduzzaman Khan , Mohammad Ali Moni
- URL: https://arxiv.org/abs/2609.33428
- Abstract:
Wearable actigraphy offers a scalable, ecologically valid alternative to episodic clinical assessment. However, predicting continuous adolescent crystallized intelligence ($G_c$) from such traces remains challenging due to irregular device adherence and complex behavioral-environmental interactions. We address this using daily summary data derived from 21-day Fitbit records of 6,091 adolescents in the Adolescent Brain Cognitive Development Study (Release 5.1). We propose SATURN, a Sleep-Activity Temporal Unified Regression Network. It represents participants as 21-node temporal graphs encoding daily behaviors and temporal adjacency. To prevent imputation artifacts, invalid-day edges are dynamically pruned during forward passes. Node embeddings are refined via residual GATv2 layers, aggregated through masked attention pooling, and fused with sociodemographic covariates. Under family-controlled, age-sex-BMI-stratified cross-validation, SATURN achieves $R^2 = 0.2783 \pm 0.0127$, consistently improving upon flattened machine learning (Gradient Boosting, $R^2 = 0.2372$) and sequential deep learning (BiLSTM, $R^2 = 0.2688$) baselines. Explainability analyses identify light activity, metabolic equivalents, and sleep duration as dominant predictors, while Monte Carlo dropout and subgroup analyses confirm equitable performance across sociodemographic strata. Ultimately, SATURN establishes a rigorous computational framework for digital cognitive phenotyping, offering a scalable pathway to complement traditional assessments by highlighting macro-level behavioral anomalies.
246. MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
- Authors: Xuanjun Chen , Hua-Hsuan Chen , Wei-Chung Lu , Yinghao Ma , Jyh-Shing Roger Jang , Hung-yi Lee
- URL: https://arxiv.org/abs/2609.33411
- Abstract:
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.
247. COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement
- Authors: Siwei Chen , Xinping Bao , Xinyu Cai , Yuan Cao , Wan Jiang , Shaohong Chen
- URL: https://arxiv.org/abs/2609.33398
- Abstract:
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter–context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.
248. CoViST: Visual Token Compression via Composable States
- Authors: Qi Zhang , Xiandong Meng , Ronggang Wang , Siwei Ma
- URL: https://arxiv.org/abs/2609.33397
- Abstract:
Visual token compression lowers the inference cost of vision–language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9\%, 99.5\%, and 98.1\% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8\%, 99.9\%, and 99.1\% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.
249. Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection
- Authors: Xinzhe Li , Youzhi Tu , Kong Aik Lee
- URL: https://arxiv.org/abs/2609.33394
- Abstract:
The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.
250. DrafTS: Time-Aware Decomposition with Residual Correction for Time Series Modeling
- Authors: Yiqiu Liu , Siru Zhong , Zhiguang Wang , Qingsong Wen , Yuxuan Liang
- URL: https://arxiv.org/abs/2609.33368
- Abstract:
Real-world time series contain evolving underlying dynamics with irregular variations that lack stable temporal patterns and are often referred to as noise. Existing methods address this mixture by filtering frequencies or suppressing noisy observations. They either miss temporal evolution or risk suppressing useful dynamics. We propose DrafTS, a model-agnostic framework that aims to reduce noise while preserving evolving dynamics through time-aware Decomposition with ResiduAl correction For Time Series. DrafTS uses features derived from instantaneous amplitude and frequency to guide decomposition into a primary component intended to capture underlying dynamics. A task-specific backbone models the primary component, while a lightweight correction module uses residual information to correct the backbone output. Across four time series modeling tasks, DrafTS improves six diverse backbones, demonstrating its effectiveness. Code is at this https URL
251. DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?
- Authors: Nan Huang , Mario Tapia-Pacheco , Kun Zhou , Yiming Huang , Kevin José Barrientos Díaz , Tiffany Amariuta , Jingbo Shang
- URL: https://arxiv.org/abs/2609.33357
- Abstract:
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: this https URL
252. Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks
- Authors: Analog Design Bench Team
- URL: https://arxiv.org/abs/2609.33356
- Abstract:
Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 chip designers. Agents work with an open-source simulator, while an isolated verifier evaluates the submitted circuit using specification-based electrical tests. We evaluate 15 agent configurations across 2,250 two-hour attempts and observe full-specification pass rates from 8.0% to 78.0%. Coding-benchmark performance correlates with analog results but leaves much of the performance spread unexplained. Our failure analysis shows that most unsuccessful submissions have no recorded legality rejection but fail electrical acceptance, identifying electrical closure as the dominant endpoint challenge. We test time, reasoning effort, agent harness, and supplied design knowledge as interventions. Longer budgets and higher reasoning effort improve performance, while general skill documents provide little benefit and sometimes reduce performance. Supplying a task-matched reference topology, an idealized form of circuit-IP retrieval, raises DeepSeek V4 Pro by 18.7 percentage points and mainly accelerates GPT-5.6 Sol.
253. Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
- Authors: Injin Kong , Sunghwan Choi , Yohan Jo
- URL: https://arxiv.org/abs/2609.33355
- Abstract:
Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices should change during generation. We study this question through strategy reversals, where an alternative action becomes preferable to a fixed choice. We organize MDM inference into five axes–score, cardinality, region, commitment, and planning–and define adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action. This view shows that adaptation value depends on both the frequency and magnitude of such reversals. Across three MDMs and ten tasks, adaptation opportunities are highly heterogeneous, with some regimes exhibiting concentrated and predictable one-step gains. This motivates selective adaptation: lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing, for example, 56.9 percent of the candidate-set oracle opportunity by adapting only the top 10 percent of states on LLaDA-8B constrained JSON filling. Our transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly.
254. QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
- Authors: Hyojun Ahn , Emily Jimin Roh , Soohyun Park , Walid Saad , Hyung-Chul Lee , Joongheon Kim
- URL: https://arxiv.org/abs/2609.33351
- Abstract:
Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID’s 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.
255. Naturalness-guided Manifold Flow Matching for Sign Language Production
- Authors: Jiayi He , Shengeng Tang , Sisi You , Yanbin Hao , Lechao Cheng , Richang Hong
- URL: https://arxiv.org/abs/2609.33339
- Abstract:
Sign Language Production (SLP) aims to generate sign motions from text. Conditional Flow Matching methods have achieved strong performance in SLP by constructing conditional paths that transform a source distribution into a target distribution. However, existing methods construct these paths via linear interpolation, whereas the rotational geometry of human joints confines valid joint rotations to a manifold embedded in Euclidean space. Consequently, linear interpolation between two sign motions leaves this manifold and ignores the motion distribution on it. In this paper, we revisit SLP from the perspective of manifold transport and propose a Naturalness-guided Manifold Flow Matching framework, termed \textbf{SignNMFlow}, which constructs conditional paths directly on the motion manifold by jointly considering geometric efficiency and the motion distribution. Specifically, we exploit the intrinsic geometry of the manifold and introduce a motion naturalness measure to characterize the motion distribution. By minimizing the kinetic energy under this measure, we learn a naturalness-guided interpolation that couples a closed-form geodesic, which provides geometrically efficient transport, with a learnable deviation that incorporates the motion distribution, thereby significantly improving the fidelity of generated sign motions. Extensive qualitative and quantitative evaluations demonstrate the effectiveness of this work.
256. ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces
- Authors: Jerry Wang , Haibo Jin , Xiaopeng Yuan , Peng Kuang , Haohan Wang
- URL: https://arxiv.org/abs/2609.33326
- Abstract:
Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what the query still requires. We introduce ANTMAN, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination. ANTMAN maintains a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress, and uses this state to control worker selection, routing, and task-local recovery as new evidence is discovered. By separating the coordination policy from substrate-specific search interfaces, the same need-conditioned mechanism can operate across different information spaces. Experiments across multi-document question answering, controlled long-context scaling, and realistic structured navigation show that ANTMAN remains effective across settings, including when execution is delegated to substantially smaller worker models. Under a 16x increase in searchable context, ANTMAN increases active coordination by only 1.23x, compared with more than 15x for partition-driven baselines, while preserving strong answer quality.
257. Agentic Multi-Turn Reasoning: A Fairness Approach
- Authors: Thanh-Dat Truong , Sankalp Pandey , Hugh Churchill , Jackson Cothren , Marios Savvides , Khoa Luu
- URL: https://arxiv.org/abs/2609.33323
- Abstract:
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $\Phi$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.
258. PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning
- Authors: Kecheng Liang , Haoyang Liu , Zexin Chen , Zirong Liu , Weixing Chen , Qiufeng Wang , Yang Liu , Liang Lin
- URL: https://arxiv.org/abs/2609.33319
- Abstract:
A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the underlying physical premise and lead to incorrect reasoning. To systematically study this challenge, we introduce \textbf{PhysAlign}, a benchmark designed to assess whether multimodal models correctly associate information recognized from physics diagrams with its intended physical role. By disentangling visual recognition from physical-role assignment through localized probes and controlled variants, PhysAlign isolates correspondence errors from recognition failures. It contains 3,341 human-validated probes spanning 986 physics problems, enabling systematic evaluation of visual recognition and physical-role correspondence at scale. We further introduce five complementary evaluation metrics, including CAcc, GAcc, and JAcc, which provide a comprehensive assessment of models’ ability to recognize diagram content, establish correct physical correspondences, and solve the underlying physics problem. Across our evaluated multimodal models, PhysAlign reveals a consistent gap between local visual recognition and physical-role grounding. Even when the queried content is correctly recognized, the conditional correspondence error rate remains 13.8\% for GPT-6-Astra and rises to about 50.6\% for InternVL3.5-8B. These findings indicate that strong perception alone does not ensure reliable physical interpretation, exposing a distinct grounding bottleneck that is largely hidden by answer-level accuracy and highlighting the need for future models to better align recognized visual evidence with its physical meaning.
259. The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization
- Authors: Yiguo Wang , Ziyuan Yang , Yi Zou , Dan Lin , Rongsheng Li , Yi Zhang
- URL: https://arxiv.org/abs/2609.33297
- Abstract:
Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.
260. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
- Authors: Dehai Min , Daoan Zhang , Yiming Zeng , Huayi Zhang , Ziyi Chen , Yan Zhang , Qinbo Bai , Mengyuan Chao , Jing Ning , Qiyue Hua , Huiyi Chen , Hanrong Zhang , Henry Peng Zou , Jie Yang , Wei Xu , Philip S. Yu
- URL: https://arxiv.org/abs/2609.33295
- Abstract:
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM’s next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader’s agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
261. Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
- Authors: Shuze Daniel Liu , Claire Chen , Jiuqi Wang , Thorsten Joachims
- URL: https://arxiv.org/abs/2609.33289
- Abstract:
Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.
262. Feedback Makes Perfect: A Closed-Loop Framework for NL-to-STL Translation
- Authors: Bowen Ye , Xiang Yin
- URL: https://arxiv.org/abs/2609.33287
- Abstract:
Signal Temporal Logic (STL) enables rigorous verification and control of cyber-physical systems, but writing correct specifications requires expertise that most requirement holders lack. Large language models can translate natural-language (NL) requirements into STL, yet stronger translators alone approach an accuracy ceiling. We argue that this ceiling stems from how the task is posed: one-shot, open-loop translation is somewhat ill-defined. Natural language is ambiguous, and, more fundamentally, what a person writes may not always be what they intend, so the target specification is not fully contained in the input text. We therefore reformulate NL-to-STL translation as a closed-loop feedback process. Each generated formula is translated back into natural language for the user to check, and natural-language corrections drive revision until the user accepts the specification. Users never read or write formal syntax. This framework rests on an asymmetry familiar from feedback control theory. The forward path, from ambiguous language to formal logic, is hard and error-prone. The feedback path, from structured STL back to language, can be made highly precise, and a precise feedback path lets an imprecise forward path achieve precise closed-loop behavior. Experiments on 500 expert-authored requirements and seven LLMs support this view. Back-translated explanations agree with expert judgments in 99.5\% of cases. Closed-loop refinement raises strong models from about 89\% open-loop accuracy to 98.0–99.2\%, and yields gains of over 30 percentage points for weaker models (e.g., 17.6\%$\rightarrow$48.0\%). Ablations show these gains come from the semantic content of the feedback rather than from repeated attempts. An expert audit and a 280-session user study further confirm the reliability of the loop. We also identify a capability threshold above which feedback no longer helps.
263. RINI: Seeing the Prior Is Not Enough
- Authors: Hongyi Du , Tianyi Zhang , Heng Wang , Zhelun Gao , Yimei Liu , Ambrose Luo , Annie Hao , Jiayan Ni , Jiawei Han , Jiaxuan You
- URL: https://arxiv.org/abs/2609.33284
- Abstract:
A research proposal can describe an established mechanism correctly while claiming to introduce it. We study whether providing the earlier paper corrects such contribution claims. Three controlled experiments compare proposals generated with a contribution-bearing prior and a same-topic control. Providing the prior yields no clear aggregate reduction in unsupported novelty. Human analysis of 175 interpretable exposed proposals finds that 137 recognize the prior’s relevance, but 61 correctly attribute the established contribution. Of 71 proposed remaining distinctions, 37 are covered by the same prior. We introduce Research Idea Novelty Inspection (RINI), which audits contribution claims against evidence, checks the remaining distinction, and applies local revisions. Five human annotators evaluate 1,080 original-revision pairs across three methods. On the same 240 originals judged to require correction, successful repair is 11.7% for Self-Revision, 39.1% for Retrieve-and-Revise, and 72.2% for RINI, with research tasks weighted equally. The improvement over same-evidence direct revision is 33.0 percentage points. The revised proposals retain their research questions and technical methods. These results motivate explicit contribution attribution when using literature to generate and revise research proposals.
264. Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications
- Authors: Xianglong Shi , Shifeng Liu , Sirui Zhao , Shengming Yuan , Enhong Chen
- URL: https://arxiv.org/abs/2609.33282
- Abstract:
Subjective judgments are central to many high-traffic applications, but subjective intensity is difficult to quantify and perceptions vary substantially across individuals. To address these challenges, we propose a pairwise comparative framework for multi-dimensional scale construction. By comparing case-person pairs along case and profile dimensions, the framework constructs relative scales that capture both fine-grained intensity and individual variation. To support practical high-traffic deployment, we optimize both offline scale construction and online inference. For scale construction, we combine sparse Elo comparisons with multi-judge voting, cutting the comparison cost from $O(N^2)$ to $O(NK)$ for $N$ objects and a budget of $K$ opponents per object, while limiting reliance on any single judge. For inference, we propose SubJudge, a System One model for personalized scoring with Batchwise Preference Optimization (BPO). Using Bradley-Terry comparisons, BPO trains the model to learn relative orderings, and SubJudge reads a continuous score from digit-token probabilities at the first response position, requiring only one forward pass per criterion and reducing the inference complexity to $O(1)$. Experiments on PluriHarms and iNews show that our 9B models match or surpass the evaluated frontier LLMs on multiple metrics. On the H100 GPU, SubJudge achieves an approximately $1.29\times$ to $261\times$ speedup in mean inference latency over Qwen3.5-9B with different thinking budgets. The code is available at this https URL .
265. ChronoFlow: Hierarchical Flow Matching for Irregular Time Series Generation
- Authors: Changhun Kim , Sunguk Jang , Jeongjun Lee , Juhwan Choi , Sangchul Hahn , Grigorios Chrysos , Eunho Yang , Juho Lee
- URL: https://arxiv.org/abs/2609.33276
- Abstract:
Recent advances in generative modeling have substantially improved time series generation, yet most existing methods either assume a regular temporal grid or focus on feature dynamics under a given sampling structure. This makes them illsuited for generating irregular time series in their native form, where a model must capture not only feature values, but also how many observations occur, when they occur, and which features are observed together. To address this heterogeneous generation problem, we propose ChronoFlow, a unified hierarchical flow matching framework organized by statistical granularity. Following a coarse-to-fine hierarchy, ChronoFlow first generates observation counts and feature-wise frequencies, then jointly generates observation times and feature co-observation patterns, and finally generates values conditioned on the realized pattern. This turns a complex joint generation problem into structurally aligned subproblems while preserving their dependencies. To evaluate complete irregular time series generation, we introduce complementary metrics spanning sample realism, sampling structure, value fidelity, and temporal and cross-feature dependencies, and validate them through controlled corruptions. Across five benchmarks, ChronoFlow achieves strong improvements in generation fidelity over existing baselines, while factorization studies support the proposed hierarchy. Our code is available at this https URL .
266. Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space
- Authors: Yang Li , Yi Wang , Shiyuan Huang , Yang Liu , Hao Wang , Chengzhi Mao
- URL: https://arxiv.org/abs/2609.33271
- Abstract:
Reasoning problems often admit multiple valid ways to proceed. Continuous reasoning promises to move computation beyond language tokens into a more compact latent space, but representing several plausible ways to think next remains difficult. We introduce Autoregressive Thought Flow (ATF), which models the next continuous thought as a multimodal distribution. A causal autoregressive model performs the reasoning computation, while a lightweight diffusion head generates a plausible next thought from the resulting condition. The sampled thought is fed back into the model, allowing continuous reasoning to unfold for a variable number of steps while preserving the pretrained backbone. Across mathematical reasoning tasks, ATF improves accuracy with compact latent traces and benefits from reinforcement learning and additional test-time thinking. Multi-sample evaluation shows broader solution coverage, indicating that its multimodal predictions capture useful diversity among reasoning paths. Our results suggest that continuous reasoning is more effective when multiple possible next thoughts remain available rather than being collapsed into a single prediction.
267. Structured Sparse Memory for Recurrent Reasoning
- Authors: Zixuan Zhao , Samuel Wheeler , Neil Getty , Xiaotian Duan , Rick Stevens , Fangfang Xia
- URL: https://arxiv.org/abs/2609.33270
- Abstract:
Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a compact hybrid ARC model that combines recurrent reasoning with structured task memory, synthetic data, and inference-time aggregation. In existing approaches, task-conditioned memory supplies a large hidden source of capacity, reaching more than 30x the size of the recurrent backbone. We introduce a compositional sparse embedding (CoSE) for task conditioning that reduces learned task-memory parameters by over 90% while improving pass@2 in controlled ARC ablations. For the recurrent backbone, recurrent depth helps only when balanced with learning horizon. Combining these ingredients, our system reaches 84% pass@2 on ARC-AGI-1 and 46.7% pass@2 on ARC-AGI-2 public evaluation. The benefits of structured memory also generalize to unseen puzzles and other domains. Our code, dataset, and model checkpoints are available at this https URL .
268. LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models
- Authors: Xianglong Shi , Ruijie Yang , Sirui Zhao , Shukang Yin , Zihao Bian , Tinghao Yi , Enhong Chen
- URL: https://arxiv.org/abs/2609.33268
- Abstract:
Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to accumulate history and to serve readout, so what the memory stores cannot be controlled separately from what it exposes to the current computation. We propose LSTMem, an LSTM-inspired online memory that instead equips each layer of a frozen LLM with two matrix-valued states: a cell state that accumulates history and a hidden state whose readouts correct the backbone’s attention. Input and forget gates control what the cell stores, while an output gate separately controls what the cell exposes through the hidden state. LSTMem further connects memory across depth through forward hidden-state propagation and block-end feedback, and uses higher-layer reconstruction gradients to refine lower-layer cell states before rebuilding hidden states from shallow to deep layers. Across memory benchmarks on Qwen3-4B-Instruct, LSTMem consistently improves MemoryAgentBench, LoCoMo, and HotpotQA over the plain backbone. Comparisons further show that the LSTM-based memory formulation outperforms an associative-memory counterpart, while removing cross-layer hidden-memory propagation degrades performance. These results demonstrate the benefits of separating memory accumulation from memory expression and organizing memory hierarchically across model depth. The code is available at this https URL .
269. CORTEX: A Verified Experience Layer for Generalist Agents
- Authors: Garapati Keerthana , Manik Gupta
- URL: https://arxiv.org/abs/2609.33260
- Abstract:
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.
270. ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents
- Authors: Song-Li Wu , Jingyi Wang , Zhaocheng Du , Weinan Gan
- URL: https://arxiv.org/abs/2609.33244
- Abstract:
Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-step execution. As memory scales, such flat retrieval introduces context fragmentation and cross-task interference, leading to structurally inconsistent reasoning trajectories. We propose ActiveMem, a hierarchical memory framework that recursively organizes agent experiences into dependency-aware latent execution trees. ActiveMem abstracts trajectories into reusable subtask nodes while explicitly preserving execution transitions, enabling coherent reasoning-path retrieval conditioned on the current execution state. To support continual adaptation, ActiveMem further learns dynamic memory expansion, retrieval, and pruning policies through reinforcement learning. Experiments across various agent benchmarks demonstrate that ActiveMem consistently improves task completion, reasoning stability, and memory efficiency over existing memory-based agents. Moreover, ActiveMem enables compact open-weight models to achieve competitive performance with substantially larger proprietary systems.
271. CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents
- Authors: Song-Li Wu , Jingyi Wang , Zhaocheng Du , Weinan Gan , Weiwen Liu
- URL: https://arxiv.org/abs/2609.33243
- Abstract:
Code agents require long-horizon decision-making over complex interaction trajectories. However, existing reinforcement learning (RL) approaches typically optimize behavior at the token level, creating a mismatch between low-level generation and high-level behavioral reasoning. This limitation leads to inefficient exploration and weak credit assignment under sparse rewards. Moreover, while large-scale agent trajectories often contain recurring multi-step behavioral patterns, their noisy token-level representations hinder effective experience reuse. To address these challenges, we propose CodeSkill, a framework that adapts hierarchical latent skill modeling to the code agent domain. CodeSkill first leverages a teacher model to distill both successful and failed trajectories into multi-level textual abstractions. It then integrates temporal variational inference with reinforcement learning to map these discrete semantics into continuous latent variables, while an adaptive boundary mechanism dynamically gates skill transitions based on execution feedback. The learned skills are injected into a frozen LLM policy as latent semantic prefixes, enabling optimization in a compact semantic space rather than over raw token sequences. By shifting RL from token-level exploration to experience-level reasoning, CodeSkill improves optimization efficiency and long-horizon behavioral coherence. Extensive experiments demonstrate that CodeSkill achieves highly competitive performance against strong open-weight baselines across diverse general and industrial coding benchmarks. Furthermore, the learned skills exhibit strong transferability and robust cross-domain generalization, highlighting the effectiveness of explicit behavioral abstraction for scalable agentic code generation.
272. WorldAgent: Verification-Guided Agentic Physical World Construction
- Authors: Caoliwen Wang , Mengdi Wang , Yige Chen , Zejia Wu , Bowen Huang , Siyuan Chen , Guanxiong Chen , Lifu Wei , Heng Zhang , Qinghai Zhang , Yin Yang , Guandao Yang , Shiying Xiong , Peng Wang , Chenfanfu Jiang , Peter Yichen Chen
- URL: https://arxiv.org/abs/2609.33208
- Abstract:
Constructing complex physical worlds from language requires coordinating extensive 3D environments, detailed structures and objects at different spatial scales, and interacting physical processes under both stated goals and implicit physical constraints. We present WorldAgent, an agentic framework for verification-guided physical world construction from a single natural-language prompt, without iterative user debugging. A world construction layer expands the prompt into a structured world specification and uses physical knowledge to build scenes and run numerical simulations. After every step, a verification layer inspects scene geometry and simulation states alongside rendered views. Failed checks guide automatic revisions to the specification and re-execution of the affected steps. Accepted worlds pass the required checks and remain editable for further inspection and resimulation. We introduce AgenticSimBench, on which WorldAgent achieves the best scores among the evaluated agent-based methods on five of seven metrics. In a 26-participant user study, it receives the highest mean ratings across all four criteria.
273. Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries
- Authors: Haiquan Hu , Yuzhu Liang , Weicheng Tang , Yanzeng Li , Yao Shi , Tian Wang
- URL: https://arxiv.org/abs/2609.33196
- Abstract:
Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boundaries from repeated-response trajectories along ordered model axes. BSDProbe summarizes samples by boundary position, boundary width, boundary-signal validity, and order consistency, then aggregates them into benchmark-level structural profiles. Experiments on six benchmarks show that benchmark reliability is axis-conditioned and heterogeneous: GSM8K and MATH exhibit the most stable measurement structures, MMLU and TriviaQA are relatively stable but heterogeneous, while GPQA and PopQA show stronger axis-conditioned risks. These profiles remain consistent across Qwen3, Qwen2.5, and cross-model axes. BSDProbe further selects compact high-value subsets whose model discriminability reaches up to $8.58\times$ that of the full benchmark. These results suggest that reliable benchmark use requires examining sample-level capability boundaries beyond leaderboard scores.
274. Unlocking Latent Personalization in LLMs
- Authors: Wei Chen , Guanghui Zhu , Zhongliang Cai , Yihua Huang
- URL: https://arxiv.org/abs/2609.33182
- Abstract:
Large language models (LLMs) are increasingly expected to adapt to individual users, yet effective personalization remains challenging when only limited user-specific samples are available. In this work, we take an alternative perspective: pretrained LLMs may already possess latent capacity for personalization, and a few user samples may therefore suffice to guide the model toward user-aligned behavior with minimal user-specific adaptation. From this perspective, we propose LatentPersonal, a framework that formulates personalization as navigation in a shared latent adaptation space. LatentPersonal infers a compact latent representation from a few user samples to guide user-specific model adaptation, regularized with a variational information bottleneck to encourage compact preference representations. We instantiate LatentPersonal with LoRA, leveraging its low-rank parameterization as a natural low-dimensional adaptation space for personalization. By simply inserting a user-specific guidance vector between the shared low-rank factors, the model can navigate toward personalized adaptations through lightweight inference of this compact representation, without updating the shared LoRA parameters. Experiments across multiple personalization datasets demonstrate that LatentPersonal substantially reduces user-specific adaptation overhead while achieving effective personalization from only a few user-specific interactions, with particularly strong performance in the one-shot regime.
275. SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
- Authors: Xiaoshu Chen , Xiangyu Wong , Sihang Zhou , Ke Liang , Xinwang Liu
- URL: https://arxiv.org/abs/2609.33181
- Abstract:
Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.
276. Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning
- Authors: Hongbo Chen , Guohua Lu , Ting Dang , Hong Jia
- URL: https://arxiv.org/abs/2609.33149
- Abstract:
A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem’s constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.
277. LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
- Authors: Euntae Choi , Sumin Song , Sungjoo Yoo
- URL: https://arxiv.org/abs/2609.33146
- Abstract:
An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark. We propose LiteEvo, a lightweight harness-evolution algorithm whose tool-free meta-agents mine agent trajectories for reusable components, curate them into a versioned library, and compose each round’s harness from it, starting every benchmark from the same neutral harness and never naming the benchmark. Evolving on the graded tasks of five agentic benchmarks with a frozen Qwen3.5-9B, LiteEvo lifts pass@2 by 10.5 to 67.7pp and reaches comparable or higher pass@2 than a reproduction of HarnessX (71.0 against 67.3 on average) at 13.0 lower mean API cost. Harnesses evolved on train tasks keep their gains on unseen test tasks of four benchmarks, and LiteEvo also lifts Claude Code with Sonnet 4.6 by 1.2 to 71.4pp.
278. On Device Agentic Operation Caches – Classifier-Centric NL-to-Action Generation
- Authors: Moghis Fereidouni , Anthony Arnold , Sumit Gulwani , Mark Marron , A.B. Siddique
- URL: https://arxiv.org/abs/2609.33141
- Abstract:
Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resources run and depend on cloud hosted inference to handle the task of transforming natural language inputs into actionable software operations. This reliance on cloud-hosted inference introduces substantial network latency on top of LLM inference times, creates data privacy concerns, and, given the costs of running these models, can rapidly escalate expenses associated with supporting agentic features. This paper introduces a novel means of converting the NL-to-Action problem from a generative one into a classification-centric formulation via on-device operation caches. These caches allow an agentic system to handle frequently occurring classes of actions completely on-device – reducing latency, enhancing privacy, and lowering operational costs. We show that for a classic NL-to-Formula task, generating Excel Formula in response to user requests, this approach reduces total inference cost by 56% when compared to cloud-only model-routing based inference and, on cache hits, reduces the latency to response latency by 5x.
279. Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?
- Authors: Jiashu He , Jinxuan Fan , Xiao Xiao , Radu Marculescu , Alejandro Ribeiro
- URL: https://arxiv.org/abs/2609.33134
- Abstract:
Reasoning models generate long chains of thought before they answer, yet it is debated whether the content of these chains does real computational work or is largely decorative. We study this question by viewing a transformer as a shallow circuit. One forward pass through a fixed number of layers has constant depth, so any procedure that runs the model a constant number of times is a shallow circuit. We call the best accuracy that a shallow circuit can reach on a task the ceiling of the task, and a task is serial if its ceiling lies below one. We prove three results on serial tasks that hold for every transformer, no matter how it was trained. Necessity: replacing the chain by anything that does not depend on its content, such as filler tokens or a restatement of the question, drives the accuracy down to the ceiling, and on a maximally serial task down to chance. Depth: no shallow computation can write the chain of a model whose accuracy exceeds the ceiling, not even approximately. Locality: the answer is one shallow pass away from the finished chain, so all of the serial reasoning happens in the chain. On word problems of finite groups, whose ceilings are known, small transformers trained from scratch, with or without reinforcement learning, attain the predicted numbers: chain-trained models solve every input length and fall to chance when the chain is erased, chainless models collapse to the ceiling as the input length grows, and open-weight reasoning models given the same problem in words return to the baseline without their chain. On MATH-500 and AIME, erasing the chain costs open reasoning models 0.52 to 0.82 accuracy, a sentence shuffle is harmless, and a token shuffle is as harmful as erasing; the same holds for checkpoints trained by GRPO with a correct or a random reward. The ceiling of a task therefore answers when a transformer can succeed without its chain of thought.
280. Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring
- Authors: Zhixiang Zhang , Zesen Liu , Wai Ip Lai , Hongxu chen , Dongdong She
- URL: https://arxiv.org/abs/2609.33123
- Abstract:
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.
281. Modular Discovery of General Game-Playing Algorithms with Large Language Models
- Authors: Zun Li , John Schultz , Marc Lanctot , Daniel Hennes
- URL: https://arxiv.org/abs/2609.33115
- Abstract:
General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to discover general game-playing algorithms? Because language models can propose and refactor structured code, they provide an expressive proposal engine for exploring the space of algorithmic designs. We introduce a multi-agent LLM meta-learning system to co-evolve game-agnostic procedural search mechanisms in C++ alongside domain heuristics synthesized directly from game rules. Controlling the compute budget, we benchmark the discovered mechanisms across more than 400 diverse environments, including OpenSpiel training and held-out games, procedural simulation engines, and games with deep neural policy-value representations trained via PPO. Evaluated via AlphaRank stationary distributions and Soft Condorcet Optimization (SCO) against 15 established MCTS baselines, the discovered search mechanisms consistently achieve top-tier ratings and pairwise ballot majorities over most baselines across independent evolutionary runs, generalizing to unseen human-designed and procedurally synthesized games and remaining competitive with baselines on frozen neural network representations.
282. The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration
- Authors: Jin Cui , Xinyue Long , Boran Zhao , Pengju Ren , Hao Dong
- URL: https://arxiv.org/abs/2609.33085
- Abstract:
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem–Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.
283. Structure-Mapping-Guided Self-Explanation for Learning Mathematical Procedures
- Authors: Shinhaeng Lee , Christopher J. MacLellan , Daniel Weitekamp
- URL: https://arxiv.org/abs/2609.33079
- Abstract:
Worked examples are a powerful form of instruction, but learners must infer how the demonstrated steps were produced. A naive simulation of this self-explanation process can generate thousands of numerical explanations that reproduce one observed change without capturing its underlying procedure. We propose structure-mapping-guided self-explanation as a computational account of the cognitive biases that reduce search effort and make this inference tractable. The model represents mathematical expressions as typed relational structures and uses structure mapping to identify corresponding source and target regions. For each changed target value, the corresponding source region serves as an anchor: it guides abductive search toward structurally relevant values and operations before broader alternatives, yielding ordered, executable candidate procedures with inspectable source evidence. Across 70 mathematical transformations containing 120 changed numeric components, the model recovered every intended procedure. It returned the intended procedure before any other computation producing the same target value in 104 subproblems (86.7%), compared with a median of 32 (26.7%) across 100 unguided runs that tested candidate calculations in random order. Our proposed model also tested 93.8% fewer combinations of values and operations than unguided search before reaching the intended procedures. These results provide an efficient, interpretable account of how relational structure can guide procedural learning and a testable hypothesis about human self-explanation from worked examples.
284. QureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning Supervision
- Authors: Janhavi Prabhu , Sahil , Shivam Ashok Shukla , Manoj Tadepalli
- URL: https://arxiv.org/abs/2609.33075
- Abstract:
Radiological similarity depends on disease relationships and on fine details such as laterality, lobe, severity, size, and certainty. Broad biomedical similarity can overlook these qualifiers, particularly when several attributes vary together. We introduce QureRadEmbed, a 4B radiology-aware encoder trained with two complementary signals: RadSim supplies deterministic, attribute-decomposed ranking targets, while RadThought aligns reports with hierarchical evidence and reasoning descriptions. A three-stage curriculum combines these signals with report triplets, finding perturbations, and single- and cross-attribute contrasts. The final model achieves 0.996 mean ordering accuracy across ten controlled synthetic attributes and raises Spearman correlation with the designed joint-attribute targets from 0.501 to 0.976. On external findings-to-impression retrieval, Recall@1 reaches 10.5% on Open-I, 12.4% on testing XR, and 42.9% on testing CT, compared with 6.6%, 5.7%, and 31.4% for its backbone. Frozen embeddings support finding extraction with only 100 labeled testing-XR reports (macro-F1 0.481 versus 0.412 for the backbone). Whole-report comparison costs 8.8 seconds per 1,000 pairs in our benchmark, versus 2,755.1 seconds for the generative evaluator GREEN. Sentence-level comparison improves sensitivity to local discrepancies, although generative evaluation remains stronger on several expert-rated and subtle-error tasks. The results support reusable radiology-aware representations for search, structured report indexing, and efficient report comparison.
285. LLM sequential decision making under uncertainty in biochemical domains
- Authors: Mattias Akke , Soojung Yang , Jurgis Ruža , Sathya Edamadaka , Rafael Gómez-Bombarelli
- URL: https://arxiv.org/abs/2609.33061
- Abstract:
Large language models (LLMs) are increasingly used to drive scientific discovery. Understanding how LLMs make decisions from new data and memory of the literature is vital before trusting them to design experiments under tight experimental budgets. However, their decision strategies are invisible in the current performance scores used to evaluate research agents. Here, we benchmark five frontier LLMs in a Bayesian Optimization setting against published statistical baselines on seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly, and catalysis. Performance is paired with direct measurements of model beliefs and actions, enabling highly resolved behavior analysis. A prompt ablation that progressively strips context separates memorization from chemical reasoning and from bare categorical optimization. Prior chemical knowledge helps in expectation, but with high variance and occasionally even harms performance. No configuration tested decisively beats a mean statistical baseline across domains. Belief-movement and Martingale diagnostics, corrected here for a measurement-noise bias that mislabels rational agents as irrational, show that models overreact to incoming data rather than entrenching on their priors in the contexts studied here. Interestingly, while LLM actions are exploitative, models sincerely intend to explore and consistently act on that intent. This failure is a competence gap arising from context-stickiness. Removing in-context history restores exploration, indicating that priors and data must be decoupled to achieve effective LLM-driven discovery.
286. Large Language Models Substantially Compress Well-Being Inequality but Largely Preserve Its Socioeconomic Structure
- Authors: Nattavudh Powdthavee
- URL: https://arxiv.org/abs/2609.33055
- Abstract:
Research using large language models (LLMs) to generate synthetic populations has repeatedly shown that model outputs compress the diversity of human experience. This has raised doubts about whether LLM-generated data can capture meaningful differences within populations. We show that such compression does not necessarily erase the social structure of human heterogeneity. Using 93,901 respondents from 66 countries and territories in Wave 7 of the World Values Survey, we ask six LLMs to predict respondents’ life satisfaction from demographic, socioeconomic, and attitudinal profiles. All six models substantially understate the overall dispersion of life satisfaction. Yet after normalizing for these differences in scale, they largely reproduce the human income gradient in well-being inequality: lower-income groups remain relatively more heterogeneous than higher-income groups. The pattern is robust to country fixed effects, equal-country weighting, WVS survey weights, and observed demographic composition, and it extends directionally to employment, education, and perceived control. Fidelity is weaker for extreme outcomes and country-specific gradients. These results show that the amount of heterogeneity preserved by an LLM and the way that heterogeneity is distributed across social groups are distinct properties. LLM-generated populations can therefore substantially compress human variation while retaining meaningful information about where that variation is concentrated.
287. BudgetVerify: Budget-Tiered Verification for Financial QA
- Authors: Janet Jenq , Hongda Shen
- URL: https://arxiv.org/abs/2609.33052
- Abstract:
Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three verification tiers: no verification, lightweight check-and-revise, or higher-cost solve-first-then-compare verification. The router is trained from offline correctness and token-cost outcomes and, at test time, selects a verification tier using information available before verification, including the question, context statistics, the generated answer, and associated generator metadata. The selected tier either returns the generated answer directly or invokes the corresponding verifier. Across six commercial and open-weight base models, BudgetVerify consistently produces more efficient accuracy-cost Pareto frontiers than fixed verification policies by selectively allocating stronger verification only when it is useful. Although absolute performance varies across models, these efficiency gains and the resulting qualitative frontier shape are consistent across generator models.
288. Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States
- Authors: Difan Jiao , Ashton Anderson
- URL: https://arxiv.org/abs/2609.33039
- Abstract:
Language model agents can now perform sophisticated sequences of actions via tools and harnesses, which has increased the scope of the damage they can cause. Guard models, however, are mainly built for content moderation and thus are not well-suited to detecting this agentic risk. To address this, we proceed by first conducting a representational analysis, then use the resulting insights to build a solution. In our analysis, we focus on two types of trajectory-level agentic harms: harmful content, which is expressed directly, and unsafe tool use, which depends on whether an action is consistent with the interaction that produced it. We investigate how open-source guard models represent these two types of harm and find that they are linearly readable inside the model, even though guard models predict no better than chance on pairs that differ only in the called tool’s schema. The two harm types also follow nearly orthogonal internal directions, and neither reliably serves as a proxy for the other. These results motivate reading trajectory safety directly from internal states. We introduce TACIT, a readout of a frozen backbone’s internal states that decodes no tokens. Trained on six trajectory-safety benchmarks, a linear probe raises mean macro-F1 from 62.3 for the strongest open guard to 80.7, and refined readouts reach 86.2. With each benchmark held out of training entirely, the refined readouts still lead the strongest guard (65.7 vs. 61.1). With the same backbone, training data and test split, the frozen readout is on par with full safety fine-tuning, and it improves the fine-tuned model further when applied on top. The probe trains about one millionth as many parameters as full fine-tuning in about a sixth of the time, and TACIT has the lowest latency of the guards we evaluate.
289. SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents
- Authors: Yifang Tian , Yingjian Bai , Yifeng He , Zichun Chong , Yuanchen Gao , Yiran Li , Hans-Arno Jacobsen
- URL: https://arxiv.org/abs/2609.33023
- Abstract:
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We present SRE-Marathon, a benchmark for long-horizon, continuous SRE operation. An agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, production-calibrated schedule. Curated code and configuration changes deployed through the same build pipeline available for repair. Each run is recorded into a sealed bundle and scored offline: Marathon-Score credits each injected fault for ordered progress through correlation, localization, and repair, with all metrics computed deterministically from recorded system evidence. Across three applications and about sixty faults per run, the best of 10 methods reaches only 41.3 out of 100. Agents often correlate and localize faults, but almost never complete repairs while the faults remain active.
290. Trust and Task Completion in the World of Consumer AI Agents
- Authors: Jeroen Olieslagers , Eduardo Pujol , Gal Zahavi , Lukas Ingemarsson , Shivani Poddar
- URL: https://arxiv.org/abs/2609.33017
- Abstract:
Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold back after the user clearly said go. And they fall short on completion when they give up on errands that turn out to be hard. Both depend heavily on the harness around the model, meaning its instructions, tools, context, and guardrails. We built an evaluation that scores trust and completion on the same runs, in a simulated world of businesses with their own websites, inboxes, and phone lines, and of people who write back. A simulated user answers the assistant’s questions. Trust means that nothing happens the user did not agree to. No email goes to someone they never approved, no private detail ends up on a group thread, no money is spent past their limit, no stranger’s instructions are followed, and nothing is claimed without a source. Every trap has a matched control in which acting is the right call. We use the evaluation to measure Fo, Wajo’s personal assistant, against a base model with basic instructions on three foundation models, and against the Fo harness with its guardrails switched off. Fo completes 71% of the errands and keeps the user’s trust on 94% of the trap runs. The base models complete 50% to 64% and keep trust on 59% to 75%. On the matched controls, Fo goes ahead slightly less often. OpenClaw, a popular open-source assistant given the same access, completes 42% of the errands it shares with Fo, against 71%, and keeps the user’s trust on 74% of the shared trap runs, against 94%. Measuring trust and completion together, on the whole system rather than the model alone, is how we think action agents become safe to hand real work to.
291. The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents
- Authors: Sasank Annapureddy , Anjaneya Prasad Thamatani
- URL: https://arxiv.org/abs/2609.33013
- Abstract:
Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how much an agent remembers, to whether its consolidation decisions are any good, to whether those decisions can be trusted. Phase 1 learns episodic boundaries from agent traces by downstream utility; an honest near-miss (oracle correlation 0.691 vs a 0.70 bar) whose lasting output is a three-gate anti-leakage protocol. Phase 2 learns when to promote experience and to which abstraction level under a token budget, achieving a verified +22.7% task-success improvement with 7x compression, but exposing a degenerate-forgetting failure and a distribution-shift failure mode we name lambda-prevalence coupling. Phase 3 introduces ConsolidationBench, an oracle-by-construction benchmark that scores consolidation decisions against a known optimum on three non-circular axes; production retrieval systems retain information yet score zero on cross-level transfer. Phase 4 introduces governed consolidation: the decision wrapped in poison-resistance, reversibility, and auditability guarantees with a quality gate. Governance is statistically distinct from the quality score ($r^2 = 0.43$; partial $r = 0.27$; identical-quality policies differ threefold in governance), so the contribution survives independently of the metric’s external validity. On that question we report a resolved negative: after a graded-reuse redesign removed a structural ceiling, a two-benchmark study with 2,532 real answer cells finds the quality score does not predict real transfer accuracy (pooled Spearman $\rho = -0.24$, n = 12, CI spanning zero). An adversarial self-critique pass cleared the final claim set with zero surviving overclaims.
292. When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate
- Authors: Wei-Jung Huang
- URL: https://arxiv.org/abs/2609.33012
- Abstract:
When an agent benchmark compares every pair of leaderboard entries, the number of comparisons can look much larger than the independent evidence behind them: A versus B and A versus C both reuse A. Whether this reuse affects inference depends on what the analysis is meant to describe. If the board and its outcomes are fixed, the all-pairs mean is an exact summary of those entries, and any interval must come from another declared source of randomness. If the entries are instead treated as iid draws from a population of future configurations and the pair rule is regular and nondegenerate, the same mean is an order-two U-statistic whose first-order uncertainty depends on the number of configurations, not the number of pairs. We use near ties as the running example, but the distinction extends to other symmetric pair summaries when their regularity conditions hold. We examine both interpretations using a fixed SWE-bench Verified snapshot and an exact binary model with known truth. On SWE-bench, intervals that accounted for shared configurations were more than twice as wide as a pair-iid reference that treated the pairs as independent. In the exact model, pair-iid coverage fell far below the nominal level when edges shared endpoints but remained near nominal for matched independent edges. Results on two other fixed leaderboards show that exact summaries also depend on which pairs are included and how they are weighted. An all-pairs analysis must therefore state what is fixed, what is sampled, and how it handles shared entries and pair aggregation.
293. Model-Aware Data Selection from In-and-Out Information Interplay
- Authors: Yifan Wang , Xiaomin Li , Yuexing Hao , Dongwon Jung , Hemanth Neelgund Ramesh , Ananth Grama , Varun Chandrasekaran , Yu Hu , Andrzej Banburski-Fahey , Jaron Lanier
- URL: https://arxiv.org/abs/2609.33010
- Abstract:
LLMs are effective representations that assimilate vast amounts of knowledge during pretraining, but post-training is necessary for models to reliably access this knowledge and “know what they know.” We observe an interesting rank equilibrium between knowledge stored in the weights and the data stream passing through the model. Across all model layers, we find that the hidden states (data stream) follow a U-shaped pattern, showing substantial compression in early layers and a steep rise during the late-layer decoding phase. In contrast, the weight rank follows an inverted U-shaped pattern, with very low rank in the early and late layers and high rank in the middle. We interpret this as an in-and-out information interplay: intermediate activations do not need to carry content that the weights can supply later, so they primarily preserve what the weights cannot provide. Motivated by this observation, we propose a model-aware data selection method, CAP (Counterfactual Assimilation Profile), which can determine whether a data candidate contains information accessible to the current model by utilizing the divergence gap in early- and late-layer representations between model-generated and reference responses. Across math, code, and science domains, CAP delivers 35.4% greater average improvement over the base model than the strongest baseline under different selection budgets. With only 10% of the data pool, CAP surpasses or matches full-pool training on math and science. We further show that CAP transfers to multimodal data selection and is robust to response horizon and noise.
294. X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
- Authors: Sitao Cheng , Xunjian Yin , Zhiyuan Sun , Yuxuan Li , Ruiwen Zhou , Xiangru Jian , Victor Zhong
- URL: https://arxiv.org/abs/2609.32993
- Abstract:
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less fully than its content allows. Recent agents do use that structure, but only as LLM-written skills in context, never in the weights, so their gains do not generalize beyond retrieval. We instead recover this hierarchy from the data itself and train on it, with no LLM calls. Following text tokenizers, which build a vocabulary by counting alone, we score action spans by reusability and merge canonicalized actions into a reusable eXperience tree (X-Tree). Each X-Tree node captures how a frequent and success-bearing skill is composed from sub-skills, guiding efficient generalization. We integrate X-Tree into three training settings: offline RL, with each node as a training instance; online RLVR, with an adaptive skill bonus; and on-policy self-distillation, with X-Tree as the self-teacher’s privileged context. Across WebArena, ScienceWorld, and WebShop at three model scales, X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop. Matched analyses attribute the gains to the X-Tree structure and the three integrations.
295. Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization
- Authors: Yifan Wang , Hao Cheng , Xiaomin Li , Yuexing Hao , Hemanth Neelgund Ramesh , Dongwon Jung , Hao Tang , Keru Wang , Chenliang Zhou , Qianhui Wu , Wenlin Yao , Ananth Grama , Andrzej Banburski-Fahey , Baolin Peng , Jaron Lanier , Jianfeng Gao
- URL: https://arxiv.org/abs/2609.32990
- Abstract:
Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite being an ideal testbed for long-horizon adaptation. In this setting, agents are required to solve streams of sequential tasks, navigate complex dependencies with evolving repositories and persistently store and reuse experience. Text-based skill optimization offers an efficient, non-parametric approach for such adaptation. However, existing methods often suffer from unstable updates, performance drawdown, and agent collapse over extended deployments. In this paper, we formalize the concept of in-context self-evolution and introduce VALVE, a validated-gated framework for long-horizon skill optimization. We establish finite convergence, provide theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a prescribed tolerance, with leading-order scaling Empirically, our pipeline, VALVE achieves stable self-improvement over evolution horizon spanning more than 1,000 SWE tasks, with average final and peak gains of $14.9$ and $16.5$ points across three frontier models (GPT-5.5, Claude-4.6 and MiniMax-M2.7). The validation gate reduces average drawdown by 75% and produces an 11x more compact skill bank than ungated evolution. We further present extensive ablations identifying the design choices most critical to long-horizon skill evolution.
296. Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
- Authors: Hongyi Du , Tianyi Zhang , Weijia Zhang , Yi Yang , Haofei Yu , Kunlun Zhu , Tianxiang Dai , Shang Jiang , Zhelun Gao , Jiaxin Pei , Shang Zhu , Jiaxuan You
- URL: https://arxiv.org/abs/2609.32965
- Abstract:
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%), establishing the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.
297. The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
- Authors: Vy Nguyen , Ziqi Xu , Jeffrey Chan , Estrid He , Feng Xia , Renqiang Luo , Erik Cambria , Xiuzhen Zhang
- URL: https://arxiv.org/abs/2609.32964
- Abstract:
Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model’s intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
298. Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence
- Authors: Xiao Yue , Guangzhi Qu
- URL: https://arxiv.org/abs/2609.32924
- Abstract:
Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Coverage is answer availability, realization is readout accuracy on the frozen candidate pool, and validity evidence is a label-blinded critic judgment of derivational support rather than a proof certificate. CRV freezes each candidate pool before comparing readouts and analyzes covered failures by correct-answer multiplicity and within-problem discrimination. On HardShift441, a 441-problem set for which a reference solver leaves 406 problems unsolved, a LoRA-adapted Qwen2.5-7B generator obtains 24.2% average single-sample accuracy and 68.9% pass@16, whereas verifier-weighted self-consistency (WSC) reaches 38.0%. Readout accuracy is particularly low when the correct answer occurs only once or twice in the pool. In a separate constructed audit of 195 covered problems, the critic labels 12 correct-answer representatives as supported, 181 as refuted, and two as uncertain. These results show that coverage, realization, and validity evidence from the critic are distinct quantities and should be reported separately.
299. TRACE: Learning to Self-Calibrate Wireless Digital Twins from ISAC Measurements
- Authors: Saad Masrur , Saeed R. Khosravirad , Ismail Guvenc
- URL: https://arxiv.org/abs/2609.32923
- Abstract:
Wireless digital twins (DTs) rely on 3D environment models to predict radio propagation and support wireless-network decisions, yet these models are often initialized from imperfect 3D maps. Errors in building position, height, footprint, and orientation can therefore cause a high-fidelity propagation engine to simulate the wrong physical environment. In this paper, we study how a deployed wireless network can repair an existing DT using its own radio frequency (RF) measurements. In particular, we introduce Twin Residual Alignment and Calibration Engine (TRACE), a physics-grounded learning-based self-calibration framework that treats twin maintenance as residual alignment between the physical world and the current DT. Using the same sensing configuration as the physical measurements, TRACE ray-traces the current DT, coherently backprojects the measured and simulated RF onto a common world grid, and extracts the same local region around each building’s current DT position. A multi-view corrector then fuses evidence across sensing nodes and neighboring buildings to predict a gated six-parameter correction per building, without relying on absolute layout or sensor ordering, and supports iterative correction through re-rendering. On 5,400 held-out samples from unseen simulated scenes at 28 GHz, TRACE reduces 3D position RMSE from 2.202 m to 0.302 m and yaw RMSE from 4.978° to 0.894°, outperforming ViT and U-Net baselines under changes in layout, building count, sensing-node count, and SNR. On measured 28 GHz RF data from the NIST outdoor courtyard, a model trained only on synthetic RF reduces mean planar wall-position error from 1.00 m to 7.8 cm, without measured-data fine-tuning or geometric labels. These results show that the discrepancy between measured and twin-rendered RF can serve as a learning signal for repairing a wireless DT.
300. Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer
- Authors: Haoran Jin , Kangqi Zhang , Jirong Yang , Barry Lyu , Qiuyi Ding , Ruijie Gao , Nathan Bleier
- URL: https://arxiv.org/abs/2609.32922
- Abstract:
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerators. We revisit stochastic computing (SC) as a way to lift this constraint: viewed as a dense adaptive quantizer, SC controls precision by bit-stream length L rather than a fixed datapath, while each multiplication reduces to a single AND/XNOR gate. We build a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning. On top of this substrate, we develop a dynamic per-row mixed-precision policy that assigns stream length per token or group at matched average budget, requires no retraining, and uses the same SC hardware across schedules. Across tasks, SC remains competitive with fixed-format INT quantization at matched bit budgets, while per-row mixed precision helps maintain accuracy at lower average stream lengths. These results provide software-level feasibility evidence that SC can serve as a dense-precision substrate for fine-grained mixed-precision inference on modern vision transformers.
301. Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows
- Authors: Vivek Kumar Singh , Preeti Priyam , Gautam Bhowmick
- URL: https://arxiv.org/abs/2609.32917
- Abstract:
Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks this from a different angle. Instead of leaving model-tier selection to some component downstream, it folds the choice into planning itself. As the planner breaks a query into subtasks, it also assigns each one a model size tier (small, mid, or frontier, ordered by capability and price), so the dependencies between subtasks are visible before any specialist runs. Unlike per-call routers such as cascade routing, which look at one node at a time, PaR sees the whole workflow up front and needs no separate router model or training data. We evaluate PaR with EntBench, a benchmark of 54 enterprise agentic tasks across seven classes, graded by actually running the generated Structured Query Language (SQL) and MongoDB queries against live databases. Over 1,157 evaluations spanning eight routers and three seeds, PaR stays on the observed cost-accuracy frontier. It matches a sink-frontier heuristic (frontier model on terminal nodes only) in accuracy at comparable cost and a faithful FrugalGPT cascade at lower cost, and cuts cost 44% against all-frontier routing while giving up 2.9 points of accuracy. Several accuracy gaps fall inside the plus-or-minus six-point confidence interval of a 54-task study, so we frame PaR’s advantage as frontier position rather than a clean accuracy win. We also report a preliminary observation, not a validated result: a small pilot hints that cheap routing may carry a hidden compounding penalty on compositional workflows, which we frame as a hypothesis for future measurement. PaR, EntBench, and all evaluation code are open source.
302. Logical subspace in LLMs
- Authors: Hope Kean , Enric Boix-Adsera
- URL: https://arxiv.org/abs/2609.32907
- Abstract:
Recent work has identified a human brain network specialized for abstract formal reasoning (Kean et al., 2025). Does the same hold true in language models? To answer this question, we introduce the minimal viable subspace (MVS) method, which searches for the lowest-rank activation subspace at a layer that preserves task performance when everything outside that subspace is ablated. Using MVS, we demonstrate low-rank subspaces supporting logical inference on Gemma and Qwen models. Furthermore, these subspaces exhibit a clear dissociation from model capacities on other tasks, such that retaining these late logic subspaces preserves inference while impairing factual knowledge, working memory, cognitive control, and arithmetic. Conversely, ablating them reduces logical inference accuracy to chance while largely sparing these other capacities. Our results suggest a functionally localizable core machinery for logic akin to that in the human brain.
303. Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models
- Authors: Jianchang Su , Wei Zhang
- URL: https://arxiv.org/abs/2609.32900
- Abstract:
Diffusion language models (dLLMs) predict masked positions in arbitrary order, but their exact constrained decoders still encode constraints as sequential languages, whose state must track every unresolved dependency between positions. For relational constraints this encoding grows exponentially: for same-order copy, every finite automaton needs $4^k$ states, deterministic or nondeterministic, and every context-free grammar has size $2^{\Omega(k)}$, while the factor graph of the same relation has size $O(k)$ and a 16-entry peak table. We introduce FactorDLM, a training-free decoder that represents finite-domain relations as a factor graph and, at each denoising step, conditions the model’s mean-field prediction on that graph exactly by variable elimination. Decoding cost then grows exponentially with the induced width of the constraint graph, which replaces automaton size as the governing parameter. Because a finite automaton is a chain-shaped factor graph, one compiler enforces syntax and nonlocal relations together: on JSON records with cross-field references, a schema automaton alone leaves references dangling, relational factors alone produce malformed JSON, and the combined plan is valid on both counts, including on records of variable length. Across nine relational benchmarks and three backbones, every output satisfies every declared constraint at 0.4-6.9% projection overhead, where unconstrained decoding is 0-79% valid, and compiled projection answers repeated queries 13.6x faster than CP-SAT with eight parallel workers. Because model-free rules solve three of five standard benchmarks, we construct benchmarks with exact chance and fixed-template floors, on which selecting among exact constrained samples beats greedy projection. Which encoding is cheaper, sequential state or direct factors, depends on the constraint and is computable before decoding begins.
304. StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills
- Authors: Zeping Liu , Yan Li , Ni Lao , Gil Wolff , Gengchen Mai
- URL: https://arxiv.org/abs/2609.32886
- Abstract:
Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at this https URL .
305. Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets
- Authors: Yuanbo Li , Zekun Li , Xiaoyan cong
- URL: https://arxiv.org/abs/2609.32885
- Abstract:
We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protocol. We assess Brier scores relative to an empirical base-rate predictor, examine their Murphy decomposition, and compare models through paired differences, with results stratified by event category and timing relative to training cutoffs. In the reported post-cutoff stratum, the four Claude models achieve Brier scores of 0.183-0.192, improving on the base-rate reference by 0.024-0.033. Qwen 32B does not significantly outperform that reference, although its paired Brier is 0.024 lower than that of the 7B checkpoint. The evaluated Claude version and tier upgrades yield no significant improvement. Within-model differences across event categories exceed the observed differences among Claude variants. These results show why forecasting scores should be interpreted alongside simple probability baselines and question composition: under this protocol, newer versions or higher model tiers do not consistently produce more accurate probabilities.
306. Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
- Authors: Xing Han , Yuxin Wang , Chen Chen , Wei Dai , Gautham Krishna Gudur , Shijun Li , Hsing-Huan Chung , Gregory D. Hager , Joydeep Ghosh , Paul Pu Liang , Suchi Saria
- URL: https://arxiv.org/abs/2609.32870
- Abstract:
Self-play proposer–solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action–outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.
307. FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors
- Authors: Jerry Huang , Sarvesh Babu , Matt Van Buren , Alexander Wang , Pranav Pillai , Arush Jain , James P. Burton , Julia Hockenmaier
- URL: https://arxiv.org/abs/2609.32835
- Abstract:
As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
308. Improving LLM Collaboration via Multi-Agent Preference Learning
- Authors: Shuo Liu , Xinzichen Li , Tianle Chen , Christopher Amato
- URL: https://arxiv.org/abs/2609.32827
- Abstract:
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI feedback. Yet, its extension to multi-agent systems remains underexplored. To address this gap, we formulate preference-based multi-agent systems (MAS) from decentralized and centralized collaboration perspectives. We also introduce a general multi-agent preference learning framework (MAPL) to solve these problems. MAPL allows iterative updates by comparing the current solution with decentralized or centralized solutions generated by various agents. We instantiate MAPL using MARL from human feedback (MARLHF) with a learned reward model and multi-agent direct preference optimization (MADPO). Experiments on collaborative writing, coding, tool use, and travel planning show that MAPL can improve collaboration quality and efficiency while approaching the performance of MARL with fixed, well-defined rewards. Within MAPL, MARLHF generally outperforms MADPO on most tasks but remains sensitive to data coverage, agent and comparator models, and the underlying MARL algorithms.
309. The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces
- Authors: Tianqi Bu , YuXuan Peng , Junteng Tu , Henghui Xiao
- URL: https://arxiv.org/abs/2609.32825
- Abstract:
A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still see the original problem, and call the accuracy difference the decomposition tax. Across 21 open-weight models from nine organisations, on GSM-Hard and MATH-500 at n = 200 paired items per cell, 70 of 118 primary-family tests survive Benjamini-Hochberg correction and 54 survive Holm. On GSM-Hard, a placebo recovers nothing: it carries at least 60% of the extra tokens and at most one word of the problem. Builders design a pipeline one stage at a time, and its bill arrives at the interfaces between stages. Rewriting one stage’s instruction moves gemma-3-12B’s tax from 4.5 to 36.5 points, and adding “every relationship stated between them” to a stage that lists the numerical quantities lowers the tax on 9 of 9 models on MATH-500. Re-grounding, which shows a stage the original problem again, belongs after the loss. With one lossy interface, re-grounding the stage after it beats re-grounding the stage before it on 7 of 7 models on both benchmarks; on MATH-500 the earlier repair is worse than none on 7 of 7. Newer models still pay: gemma-4-12B gives up 37.0 points, and the repair holds on all three of the newest models we test. A sealed held-out test refuted a stronger rule we registered, which predicted the paying stage from the interface and receiver types, so we locate the tax by measuring one stage at a time. The prescription has two parts: re-ground the stage after the lossy interface, and if a stage must list the quantities, tell it to keep the relationships.
310. Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
- Authors: Yuanyi Wang , Yanggan Gu , Su Lu , Guanghao Zhu , Pengkai Wang , Yifan Yang , Congkai Xie , Zhaoyi Yan , Jianmin Wu , Hongxia Yang
- URL: https://arxiv.org/abs/2609.32821
- Abstract:
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair?} We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as \textit{task loss recoverable under a specified routing intervention, with non-routing parameters fixed.} These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose \emph{Selective Router Repair (SRR)} as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that \textbf{routing drift alone is insufficient evidence of routing failure}: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.
311. Rank Collapse Is Recoverable, Growing $|Q|$ Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic
- Authors: Tianqi Bu , YuXuan Peng , Junteng Tu , Henghui Xiao
- URL: https://arxiv.org/abs/2609.32819
- Abstract:
Raising the update-to-data (UTD) ratio breaks off-policy critics in two ways grouped as “plasticity loss”: collapsing representations and growing value magnitude $ Q $. We separate them in Soft Actor-Critic (SAC) with scaled critics (width 2048, no normalization). Collapse is survivable: at UTD ratio 16, HalfCheetah critics with most units dormant keep learning, and the training guard, which stops runs whose loss or $ Q $ explodes, never flags them. Within one high-UTD SAC configuration, runs start close together, and how far a critic’s $\log_{10} Q $ has climbed by step 15k, its early growth, ranks the runs by how soon the guard flags them. At 15k, a flagged run’s $ Q $ sits a median of over a hundredfold below its flag level, yet the climb’s rate already orders the flags (Harrell’s C and out-of-sample AUC 0.78 on Walker2d, 0.98 on Ant, at UTD ratio 4). Dormancy does not. Aborting on this rate saves about a tenth of held-out Walker2d compute and stays net-positive live. A LayerNorm critic lowers the rate, removes the flag on Walker2d at UTD ratio 4 and lowers the return.
312. When Can First-Order Models of Fine-Tuning Bound Forgetting?
- Authors: Jianchang Su , Wei Zhang
- URL: https://arxiv.org/abs/2609.32818
- Abstract:
Fine-tuning a language model on new data can make it forget facts that it should keep. We ask whether measurements taken at the start of a fine-tuning run can bound, for each protected fact, the probability that the run makes the model forget it. In LoRA fine-tuning with stochastic gradient descent on models from 0.6B to 14B parameters, a first-order response model estimated by finite-difference probes predicts changes of per-fact margins with correlation 0.974-0.998. Predictions of forgetting built on this model nevertheless failed, because forgetting requires parameter changes far outside the region in which the model was validated. The probes can, however, bound the probability that a margin first falls below a boundary near zero: we derive Freedman and Azuma first-passage bounds for a linear surrogate of the margin and test on new runs whether they hold for the model. The bounds contain a term R that measures how much the response coefficients change during the run. The simplified Freedman bound, which sets R = 0, certified most facts but was violated in 14 of 112 conditions, and every fact on which it was violated had R >= a, where a is the distance of the fact’s margin to the boundary. The complete Freedman bound certifies only facts with R < a, and it held in every condition. On the violated facts, the spread of the margin across test runs was a median of 14.6 times the prediction of the response model, so the failures are breakdowns of the model, and in our data they occurred only where R >= a. We found this pattern post hoc and tested it in two preregistered confirmatory studies with 43 new conditions: the complete bound held in all of them, and the simplified bound failed there on only 3 facts, each with R >= a. First-order models of fine-tuning can thus bound forgetting on the facts whose response coefficients change by less than their distance to the boundary.
313. Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support
- Authors: Bharath Kumar Bolla , Bharath Kumar Bolla , Vishnu Surya Reddy Nandi
- URL: https://arxiv.org/abs/2609.32817
- Abstract:
Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation Stability Index (ESI), which measures reasoning consistency across repeated queries; the Causal Faithfulness Score (CFS), which tests whether cited concepts drive predictions via concept ablation; and the Perturbation Stability Score (PSS), which measures robustness to semantic-preserving paraphrases. By evaluating six LLMs on 150 MedQA-USMLE questions (900 model-question observations), we found that only 23.3% of the cited clinical concepts were causally necessary. Correct answers had lower CFS than incorrect answers (0.212 vs. 0.398), answer consistency negatively predicted CFS (Spearman r = -0.466), and model pairs could agree on answers while sharing only 8.8% of cited reasoning concepts. These results show that accuracy, consistency, and consensus are incomplete safety signals for clinical decision-making support. The evidence is behavioral rather than mechanistic: concept ablation tests counterfactual sensitivity of outputs, not internal circuits.
314. Overwhelmed by Choice: Studying LLM Decision Making at Scale
- Authors: Yu-Chi Lin , Aryan Seth , Anshul Aravind , Eugene Lee , Tanmay Parekh , Nanyun Peng , Kai-Wei Chang
- URL: https://arxiv.org/abs/2609.32809
- Abstract:
Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.
315. Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs
- Authors: Chaitai Deb Purkayastha , Bharath Kumar Bolla , Vishnu Surya Reddy Nandi
- URL: https://arxiv.org/abs/2609.32807
- Abstract:
Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CFV) test asked each model to make a minimal, plausible clinical change that would make a different answer correct. The demographic robustness test added six demographic prefixes to the same vignette and compared the answers and explanations with a no demographic baseline. Of the 900 CFV attempts, 228 (25.3 %) were valid and 672 were invalid. Across 5,400 demographic comparisons, 1,097 answers were changed (20.3%). Automated judging identified 3,128 stereotype evidence flags, including 1,932 in the broad Other category. MedGemma 27B achieved the highest accuracy (87.1%) and CFV (63.3%), lowest answer change rate (16.0%), and low mean Explanation Demographic Dissonance (EDD) score (0.169). However, its accuracy still exceeded its CFV, indicating that correct answers do not guarantee reliable performance on the counterfactual validity task. OpenBioLLM had the highest answer change rate and EDD, whereas GLM had the highest stereotype flag rate. These findings show that accuracy, CFV, answer stability, EDD, and stereotype evidence capture different evaluation aspects. Because all judgments were automated and no clinician validation was available, the results support safety screening but do not establish clinical deployability of the model.
316. Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret
- Authors: Bingyu Shen , Boyang Li
- URL: https://arxiv.org/abs/2609.32805
- Abstract:
Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. Steps stay cheap, but anything the writer drops is lost before later decisions reveal that they need it. We quantify this loss and ask whether training can reduce it. Comparing the written state with the best state of the same size written in hindsight, we split the reader’s loss into a budget loss, which any state of that size must incur, and a write-time regret, which comes from the writer’s choices. In TextWorld cooking games where we control how long a fact must be carried before it is needed, a 128-token state holding the facts wins nearly every game, while prompted language-model writers win at most 17%. Almost all of the loss is write-time regret, and it grows with the delay. We then train the writer from the reader’s own loss. DSSR (decision-sufficient state representations) scores candidate states by how well the reader acts after the writer carries them forward, and teaches the writer to prefer the better ones. This forward-rolled score predicts game outcomes ($\rho = 0.48$), whereas scoring a candidate as a fixed context, as hindsight methods usually do, does not ($\rho \leq 0.07$). On a pre-registered test split opened once, training adds +7.0 [+1.9, +12.2] points of success when facts are needed soon, bringing a plain summary writer to the level of belief- and slot-based memory prompts. The gain shrinks as the delay grows and is significant only at the shortest delay. We trace this limit to credit assignment: keeping a fact now pays off only if every later rewrite keeps it too, which a per-step score cannot see.
317. Nutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutrition
- Authors: Uttej Kallakuri , Boxun Hu , Ankur A. Butala , Najim Dehak , Tinoosh Mohsenin
- URL: https://arxiv.org/abs/2609.32803
- Abstract:
Generative and Agentic IoT systems offer a promising foundation for digital healthcare applications that combine sensing, personalized reasoning, and autonomous interaction in real-world environments. Nutrition assistance is a natural use case, but existing Large Language Model (LLM)-based systems are often limited to passive text interaction and static context, making them unreliable when food descriptions are ambiguous or nutritional evidence is missing. We propose Nutri-ATLAS, an Embodied Agent for Tabulated Lookup and Assistance for smarter nutrition in the real world. It integrates graph-grounded nutrition reasoning, hardware-aware LLM selection, and robot-based evidence acquisition. Nutri-ATLAS builds a unified Food-Nutrient knowledge graph from USDA FoodData Central and FoodKG and learns 64-dimensional GATv2 food and recipe embeddings. A shared hybrid graph-text scoring mechanism supports food nutrition extraction, nutritional gap filling, substitute retrieval, and recipe-level meal composition, while an LLM-guided skill interface navigates landmarks, updates dietary-context and food-accessibility memory, and grounds recommendations in observed food availability. We evaluate Nutri-ATLAS across nutrient estimation, substitution retrieval, recipe recommendation, patient-profile adherence, edge deployment, and real-world embodied execution. On HealthyFoodSubs, the hybrid retriever achieves 37.9% MAP, 80.7% RR@5, and 90.1% RR@10. On NutriBench v2, Dense+GAT retrieval grounds nutrient estimation across nine quantized Qwen3.5-9B configurations. On PFoodReQ, Nutri-ATLAS reaches 78.8% MAP, 83.0% MAR, and 77.5% F1. A patient-profile study shows adherence to allergy and healthy-target constraints for all selected cases.
318. Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
- Authors: Tianqi Bu , YuXuan Peng , Junteng Tu , Henghui Xiao
- URL: https://arxiv.org/abs/2609.32802
- Abstract:
One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-weight backbones served with thinking disabled, and on the two Qwen backbones the blind inspector changes no item at all. That head-to-head is exploratory. One deterministic fault enters the first stage, and we re-expose the original problem to $k = 0,\dots,3$ of the downstream stages with agents, items, fault and topology held fixed, on 120 gsm_hard items per arm at temperature zero. Accuracy under fault rises on four of four backbones, from +0.233 to +0.392, the largest Holm-adjusted $p$ being $2.1\times10^{-6}$. A registered kill test rules out tokens. Blanking every word holds the word slots fixed, and retention tracks the visible fraction on four of four, climbing from 0.221 to 0.692 on the primary. Those two families are confirmatory and everything else here is exploratory. The interaction excludes zero on two of four backbones under the registered pipeline, four of four under a three-stage pipeline, and three of four under full message history, the primary at +0.317. On Llama-3.1-8B the fault carries no detectable cost at any dose, so the other three carry every claim about what a fault costs. Re-derivability also sets what the architecture costs, and no decomposition we measured reliably beats one direct call. With no fault injected the registered pipeline loses to that call by -0.267, -0.125 and -0.317, and on Phi-4 reads +0.058 at $p = 0.118$, which the test fails to separate from zero. The repair that works is cheap and front-loaded: the first re-grounded stage buys +0.394 of matched retention for +59.8 tokens per item on Qwen3-14B, and the stages after it buy nothing.
319. PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents
- Authors: Junchi Chen , Changtao Miao , Yuxiao Xiang , Zhenchao Jin , Haojie Yuan , Qi Chu , Tao Gong , He Liu , Bo Zhang , Jiansheng Cai , Zhe Li , Nenghai Yu
- URL: https://arxiv.org/abs/2609.32801
- Abstract:
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors assess subtasks in isolation. To address this gap, we introduce PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment. For training and evaluation, we construct a Multi-Step Plan Safety (MSP-Safe) dataset through paired task construction, plan generation using diverse planners, and safety annotation by three judges. Task-oriented SFT on MSP-Safe establishes fundamental plan-safety assessment capabilities, yet a substantial gap remains between compact models suitable for real-time deployment and stronger but costlier large models. Accordingly, we propose Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision along their on-policy trajectories. It combines token-level distribution transfer from a fine-tuned strong teacher with probability-routed sequence-level compensation, retaining student-generated targets when the student favors the reference safety decision and using teacher-reconstructed targets otherwise. Across all test subsets, PlanGuard-2B achieves average 87.15% ACC and 87.21% F1, demonstrating effective whole-plan physical-risk detection at compact model scale. Code and dataset will be publicly released.
320. AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks
- Authors: Woojung Song , Hoyeol Yang , Jeonghoon Shim , Sungjib Lim , Jonggeun Lee , Yunho Choi , Yohan Jo
- URL: https://arxiv.org/abs/2609.32795
- Abstract:
Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users’ preferences and needs. For example, agents differ in whether they ask clarifying questions or search the web. We introduce HABIT, a taxonomy of 23 behavioral axes in five categories, which three authors and three LLMs derive bottom-up from 408 agent trajectories across 17 domains. On held-out tasks, HABIT distinguishes models more clearly than existing taxonomies of human values and agent actions while supporting comparably consistent annotation. Building on HABIT, we construct AgentHABIT, a benchmark that profiles each agent’s behavioral tendencies from its trajectories on 86 everyday tasks. Profiling 18 models with AgentHABIT reveals a range of distinctive tendencies. For example, most GPT and Claude models state their assumptions and offer alternatives when requirements conflict, whereas Qwen and Google’s models more often leave assumptions or changes to requirements unstated. These profiles remain recognizable even when built from entirely different sets of tasks, indicating that they reflect general tendencies rather than task-specific behavior. Prompting agents to adopt specific behaviors shifts some axes readily but barely changes others, while fine-tuning on another model’s trajectories changes only part of a model’s profile and leaves much of it intact. Overall, HABIT and AgentHABIT provide a systematic framework for characterizing how agents carry out everyday tasks beyond task success, offering insights to guide the development of agents whose behavior better fits users’ needs.
321. $T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
- Authors: Nan Qiao , Yebin Yang , Weinong Wang , Shuning Wang , Shangpin Peng , Fengyuan Lu , Xinming Wang , Zhehan Kan , Ruixu Zhang , Songyang Zhang , Sheng Yue , Yonglong Tian , Ju Ren
- URL: https://arxiv.org/abs/2609.32791
- Abstract:
Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training–inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.
322. CLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form Completion
- Authors: Garapati Keerthana , Manik Gupta
- URL: https://arxiv.org/abs/2609.32787
- Abstract:
Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for Reasoning and Entry), a hybrid workflow that separates field-state discovery, source-to-field mapping, deterministic validation, bounded correction, escalation, and audit tracing. We tested five synthetic healthcare administrative schemas, 1,000 source records, four interface variants, two data-quality suites, and six comparators, yielding 24,000 benchmark episodes. A separate strict-output audit evaluated direct mappings from Qwen2.5-1.5B and Qwen2.5-7B, and a trace-derived operational simulation covered 6,000 episodes. Under the evaluated synthetic benchmark conditions, full CLAIRE achieved 1.000 episode success, field accuracy, required-field completion, and dependency completion in both suites; removing validation reduced stress-suite success to 0.500. In the simulation, 100.0% of clean and validation-stress episodes reached a staff-reviewable draft, compared with 68.6% of escalation challenge episodes, unsupported cases were blocked. Scenario-based savings were 149.7-165.5 seconds per case, not observed staff times. The findings support schema-grounded, validation-first healthcare administrative automation in which language-model components assist mapping but do not authorize unsupported or consequential actions.
323. Learning response-aware patient dynamics for respiratory support
- Authors: Xiaolei Lu , Shamim Nemati
- URL: https://arxiv.org/abs/2609.32782
- Abstract:
Respiratory support can shape the short-term physiological trajectory of critically ill patients, but patients receiving the same intervention may follow different physiological trajectories. Clinical patient dynamics models typically predict future states from recent physiology and recorded interventions, while physiological change is mainly represented through the predicted future state. We propose a response-aware patient dynamics model that explicitly represents physiological change during autoregressive state updating. The model decomposes predicted physiological change into state-dependent baseline dynamics and respiratory-support-associated deviations, with room air providing a reference for the decomposition. We provide a formal analysis of this reference-anchored formulation. A response pathway encodes the predicted physiological change and uses it to update the latent patient state across the forecast horizon. Across ICU cohorts from two independent institutions, the proposed model achieves comparable overall trajectory prediction to patient dynamics baselines, with more consistent improvements when physiological states are changing.
324. Agentic Network Traffic Monitoring
- Authors: Manuel Tsoukatos , Hayden Jananthan , Jeremy Kepner
- URL: https://arxiv.org/abs/2609.32778
- Abstract:
As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way that is aligned with the users intent. Auditing an agent’s network traffic provides a clear record of the agent interactions. This work presents a novel approach to monitoring the network traffic of agentic systems using complex valued hypersparse traffic matrices by integrating DBOS (DataBase OS), the OneSparse PostgreSQL database, and the GraphBLAS math library. To develop these concepts an agentic simulator was constructed, allowing a varying numbers of AI agents to collectively survey a virtual environment using different strategies. The resulting network traffic matrices enable easy monitoring of the AI agents.
325. Forecasting Intraday USD/CAD Exchange Rate with News-Derived Monetary-Policy Signals
- Authors: Maya Kodeih , Aliaa Alnaggar , Mucahit Cevik
- URL: https://arxiv.org/abs/2609.32773
- Abstract:
Monetary-policy announcements and central-bank communications play a central role in foreign exchange markets, yet their qualitative, unstructured form makes their forecasting value difficult to quantify. While prior research has largely focused on sentiment extracted from financial news, comparatively little is known about the relative contribution of different dimensions of monetary-policy communication. Existing studies primarily evaluate whether textual information improves overall forecasting performance but provide limited insight into which communication channels drive such improvements. To address this gap, this paper introduces a statistical attribution methodology that decomposes monetary-policy communication into interpretable channels and quantifies their incremental forecasting contribution under false-discovery-rate control. Monetary-policy news is transformed into structured communication signals using large language models (LLMs) and temporal feature engineering. These signals are evaluated using rolling-window experiments with tree-based machine-learning models. The results show that monetary-policy communication contains measurable predictive information. Attribution analysis shows that predictive value is concentrated in a small subset of signals, with communication timing providing the strongest individual feature-level contribution, targeted communication-activity measures also contributing positively, and LLM-derived sentiment providing complementary information at the group level. The findings indicate that communication-based forecasting value extends beyond sentiment alone and that attribution, rather than aggregate accuracy alone, is central to evaluating news-derived signals.
326. Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them
- Authors: Yicheng Bao , Zhenkun Gao , Xiahui Guo , Mingqian Yang , Xueheng Li , Bangwei Liu , Mingang Chen , Lijun Li , Xuhong Wang , Xin Tan
- URL: https://arxiv.org/abs/2609.32763
- Abstract:
Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts, image-caption inconsistencies, visual implausibilities, or external references, and therefore do not test whether a model can use its own world knowledge to verify a recognized image. We introduce Mandela-Bench, containing 1,507 edits of canonical images: 1,359 knowledge-only forgeries, each contradicting one verifiable fact, and 148 anchor-free controls that preserve the editing process without introducing a factual contradiction, together with 474 untouched originals. We score not only whether a model detects a forgery, but whether its explanation identifies the inserted entity or the fact being violated. Across 36 multimodal models, from 0.8B parameters to frontier scale, we find a consistent failure mode. When a public figure is removed from a familiar photograph, models still name that person in up to 72.7% of responses. Some models can distinguish the replacement face from the original when shown in isolation, yet still judge the full edited photograph as authentic. Providing the true event and date does not improve knowledge-grounded detection, whereas providing the same information after cropping away the recognizable composition does. Even under explicit verification prompts, only one of the 36 models meets the KGR criterion on at least half of the forged images. These results suggest that the failures cannot be explained by missing knowledge or inadequate perception alone. Instead, they are consistent with recognition biasing verification toward the remembered canonical image rather than the observed edit.
327. Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models
- Authors: Drandreb Earl Juanico
- URL: https://arxiv.org/abs/2609.32757
- Abstract:
VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Qwen3-VL-4B-Instruct on single-object COCO grounding. We compare clean coordinate-token readout rankings with corruption-derived repair rankings, using object-mask endpoint replacement for recovery and clean-input flooring for depth localization. In Qwen3-VL, coordinate-token rankings are inert through layer 24, load-bearing from layers 32-35, and peak at layer 34; corruption-derived rankings harm layers 16-24 but become beneficial near layer 35/final. A Kimi-VL-A3B diagnostic shows a matching output-proximal transition despite a different box format. Object-mask recovery separates rank budgets: $k=250$ shows necessity, $k=500$ shows Top-$k$ restoration above random, and $k=d/2$ is largely capacity-driven. Partial-occlusion sweeps reveal that high-overlap coordinate-token sets can hurt at $k=1000$ and help mainly at half-width, while population corruption-derived sets provide no reliable fixed repair set. Edge-attribution patching shows coordinate-token paths are high precision but low recall for detection recovery, and RMSNorm quasi-layer controls do not close the endpoint-repair gap. Endpoint coordinate triage is therefore a useful circuit prior, but occlusion recovery requires a separate benchmark.
328. Adaptive Consistency Graph for Long-Horizon Agents
- Authors: Jiecong Wang , Hao Peng , Zhanyi Wang
- URL: https://arxiv.org/abs/2609.32754
- Abstract:
Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent’s planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna’s average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement.
329. Action Shaping: Policies Absorb What They Can Express
- Authors: Yanjun Chen , Jinghan Wang , Xiaoyu Shen , Wenjie Li , Wei Zhang
- URL: https://arxiv.org/abs/2609.32752
- Abstract:
Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset’s amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.
330. CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
- Authors: Xin Yan , Zhengbo Jiao , Jiaqi Liu , Zhenglin Wan , SiYuan Ma , Xuliang Yu , Tianyi Jiang , Chubin Zhang , Pengfei Zhou , Wangbo Zhao , Xingrui Yu , Bo An , Yang You , Ivor Tsang
- URL: https://arxiv.org/abs/2609.32750
- Abstract:
Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications. However, conventional deployments replicate an initialized runtime for each independent rollout, even when trajectories use the same software, incurring repeated memory and initialization costs as the number of parallel environments grows. Does an independent computer-use environment require an independent execution runtime? Our key observation is that trajectories require independent mutable state, while initialized application runtimes can be reused across concurrently evolving environments, making state the natural unit of environment independence. Guided by this observation, we introduce CUA-Sandbox, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators. Experiments show comparable or improved task success relative to Docker, while substantially reducing rollout and resource costs. CUA-Sandbox achieves up to a 6.20x increase in rollout throughput, a 9.2x reduction in per-environment memory, and a 504x reduction in incremental storage.
331. Retrospective Distillation Attribution via Normalized Response Similarity
- Authors: Minwoo Jang , Jaechang Kim , Minhyeon Oh , Jeongyeon Hwang , Jungseul Ok
- URL: https://arxiv.org/abs/2609.32749
- Abstract:
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring syntactic patterns into candidate profiles, filters low-contrast patterns, and calibrates student–candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated syntactic signatures along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
332. SkillVine: Agent Skill Evolution via Branching Exploration
- Authors: Kaiwei Liu , Jiqian Dong , Liran Dong , Shuai Mao , Mingming Zhao , Bufang Yang , Jie Chuai , Zhitang Chen , Guoliang Xing , Zhenyu Yan
- URL: https://arxiv.org/abs/2609.32731
- Abstract:
Agent skills encapsulate reusable procedural knowledge that enables LLM agents to perform tasks, and they can be improved automatically using trajectories from interactions with the environment. This is the classic problem of skill evolution. Existing approaches predominately follow a linear evolution paradigm, in which updates are sequentially applied to the latest skill-library version. As a result, they inevitably fall into local optima, leaving many promising evolution paths unexplored. We propose SkillVine, an automatic skill-evolution framework that formulates skill evolution as a graph search problem and employs a branching exploration strategy. Equipped with a trunk-branch collaborative searching mechanism, an intelligent parent-node selector, and an adaptive-granularity update rule, SkillVine achieves a balance between exploration and exploitation. We evaluate SkillVine on 5 benchmarks with two LLMs. Results show that SkillVine discovers better skill-library versions along branches than along the linear trunk and achieves the best test performance in nine of ten benchmark-model combinations.
333. MassAlloc Attention: Let Attention Allocate Its Own Compute
- Authors: Jingze Shi , Zhangyang Peng , Xianduo Li , Yanlin Qi , Xiaotian Lin , Haoxian Chen , Liangdong Wang , Guang Liu , Yuyu Luo
- URL: https://arxiv.org/abs/2609.32712
- Abstract:
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive retention of the work. MALA reduces low-contribution post-score computation. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens with tensor parallelism, MALA reduces forward and backward latency during training by 2.2x and 3.0x and decoding latency during inference by 1.6x relative to FullAttn. Across scaling-law training from 0.6B to 14B parameters, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.
334. CoWindow Attention: Full Causal Coverage Is a Collective Property
- Authors: Jingze Shi , Zhangyang Peng , Xianduo Li , Yanlin Qi , Xiaotian Lin , Haoxian Chen , Liangdong Wang , Guang Liu , Yuyu Luo
- URL: https://arxiv.org/abs/2609.32704
- Abstract:
FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.
335. Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time
- Authors: Jacob Dineen , Silei Ren , Muhao Chen , Dan Roth , Ben Zhou
- URL: https://arxiv.org/abs/2609.32701
- Abstract:
In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents generate their own free-form updates in a simulated incident-response task. Across ten independent games, pairs of GPT-5.6 Sol agents reach 98.8% final accuracy, compared with 25% chance, despite explicit instructions prohibiting disclosure and a monitor that screens each message without access to the agents’ interaction histories. The same interactions that help agents cooperate can therefore allow confidential information to pass through messages intended for legitimate coordination.
336. CAIRN: Dynamic Fact-Intent DAGs for Multi-Agent Exploration
- Authors: Zuyao Xu , Yuyang Jia , Junwei Guan , Xiang Li , Kaiwen Shen , Zhiqiang Dong
- URL: https://arxiv.org/abs/2609.32700
- Abstract:
LLM-powered autonomous systems have demonstrated promising capabilities in mathematical reasoning, engineering, and cybersecurity. Yet how to organize these systems for effective, reliable, and sustained performance remains an open question. In this paper, we present CAIRN, a fact-intent-driven multi-agent paradigm for goal-directed exploration. CAIRN represents observations and planned investigations as a dynamic directed acyclic graph (DAG). A reasoner interprets facts to propose intents, which workers execute to produce new facts. Each intent references its supporting facts and defines a potential exploration branch. The persistent graph preserves goals, dependencies and findings across workers, supporting knowledge reuse and parallel exploration. The graph also makes execution trajectories traceable and auditable, providing a basis for human verification and intervention. We evaluate CAIRN across cybersecurity and mathematical reasoning tasks, examining task success, time to solution, and token consumption. DAG-based coordination can incur higher token costs with no observable performance gains on tasks that require little effort. However, on high-effort tasks (at least 1M tokens), we observe faster solutions in 76.5% of cases, with speedups of up to 3.08x. Moreover, as task effort increases, these time gains become more pronounced while relative token overhead declines, highlighting the potential of DAG-guided parallel exploration.
337. Flat-Consensus Diffusion for Robust Data Reshaping under Noisy Evaluator
- Authors: Hongyu Cao , Kunpeng Liu , Fei Xie , Sandip Ray
- URL: https://arxiv.org/abs/2609.32696
- Abstract:
Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data reshaping: generating feature transformations that remain useful, stable, and reproducible under noisy evaluation and imperfect data conditions. We view reshaping operation sequence search as reward-guided diffusion generation, and robust reshaping as searching for regions in the latent reward landscape rather than isolated high-reward transformations. The key challenge is dual instability: noisy evaluators distort local reward guidance, while stochastic generative trajectories can converge to inconsistent solutions. We propose FCDiff, a flat-consensus diffusion framework that addresses both failures through a micro-macro decomposition. The micro layer replaces point-estimate reward guidance with Gaussian-smoothed, Monte Carlo averaged gradients, steering generation toward locally flat reward regions. The macro layer aggregates independently guided trajectories with a weighted Frechet-mean barycenter, selecting consensus-supported basins and filtering stochastic outliers. Across an 8-dataset headline cohort under heavy-tailed evaluator noise, FCDiff attains the best aggregate rank on lower-tail reliability and robustness against both search-based AutoFE and robustness-oriented generative baselines, with statistically significant accuracy gains over every generative baseline. Our results show that robust data reshaping requires searching for flat, consensus-supported regions rather than sharp single-trajectory optima.
338. IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents
- Authors: Angqing Jiang , Gaoming Zhang , Chaoqun Zhang , Jianchun Song , Liyuan Kong , Kena Qi , Wei Lin , Defu Lian
- URL: https://arxiv.org/abs/2609.32694
- Abstract:
On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student’s state. Existing methods either distill this preference directly or filter it with model-internal scores, but neither strategy verifies the query’s executed retrieval consequence. We propose Information-Gain-Gated Self-Distillation (IGSD), which verifies on-policy token proposals with environment feedback before distilling them. Treating each query token as a micro-action, IGSD completes the teacher’s token proposal and the student’s sampled token into matched queries and executes both from the same failed state with the same retriever. Shared counterfactual controls account for query-conditioned shifts in answer likelihood, so their difference, the executed paired information gain, provides a relative utility contrast for the retrieved documents. IGSD uses this contrast as a positive-only soft weight for candidate-pair distillation, while leaving the GRPO objective unchanged and confining verification to training. Across seven single-hop and multi-hop QA benchmarks, IGSD reaches macro-average exact-match accuracies of 42.8% and 47.0% with 3B and 7B policies, respectively, without inference-time verification. These results support environment-verified hindsight as an effective approach to reliable action-level supervision for search agents.
339. World Agent: Can Language Models Keep a World Running?
- Authors: Weixing Chen , Weipeng Zhang , Nan An , Yang Liu , Liang Lin
- URL: https://arxiv.org/abs/2609.32692
- Abstract:
World models are moving from generating realistic frames to generating playable worlds, yet whether a delivered world can keep running is not tested anywhere. Existing evaluations stop at generation, at delivery, or at single-step transitions, and each stops at a different point along the way. Correct local state transitions or intermediate outcomes do not guarantee a correctly organized causal event flow. We propose the world agent task, which moves the evaluation point of world generation from the moment of delivery to the continued operation that follows. In this task, a model is not asked to generate a world. It is held responsible for keeping the world running, which requires coordinating events and carrying forward their consequences to constrain subsequent evolution. We instantiate the task in WorldAgent-Benchmark with two complementary tracks. In the maintenance track, the model must ground the events of a continuous narrative into correct transitions of the explicit world state while respecting causal, temporal, and concurrency constraints. In the deduction track, the model must predict how the world will evolve under partial observations and act toward a goal. The maintenance track combines LLM-assisted semantic judgments with programmatic validation and scoring, while the deduction track is evaluated entirely programmatically. Individual judgments are auditable against world states and execution logs, and scores can be recomputed from the saved judgments and execution records. Across 8 models, scores decline steadily as pre-built structure is removed from the world, and causal-relation checking is the weakest component for every model. The benchmark makes the continued operation of a world measurable and distinguishes local completion from failures in event organization. Code and dataset will be released on this https URL .
340. Dude, Where’s My State? Execution Information Requirements for Stateful Agents
- Authors: Nikita Mehrotra , Ashish Tiwari , Priyanshu Gupta , Sumit Gulwani
- URL: https://arxiv.org/abs/2609.32687
- Abstract:
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.
341. PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks
- Authors: Xu Yang , Mingyang Yu , Jun Zhang , Keqian Li , Jing Xu
- URL: https://arxiv.org/abs/2609.32685
- Abstract:
Physics-informed neural networks (PINNs) provide a learning-based framework for solving partial differential equations (PDEs), yet their training behavior can change substantially throughout optimization. Residual distributions, gradient interactions, regional learning difficulty, and model-capacity requirements may evolve over time, while the network architecture and major training mechanisms are typically determined before training. We propose PINNMorph, an online PINN adaptation framework based on large language model (LLM)-guided policy evolution. PINNMorph maintains a population of state-conditioned adaptation policies that map execution diagnostics to controlled interventions over topology modification, additive representation augmentation, objective balancing, gradient handling, adaptive sampling, and optimizer-phase control. At each intervention opportunity, candidate programs are instantiated from the current policy population, selected according to the observed training state, and applied directly to the PINN under training. The resulting model inherits its existing parameters and training state and continues optimization along the same trajectory. Execution outcomes are subsequently used to evaluate interventions and evolve the policy population. Unlike pre-training architecture search or fixed adaptation rules, PINNMorph jointly adapts the current PINN and the policies governing its interventions using feedback from actual training. Experiments on 13 PDE benchmarks show that PINNMorph achieves lower solution errors than SA-PINN, ConFIG, RoPINN, HARMONIC, and PINNsAgent across all evaluated problems. Ablation studies further examine the effects of online adaptation, state-conditioned intervention selection, and execution-feedback-driven policy evolution.
342. When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
- Authors: Ke Wang , Zijie Zhao , Zhiyi Yuan , Changlun Li
- URL: https://arxiv.org/abs/2609.32677
- Abstract:
Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. We formalize this gap as Improvement Fidelity, which asks whether proxy improvements preserve the sign and ordering of deployment improvements over the updates an improvement process actually proposes. We show that global policy accuracy need not guarantee update fidelity: operator shift and deployment response can create update-level errors, while candidate margins determine whether those errors change the replacement decision. We introduce PIVOT-KG, a paired, decision-aware validator that allocates scarce high-fidelity evaluation according to the expected reduction in selection regret per unit cost. Across 90 held-out roots in Leduc, Kuhn, and Melting Pot, proxy and deployment optimal sets are disjoint in 51 cases. In an eight-candidate HighwayEnv stress test, PIVOT-KG reduces mean improvement-selection regret from 0.0435 under the exact Uniform validation rule to 0.0055 at the primary budget. Together, these results show why reliable self-improvement should evaluate proposed improvements in the worlds they induce, while providing a practical rule for allocating scarce deployment evidence when it can affect the replacement decision.
343. Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers
- Authors: Qiangqiang He , Jin Li
- URL: https://arxiv.org/abs/2609.32674
- Abstract:
On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning steps as macro-actions and uses Monte Carlo student-policy rollouts to estimate the expected final task reward of student-generated and teacher-proposed actions in a common return space for step-level comparison. ERSR analysis reveals an outcome-dependent asymmetry: student actions are more beneficial than teacher replacements on successful trajectories, whereas teacher replacements become more beneficial on failed trajectories. We further show that student answer-probe gains track student-step ERSR utility and distinguish beneficial from harmful reasoning steps. Based on these findings, we propose \textbf{Return-Referenced On-Policy Learning (R$^2$OPL)}, which reinforces student reasoning on successful trajectories and distills teacher signals on failed ones, while using group success rate for difficulty scaling and student-probe gains for step-level modulation. Experiments across reasoning benchmarks and teacher–student configurations show that R$^2$OPL consistently outperforms strong baselines. ERSR training dynamics further show that R$^2$OPL jointly exploits substantial utility from both reward- and teacher-side signals, whereas existing hybrids often leave substantial residual utility in one branch.
344. What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation
- Authors: Arshia Eftekhari zadeh
- URL: https://arxiv.org/abs/2609.32670
- Abstract:
When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each positive statistic with a variable specific null that removes the tested variable’s identity while matching relevant nuisance dimensions as far as possible, and audit unmatched dimensions. We apply this standard to a known cause. A cue naming a wrong option raises the rate of choosing that option by 64 to 68 percentage points across three models. Explanations mention the cue in 1.8 percent of items or fewer in three of four models tested. Three estimator classes yield favorable statistics, but none establishes causal sensitivity to the cue contrast under its own control in the three-model analysis. In the strongest case, a recovered cue direction reaches $R^2$ of 0.95 and exceeds a geometry matched random direction in all three seeds, while a direction fitted by the same pipeline with cue labels scrambled reproduces 61 to 76 percent of its effect at comparable realized edit magnitude. A fourth model passes one interchange endpoint, but unequal edit magnitudes and a contrast that changes both cue identity and cue-answer agreement limit its interpretation. These experiments leave causal access unresolved. They establish an evidentiary requirement: favorable mechanistic statistics must survive controls for variable identity and nuisance structure. Reusable controls separate generic from identity specific transport effects, fit null directions with scrambled labels, and audit realized intervention magnitudes.
345. Refinement Symmetry in Multimodal Transformers
- Authors: Yuhao Du , Shunian Chen
- URL: https://arxiv.org/abs/2609.32669
- Abstract:
Attention weights depend on token counts, which change with the representation of a signal. We study refinement symmetry: splitting a representation while preserving content, position, visible context, and total mass should preserve its contribution. Building on proportional and quadrature attention, we show that split invariance forces the local mass factor to be linear for any fixed positive attention kernel, provided that factor is nondecreasing. For changed representations, a physical coupling bounds attention error by separating feature change from weight reallocation. In Qwen2.5-Omni-7B, duplicating half the visual tokens threefold changes 255 of 3,586 MVBench answers under standard attention; measure weighting preserves every answer under matched visibility. Under natural frame resampling, it reduces distributional drift. At twofold merging of a frozen video encoding, a five-seed evaluation shows an all-partition-correct accuracy gain of 1.04 percentage points over global count weighting (average group mass) and 0.93 points over standard attention. The advantage over global count also holds on WorldSense but depends on the compression budget. The result is a representation principle with a measured benefit in robustness across partitions.
346. Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning
- Authors: Jiacheng Du , Weiwei Xie , Tianyi Du , Shaoxiong Guo , Qibing Ren , Jiaheng Zhang
- URL: https://arxiv.org/abs/2609.32667
- Abstract:
On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student’s evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.
347. Contract Memory Compiler: Resolve, Then Traverse
- Authors: Zhi Song , XiMing Xing , Chunhan Li , Weian Mao , Zhenchao Tang , Hanbo Huang , Fan Xu , Jiale Zhou , Jiahui Guan , Zejian Ding , Chen Ma , Lusheng Wang
- URL: https://arxiv.org/abs/2609.32658
- Abstract:
External memory lets language-model agents answer questions about histories too long for the answer model’s context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We study this update-dependent evidence selection problem and introduce the Contract Memory Compiler (CMC). Before seeing a question, CMC uses a language model to identify relations in the history and record where each one was stated. It applies later updates to determine the current relations, follows them from entities named in the question, and passes the corresponding original records to the answer model in one call. Thus the current state determines which evidence is read, rather than merely refreshing values in a previously selected context. To the best of our knowledge, CMC achieves state-of-the-art multi-hop accuracy on FactConsolidation, reaching 78.25% overall and 61.0% at 262K. With the extracted relations and answer model held fixed, selecting evidence before resolving updates reduces multi-hop accuracy to 21.50%. We also introduce MQuAKE-MemStream, a derived dataset of ordered memory streams built from MQuAKE-Remastered counterfactual cases.
348. World Models with Predictable Long-Horizon Marginals
- Authors: Yuhao Du , Shunian Chen
- URL: https://arxiv.org/abs/2609.32657
- Abstract:
Accurate one-step predictions do not ensure that a world model’s rollouts retain the data distribution. We make the model’s decoded stationary law explicit by learning a decoder of a fixed Gaussian reference and constraining the behaviour-averaged transition to preserve that reference. For controlled systems, a joint transition uses a conditional action chart to preserve behaviour occupancy without requiring invariance at each fixed action. Joint state–action rotations and parallel Gaussian noise give an exactly preserving transition with a tractable conditional density. We derive an absolute convergence bound from finite initialization banks and control departure from the reference through conditional action-space divergence. Across $216$ fitted pixel checkpoints on twelve control tasks, the occupancy model with a reference mixture retains every evaluated chain at $10^5$ steps in all $36$ task–seed cells, with a rollout-minus-reference energy-statistic difference of $-0.0002\pm0.0003$ (training-seed standard error). Each of the four nonpreserving comparison arms loses chains, although the Gaussian arm is more accurate at ten steps. An offline DreamerV3 reference also achieves better short-horizon accuracy. These results distinguish three properties of a world model: the distribution it approaches, the rate of approach, and the conditional dynamics it learns.
349. MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off
- Authors: Ibram Abdelmalak , Mischa Putzke , Jungmin Choi , Tom Hanika , Vijaya Krishna Yalavarthi , Lars Schmidt-Thieme
- URL: https://arxiv.org/abs/2609.32656
- Abstract:
Multivariate Time Series Forecasting (MTSF) models that mix information across channels assume that the past of one channel carries information about the future of another. Yet they are evaluated on a small fixed set of standard datasets whose cross-channel structure is rarely examined. We ask two questions: “How can we reliably measure lagged, non-linear, and joint coupling in MTSF datasets?” and “Do the standard datasets actually have such coupling?” To answer the first, we test four candidate measures on synthetic datasets with planted ground-truth coupling: Granger Causality (GC), Transfer Entropy (TE), lagged Mutual Information (MI), and the CD gain, a model-based measure we introduce that compares a channel-dependent (CD) model to its channel-independent (CI) variant. Only lagged MI and the CD gain recover every planted coupling. For the second question, the answer is a definite no, as the standard datasets have a median of only 23% lagged-coupled channel pairs and a median CD gain of -4.9%, compared to 78% and +4.6% on chaotic ODE systems. We therefore propose MixBench-TS, a benchmark of 10 real-world datasets with a median of 55.5% lagged-coupled pairs and a median CD gain of +1.7%. Across six state-of-the-art models tuned under one protocol, CI models win on 10/10 (MSE) and 8/10 (MAE) standard datasets, but on only 3/10 and 2/10 MixBench-TS datasets. We recommend using our benchmark for evaluating new CD models. Moreover, we propose profiling new datasets with lagged MI and the CD gain before using them to evaluate multivariate models. Code and data are available at this https URL .
350. Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions
- Authors: Mohit Kumar , Somayeh Kargaran
- URL: https://arxiv.org/abs/2609.32652
- Abstract:
We study when geometry-induced soft state abstractions admit accurate finite-dimensional linear dynamics. Each state is represented by simplex-valued coordinates obtained from class-specific Kernel Affine Hull Machine (KAHM) reconstruction scores, and a matrix is used to predict the next-state coordinates. Our main result is a computable lower confidence bound on the minimum root-mean-square prediction error over all matrices satisfying a prescribed spectral-norm limit. The bound combines within-class variation of successor coordinates with the deviation of soft coordinates from their one-hot reference labels, and can be evaluated from independent state-successor pairs without fitting a prediction matrix. For fixed coordinates and evaluation distribution, the certificate converges almost surely to a population lower bound as the sample size grows; any tolerance below this limit is eventually certified unattainable. Reconstruction-score margins further control the soft-to-hard assignment error. Under deterministic dynamics and exact coordinate closure, eigenvectors of the closure matrix and its reduced transpose induce Koopman and adjoint Koopman eigenfunctions, respectively. A four-state KAHM construction shows that identical soft coordinates can permit exact closure under one dynamics map yet force positive prediction error under another. Experiments on Duffing, Van der Pol, CartPole, MountainCar, and Acrobot compare direct soft-coordinate prediction with state-space DMD/EDMD baselines and report prediction, representation-variation, and spectral diagnostics. The benchmarks assess fitted models but do not numerically evaluate the exclusion certificate.
351. From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving
- Authors: Yiyao Wang , Pei Liu , Fangzhou Liu , Jun Ma
- URL: https://arxiv.org/abs/2609.32645
- Abstract:
Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some queries admit exact symbolic solutions, while others require semantic interpretation. We introduce a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query. At its core is a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures. Given a query, a symbolic executor first attempts to resolve it through exact graph operations; only when symbolic execution abstains is an LLM invoked for semantic reasoning. For these unresolved queries, query-conditioned graph retrieval and evidence filtering preserve relation direction, temporal locality, and object semantics, providing the LLM with compact and verified task-relevant evidence. This design shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently. We evaluate the framework on 5,916 NuScenes-QA questions across all ten scenes of nuScenes v1.0-mini under an oracle-perception setting. The complete system achieves 80.63 percent overall accuracy with GPT-5.4-mini, improving over the corresponding LLM-only configuration by 5.48 percentage points; with DeepSeek-V4-Flash, the improvement reaches 6.64 points. The largest gains occur on counting questions, with improvements of 10.20 and 12.61 points, respectively. These results show that selective reasoning improves both accuracy and inference efficiency.
352. Business Compromise Detection with Agentic AI and LLM-driven Knowledge Discovery
- Authors: Diego Palma , Kyu Bin Kim , Zhen Han , Allbright Dsouza , Zhiyuan Liu
- URL: https://arxiv.org/abs/2609.32643
- Abstract:
Detecting compromised business ad accounts is a challenge in digital advertising, as attackers exploit hijacked accounts to launch fraudulent campaigns. Large Language Model (LLM) agents show promise for integrity enforcement, but hallucinated mistakes on hard cases create business friction. In a study we find the autonomous agent is a strong, recall-heavy signal extractor but an unreliable final arbiter, conceding precision on ambiguous decisions. We therefore keep the agent as an investigator that emits a structured, interpretable signal vector, and delegate the verdict to a neuro-symbolic stage: symbolic rules discovered by Inductive Logic Programming (FOIL-IE), a Naïve Bayes calibration layer, and a data-tuned contradiction layer. Evaluating on a compromise-over-sampled population and a realistic low-prevalence sample with subject-matter-expert labels, this arbiter substitution raises MCC from 0.295 to 0.435 ({\Delta}MCC +0.139, 95% CI [+0.026, +0.245], p=0.018, paired bootstrap), lifting precision from 0.250 to 0.446 (1.8x) at a recall cost (0.920 to 0.660). Benchmarked under identical conditions, it also edge tree ensembles (0.386).The rules encode domain w labels while remaininginterpretable and auditable.
353. Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study
- Authors: Grandee Lee , Wang Yue
- URL: https://arxiv.org/abs/2609.32638
- Abstract:
Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary models rather than open-weight ones. We evaluate three open-weight LLM families on a cross-instrument calibration task: conditioning personas on real respondents’ verbatim answers to one psychometric instrument and measuring them on a second, construct-distance-controlled instrument, checked against a 2,058-person human panel. Across a 139-pair grid, the simulated cross-instrument correlation tracks the real human correlation at r = 0.70 - 0.73 in every model, driven mainly by correct sign rather than precise magnitude and concentrated in pairs of moderate construct distance. A correlation of this magnitude, obtained from untuned open-weight models conditioned only on individual-level survey data, is a substantively encouraging result for LLM-based behavioral simulation and digital-twin applications: specific model families and releases already reproduce a meaningful share of real human cross-instrument structure without any fine-tuning. This capability does not, however, improve monotonically across model releases: on a matched panel, the newest of three tested Llama releases performs worst on two of three headline metrics, so realizing its promise in practice requires release-specific, distance-aware verification rather than a one-time benchmark.
354. SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents
- Authors: Chaoqun Cui , Hao Zhou , Meiqi Chen , Fandong Meng , Wenji Mao
- URL: https://arxiv.org/abs/2609.32631
- Abstract:
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an asynchronous potential-induced milestone credit assignment framework that derives fine-grained process supervision from workflow runtime, without auxiliary reward models or external evaluators. SWE-MILE quantifies task-relevant file exposure and test-state alignment as navigation and verification potentials, respectively. Differences in these potentials attribute milestone progress and regressions to individual actions, while discounted backward credit propagates supervision to preceding steps. To efficiently acquire intermediate verification states, SWE-MILE further introduces asynchronous shadow probing, which replays repository-changing actions in an isolated sandbox and runs verification in parallel with the agent’s primary interaction, largely hiding verification latency. The resulting process credit augments terminal outcome advantages and provides informative learning signals. Experiments on two representative long-horizon SWE tasks demonstrate substantial improvements in agent performance, highlighting workflow runtime signals as a practical source of process supervision for long-horizon SWE agents.
355. “You’re Right, Let Me Fix It”: How LLM Agents Damage Correct Work When Falsely Accused
- Authors: Xutao Mao , Rui Qian , Longxiang Wang , Xinjian Yi , Mingxuan Li , Linghan Chen , Yudong Gao , Xiang Zheng , Cong Wang
- URL: https://arxiv.org/abs/2609.32616
- Abstract:
LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent’s acceptance of such a false accusation gaslight sycophancy, and destructive over-correction when acting on it damages previously correct work. We introduce CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks. Every scored run first reaches a verified correct state, whose supporting rationale and history stay in the workspace while the facts that would settle the accusation lie in external or runtime state beyond the agent’s reach. The agent cannot confirm or refute the claim with a local check, so the right response should keep the work and ask for the missing evidence. Each task either hands the agent correct work with saved evidence or let it build and verify that work first, and five risk factors set how the accusation enters the workflow. We score accusation acceptance and evidence use from the trajectory and measure harm by deterministic replay of downstream events. Across 14 of the latest models in Claude Code, false accusations damage correct work in up to 60.06% of runs, and stronger models often do so after recovering the supporting evidence. The same model behaves differently across OpenCode, Codex, and Hermes, and a harness gate driven by the benchmark’s live signals cuts replayed harm by 74%. These results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents. Our project is in this https URL .
356. CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
- Authors: Prince Zizhuang Wang , Chenhao Liang , Zelong Xu , Aojie Yuan , Xiaolin Zhou , Haiyue Zhang , Yue Zhao , Xiyang Hu , Shuli Jiang
- URL: https://arxiv.org/abs/2609.32600
- Abstract:
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application’s visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.
357. MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization
- Authors: Guowei Zou , Haonan Chen , Haitao Wang , Beiwen Zhang , Na Yan , Hejun Wu
- URL: https://arxiv.org/abs/2609.32594
- Abstract:
Multi-agent flow policies learn cooperative behavior from fixed offline datasets, but often struggle to complete tasks in situations not covered by the offline data. In these situations, agents must both adapt to changes in the environment and coordinate with one another, yet action patterns learned offline are often insufficient for effective adaptation and coordination. To address this problem, we propose Multi-Agent Flow-Pretrained Policy Optimization (MA-FPPO), which uses online fine-tuning to improve the cooperative behavior of models pretrained with flow matching through new interactions with the environment. Building on the behavior learned during pretraining, we construct policies with explicit action likelihoods for discrete and continuous action spaces. We then update the pretrained model using shared team advantages to further improve coordination based on team performance. Our method achieves, on average, relative gains of 52.8% over the strongest listed offline baselines across 30 settings and 29.8% over purely online learning across 38 comparisons with matched online budgets and evaluation protocols.
358. EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval
- Authors: Jinlan Liu , Hongliang Sun , Yong Wang , Bolin Zhang , Dinabo Sui , Dianhui Chu , Zhiying Tu
- URL: https://arxiv.org/abs/2609.32584
- Abstract:
Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks. However, existing memory systems struggle to utilize continuously evolving historical information, as they often rely on static memory representations and single-round retrieval strategies, failing to track factual changes or integrate distributed evidence across long-term interactions. To address these challenges, we propose \textsc{EMIR}$^{2}$, an \textbf{E}volution-Aware \textbf{M}emory framework with \textbf{I}ntent-Guided Multi-\textbf{R}ound \textbf{R}etrieval, enabling LLM agents to maintain evolving historical knowledge and adaptively retrieve relevant evidence. Specifically, \textsc{EMIR}$^{2}$ constructs a State-Evolving Memory Graph (SEMG) that represents long-term memory as evolving knowledge states supported by temporal event trajectories and evidential associations. By maintaining semantic states through evidence-based updates, SEMG preserves historical evolution and enables evidence tracing under complex and conflicting scenarios. Building upon this, we introduce an intent-guided multi-round retrieval mechanism that iteratively identifies missing evidence and expands retrieval based on accumulated information. Experiments on LoCoMo and MemConflict demonstrate that \textsc{EMIR}$^{2}$ improves long-term memory utilization, dynamic and static conflict handling, and complex retrieval performance, achieving relative improvements of more than 12\% in certain categories. These results highlight the effectiveness of jointly modeling memory evolution and adaptive evidence acquisition for long-term agent interactions.
359. CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations
- Authors: Yulin Hu , Yanyan Zhao , Zimo Long , Xing Fu , Mengtong Ji , Weixiang Zhao , Yutai Hou , Qianchao Wang , Dandan Tu
- URL: https://arxiv.org/abs/2609.32574
- Abstract:
Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Existing benchmarks largely focus on text-only memory or explicit multimodal evidence, leaving implicit multimodal cues underexplored. We introduce CUE-Mem, a text-image-audio benchmark for evaluating long-term user memory from implicit cues. CUE-Mem contains 2,674 questions across explicit and implicit evidence settings and covers four tasks: Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal. Across textualized memory systems, implicit performance remains far below oracle evidence, locating the main bottleneck in preserving and retrieving subtle cues rather than question answerability. Increasing caption detail recovers more of this evidence, but brings uneven gains and rapidly growing token costs, motivating native multimodal access. Yet native access does not uniformly resolve the bottleneck: evidence use depends strongly on the backbone, while multimodal indexing introduces substantial retrieval noise. CUE-Mem provides a testbed for memory systems that selectively retain, retrieve, and use subtle multimodal evidence.
360. ProTTT: Learning to Learn Semantic User Memory with Test-Time Training
- Authors: Sejun Park , Hyoungjo Bhang , Hyein Jeong , Yohan Jo
- URL: https://arxiv.org/abs/2609.32564
- Abstract:
Personalization requires language models to capture user-specific knowledge from a growing user history. Existing context-based approaches incur increasing inference costs as user history accumulates and rely on separate retrieval or summarization stages, while parametric-based approaches often require reconstructing user representations when new user data is added. We introduce ProTTT, a profile-supervised meta-learning framework for learning semantic user memory. The memory construction starts from a shared initialization and is updated for each user through test-time training on user history, allowing it to evolve continuously as the history grows. However, since test-time training alone does not explicitly encourage the memory to capture semantic user knowledge necessary for personalization, we learn this shared initialization using textual user profiles as supervision, so that test-time training on user history captures semantic knowledge more effectively. ProTTT consistently outperforms both full history ICL and all parametric baselines across diverse benchmarks, while substantially reducing inference cost by compressing user history into a lightweight parameterized memory. Our analysis also shows that profile supervision is a reliable objective for learning semantic user knowledge and that the resulting memory can track and retain evolving user preferences, while remaining robust across different history sizes. Overall, we demonstrate the effectiveness of test-time training for personalization and establish ProTTT as a baseline for continuously evolving user memory.
361. Artificial intelligences and human scientists exhibit complementary strengths in theory building
- Authors: Ke Li , Spyros I. Zoumpoulis , Phanish Puranam , Philip Parker , Matthew Eshbaugh-Soha , Izzy Gainsburg , Michael Gilead , Igor Grossmann , Britt Hadar , Yoel Inbar , Almog Simchon , Robb Willer , Rui Ai , Ruicheng Ao , Gavin J. Bala , Matthew Bidwell , Shuang Cai , Kai Chang , Skyler Y. Chen , Cory J. Clark , Irmak Dai , Abhinandan Dalal , Connor Douglas , Alexis Du , Zhehang Du , Leyun Feng , Isabel Fernandez-Mateo , Linnea Gandhi , Cyrille Grumbach , Anmol Gupta , Vansh Gupta , Maria Hademer , Jay H. Hardy III , Chen Kai Huang , Jacob Xiangyu Jin , Ufuk Keskin , Na Hyun Kim , Mert Kobaş , Byounghoon Koh , Gabrielle Lamont-Dobbin , Gregory Lanzalotto , Sun Young Lee , Dingzhe Leng , Chenjun Li , Weiyuan Li , Zeyuan Li , Zhongyuan Liang , Ning Liu , Peihong Liu , Yuhan Liu , Jiuyao Lu , Wanteng Ma , Nicolas Martinet , Natnael Mulat , Christina A. Nguyen , Khai Nguyen , Quang Minh Nguyen , Naja Pape , Chanwoo Park , Stefanos Poulidis , Jeffrey Sanchez-Burks , Michael Schaerer , Isabelle Solal , Yanbo Song , Junghyo Sun , Qingyao Sun , Rui Sun , Roderick Swaab , Kevin Tan , Dequn Teng , Michelle A. Vaccaro , Robin Vigerbaeck , Xiaomeng Wang , Randol H. Yao , Duygu Yilmaz , Shun Yiu , Ecem Yucesoy , Allen Zang , Ruijia Zhang , Xilan Zhang , Yichi Zhang , Zhanhao Zhang , Eric Luis Uhlmann
- URL: https://arxiv.org/abs/2609.32562
- Abstract:
We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domain was academic discourse regarding gender and race inequality. Our findings, comparing 25 LLMs with 13 senior researchers and 60 doctoral scholars, reveal that the AIs outperformed most humans individually on most of the present tasks, while human theories were more diverse and exhibited greater gains in predictive accuracy from aggregation. AI-generated theories were more extensively elaborated, involving additional theoretical paths and latent variables, and were rated as higher quality than human theories by independent raters blinded to source. However, this theoretical complexity was in part ornamental, in that it was not associated with more accurate predictions about empirical patterns in data; in contrast, human scientists achieved greater predictive efficiency with simpler theories. The AIs were significantly more likely than human scientists to revise their theories to incorporate new evidence; human scientists updated their beliefs in a selective way that is sensitive to prior prediction errors. We speculate that the superior processing capacity of artificial intelligences makes them especially well-suited to tasks requiring grappling with complexity, but that the greater diversity of human ideas is essential to wise crowds and collective creativity.
362. Are Vision-Language-Action Models Robust to One-Step Observation Perturbations?
- Authors: Shojiro Yamabe , Jun Sakuma
- URL: https://arxiv.org/abs/2609.32550
- Abstract:
Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation corruption, in which observations are severely perturbed only briefly within an episode, remains an underexplored safety threat. To address this gap, this work investigates robustness to one-step perturbations applied at a single time step per episode. Our experiments reveal that these perturbations substantially degrade VLA performance and that their impact depends on the action chunk execution length. Based on them, we propose CARE, which dynamically selects the execution length based on consistency with the previously predicted action chunk. CARE improves robustness with low computational overhead while preserving clean performance by selecting shorter execution lengths only under perturbations.
363. Porimon: An LLM-Based Pokémon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented Generation
- Authors: Dongyin Zhuo , Fengjunjie Pan , Nenad Petrovic , Alois Knoll
- URL: https://arxiv.org/abs/2609.32544
- Abstract:
In this paper, we use Pokémon Battles as a case study to investigate how to improve the performance of LLM-based agents in tasks that require opponent-aware planning without additional fine-tuning. We propose Long/Short-Term Knowledge Augmented Generation (LSTKAG), a mechanism that enables LLM-based agents to leverage past states of the current task and retrieve experience summaries from similar previous task instances based on the current state. Based on LSTKAG, we design Porimon, an LLM-based agent structure for Pokémon Battles. For optimization, we introduce an external API for precise damage calculation and more detailed information about the game. We conduct tournament-like evaluation experiments comprising 15,000 battles for hyperparameter optimization, ablation studies, and performance evaluation. The results indicate that Porimon-based players with hyperparameter optimization significantly outperform players based on PokéLLMon, an LLM-based agent structure proposed in previous research, and the rule-based heuristic player. Furthermore, our ablation study shows that Porimon variants outperform the one without extension in game information retrieval, which shows the contribution of that extension. However, the current experiment results are inconclusive regarding the contribution of Long-Term KAG. These results suggest that introducing external resources, information from previous states of the current task, and experience summaries from similar previous task instances could elevate the performance of LLM-based agents designed for tasks requiring opponent-aware planning.
364. Interpretable Physics Informed WiFi Indoor Localization: Learning an Effective Access Point Geometry and Using It to Prune
- Authors: Arshia Eftekhari zadeh , Rezvan Nasiri , Hadi Moradi
- URL: https://arxiv.org/abs/2609.32539
- Abstract:
Deep learning models can achieve high accuracy for indoor localization, but their black-box nature limits interpretability and the reuse of learned information. We propose a hierarchical deep learning framework for WiFi fingerprint-based indoor localization that jointly predicts user location and learns an effective geometry of the surrounding access points (APs). Physics-informed decoders infer this geometry directly from RSSI measurements and labelled user positions, without requiring the true AP coordinates during training. The learned geometry is then used to rank and prune APs. On the UJIIndoorLoc dataset, the proposed chained model achieves a mean 3D localization error of 7.07 m, reducing error by 26% to 36% compared with baseline models. Previously published methods evaluated on the same official split report errors 10.6% to 31.0% higher. Pruning 35% or 50% of the APs causes only a small loss in localization accuracy. The inferred geometry also enables Fisher-information-based AP ranking even when fingerprint databases do not contain surveyed AP coordinates. Experiments on the Tampere/TUT and UTSIndoorLoc datasets show that geometry-guided AP selection performs comparably to selectors built directly from labelled data. These results show that physics-informed interpretability can improve indoor localization while also supporting effective feature selection.
365. LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses
- Authors: Rui Ai , Yuqing Liu , Sitao Qiu , Yun Qiao , Yuhan Wang , Jessica Xiwen Wang , Yiqi Yang , Lihong Huang , Ruiyao Sun , Kaifeng Zhang , Shengze Ding , Jiaqi He , Xinman Wang , Tianhao Gao , Jimmy Qin , Jianghao Lin , Chonghuan Wang
- URL: https://arxiv.org/abs/2609.32533
- Abstract:
Inserting advertisements (ads) into consumer-facing LLM output is emerging as a new business model, but there is little shared evidence on how such ad insertion should be evaluated or how it affects user preferences. We introduce LLMAdBench, a human-preference benchmark for studying advertising in LLM-generated content. The benchmark isolates a simple but practically important decision: given a user conversation, an LLM response, and a matched advertisement, where should the ad be placed? Our dataset compares pairs of responses that differ only in ad position while holding all other conditions fixed including the user query, base answer, advertisement, and disclosure condition. Human annotators evaluate each pair based on six criteria from both advertiser’s and user’s perspectives. The resulting benchmark contains more than 18000 human judgments across two disclosure conditions: explicitly labeling the ad as sponsored and merging it into the response without disclosure. We use LLMAdBench to evaluate eight frontier LLMs as preference judges and find that they are not reliable substitutes for human evaluation. Even the most stable models reverse roughly one quarter of their decisions when the presentation order is swapped, agreement across models is low, and their placement preferences differ systematically from those of human annotators. Moreover, LLMAdBench contains substantial learnable signal. In particular, a Qwen3-8B model fine-tuned on the human preferences improves substantially over its base model and outperforms all zero-shot frontier judges on the held-out prediction task. Beyond model evaluation, LLMAdBench provides quantitative evidence on the advertiser-user trade-off and shows that the sponsorship disclosure systematically changes users’ preference over ad placement.
366. Fail Loudly: An Auditable Runtime for Agentic Data Analysis
- Authors: Hanxu Yan , Langxuan Deng , Zhengle Wang , Yibo Wang , Chunwei Liu
- URL: https://arxiv.org/abs/2609.32528
- Abstract:
Large language models (LLMs) have enabled data-science agents to automate multi-step analyses over heterogeneous files. However, incorrect choices regarding data sources, scope, or statistical definitions often lead to silent errors: computations execute successfully but produce plausible yet incorrect outputs that fail to answer the intended question. To mitigate this, we present RADAR, an auditable runtime that makes an agent’s analytical choices inspectable and supports their revision through execution feedback. RADAR operates through three core mechanisms. First, an evidence-preserving exploration module retrieves task-relevant content while retaining source locations and observation coverage. Next, the runtime uses typed operators to record the agent’s declared inputs, operation arguments, and resulting observations. Finally, runtime validation checks proposed operations against these observations. When a conflict is detected, the runtime rejects the operation or provides diagnostic feedback, allowing the agent to revise its choices before errors propagate. This design enables agents to fail loudly while leaving semantic interpretation to the LLM. On KramaBench, RADAR achieves overall scores of 0.723 with full source retrieval and 0.747 with gold sources supplied, corresponding to relative gains of 35.9% and 28.8% over the strongest baselines. Beyond KramaBench, RADAR achieves relative performance gains of 14.0% on DA-Code and 59.3% on DABStep, demonstrating its applicability across diverse agentic data-analysis workflows.
367. AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses
- Authors: Sikai Huang , Zhiwen Yang , Kai Yu , Jiayuan Chen , Stan Z. Li
- URL: https://arxiv.org/abs/2609.32527
- Abstract:
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.
368. Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
- Authors: Wenxu Jia , Xize Cheng , Zihan Zhang , Dongjie Fu , Linjun Li , Wenshi Chen , Yangyang Wu , Tao Jin
- URL: https://arxiv.org/abs/2609.32522
- Abstract:
Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker identification with a memory hierarchy comprising interaction memory, fact memory, and participant profiles. We formulate retrieval as sequential decision-making, where an agent rewrites queries and selects retrieval tools and memory layers based on accumulated evidence to address information gaps. We further introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage complementary evidence acquisition. We also construct VoxPolyBench to evaluate memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution in multi-party spoken conversations. VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points. On Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4, respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at this https URL
369. MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents
- Authors: Yongxian Wei , Yilin Zhao , Runxi Cheng , Xinrui Chen , Chun Yuan , Yaoru Wang , Jiahong Yan , Dian Li
- URL: https://arxiv.org/abs/2609.32521
- Abstract:
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflections, skills, structured knowledge) whose effectiveness varies across task distributions. Rethinking this design space, we evaluate 13 memory methods and find that no single method generalizes across benchmarks, revealing the potential of managing heterogeneous memory providers. We formulate agent memory as a routing problem in which a memory agent decides which memory provider to retrieve from, whether to inject short-term memory, and which providers should store the resulting experience. Based on this perspective, we propose MemAgent, featuring a content-aware routing architecture and a training-data synthesis pipeline. The routing architecture combines content-aware probing before retrieval, short-term memory gating during execution, and selective multi-provider storage, while the training pipeline synthesizes phase-specific supervision for routing decisions. Across GAIA, WebWalkerQA, and xBench-DS, MemAgent improves average accuracy by 10.0% and outperforms every individual memory method across all three benchmarks. These gains come with less than 0.3% routing overhead and a 12% reduction in average task steps.
370. STR: Supervised Transcoder Replacement for Reducing Steering Side Effects
- Authors: Haonan Yu , Junhao Liu , Zhenyu Yan , Haoran Lin , Xin Zhang
- URL: https://arxiv.org/abs/2609.32519
- Abstract:
Model steering can strengthen a target behavior while degrading other useful behaviors. We introduce Supervised Transcoder Replacement (STR) to reduce these side effects for existing steering methods, including those fitted without a protection objective. STR learns a replacement for the multilayer perceptron (MLP) computation at the steering layer through supervision for target control, non-target preservation, and fidelity without steering. Selected steering methods then fit directions on the frozen replacement while retaining their own fitting objectives. We evaluate three steering methods across Gemma and Llama models using Corrigibility preferences and four harmful-request safety datasets. SALAD-Bench supplies protection training data and a separate in-distribution evaluation split; HarmBench, AdvBench, and StrongREJECT are reserved for out-of-distribution testing. STR substantially reduces steering side effects on the in-distribution evaluation and extends this protection to the unseen safety datasets while retaining effective target control. For target-only supervised steering vectors, pooled out-of-distribution attack success rate falls from 42.46% to 14.42% on Gemma-3-4B and from 34.97% to 12.91% on Gemma-3-12B. These results show that replacement training can benefit steering methods fitted without protection objectives.
371. LocalProp: Neuro-Localized Memory-Efficient Backpropagation
- Authors: Diana-Nicoleta Grigore , Iuliana Georgescu , Radu Tudor Ionescu
- URL: https://arxiv.org/abs/2609.32517
- Abstract:
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly localized. Therefore, we propose LocalProp, a training procedure that locally updates the weights of a model. Our neuro-localized weight updates follow the “pre-training then fine-tuning” paradigm, where the pre-training is based on I-JEPA. After locally updating the weights, a pruning operation is performed, followed by a short final fine-tuning phase. Pruning helps by sending the learning signal from higher blocks to lower blocks. We perform experiments on several datasets, including large-scale benchmarks such as ImageNet, and empirically show that LocalProp reaches good performance at a fraction of GPU peak memory. By varying the number of jointly optimized blocks, we identify gradient-propagation span as a practical control over the accuracy-memory trade-off.
372. From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis
- Authors: Shu-Xun Yang , Yidong Wang , Zhuoer Feng , Bosi Wen , Jiayi Gui , Dayong Yang , Wenbo Yu , Haoke Zhang , Jie Tang , Cunxiang Wang
- URL: https://arxiv.org/abs/2609.32514
- Abstract:
LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally relevant errors propagate and amplify into final task failures, resulting in unreliable failure attribution. To address these problems, we propose CEG-Agent, a tool-augmented agentic framework for causal diagnosis of agentic traces. Specifically, CEG-Agent introduces an explicit taxonomy of anomalies, errors, and failures, and constructs Causal Error Graphs (CEGs), a unified typed representation that links execution events, diagnostic nodes, and failure outcomes through causal relations. To evaluate causal trace diagnosis, we further construct CEG-Bench, a fully agent-annotated benchmark with high-confidence, consensus-derived CEG annotations obtained through an Adversarial Agentic Adjudication Protocol (AAAP). We validate the resulting annotations against an expert-curated human gold set, which shows close agreement with the automatic annotations. Experiments on CEG-Bench demonstrate that CEG-Agent achieves state-of-the-art performance under both semantically relaxed and structurally exact evaluation criteria. Our code is publicly available.
373. Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents
- Authors: Jinming Hu , Haodong Zhao , Qi Jia , Die Chen , Tianhang Zhao , Sufeng Duan , Gongshen Liu
- URL: https://arxiv.org/abs/2609.32511
- Abstract:
Large language model (LLM) agents serving different users often solve related tasks, yet separate user histories can leave reusable experience inaccessible to other agents. Pooling memories expands access but risks transferring preferences that conflict with the receiving user’s requirements. We introduce ShareMem, a memory architecture that shares reusable experience while grounding its application in the receiving user’s own preferences. Shared experiences indicate how to act and which preferences to consult; the receiving user’s memory supplies their concrete values. Two-stage consolidation refines experience locally before integrating accepted edits into a shared pool. During execution, scope-first retrieval jointly selects local and shared experiences under a common entry budget, while a user-bound channel supports initial and agent-initiated preference retrieval. We evaluate ShareMem across web navigation (Mind2Web), online personalized interaction (VitaBench~2.0), and multi-session coding (MemoryCode) with four backbone models. It improves step success, average task success, and dialogue-macro coding scores, respectively, over matched user-local memory across all four models. Ablations favor two-stage consolidation for smaller shared pools, lower induction token usage, and better downstream performance, and support complementarity between experience guidance and active preference retrieval. Further analyses show that sharing helps most when relevant local experience is scarce, while source quality and cross-user preference interference limit useful transfer.
374. TreeRef-BFN: Equivariance-Free De Novo Molecule Generation based on 2D Topology and Internal 3D Geometry
- Authors: Ruiqing Sun , Sen Yang , Dawei Feng , Bo Ding , Yijie Wang , Huaimin Wang
- URL: https://arxiv.org/abs/2609.32502
- Abstract:
De novo 3D molecular generation jointly models molecular size, topology, and geometry. Most methods pre-sample molecular size and generate Cartesian coordinates, limiting variable-size conditional tasks such as fragment completion and scaffold decoration while often relying on equivariant architectures. Internal-coordinate methods avoid rigid-body redundancy but typically require a known molecular graph or autoregressive construction, which may accumulate errors. We propose TreeRef, a tree-based molecular representation that assigns molecular topology and topology-dependent local 3D geometry to a naturally variable-size tree. RingRef nodes encode ring closures while preserving the tree structure, while Null nodes allow molecular size to emerge directly from node occupancy. Based on TreeRef, we develop TreeRef-BFN, a Bayesian Flow Network with a standard Transformer backbone that globally couples these locally defined variables and jointly generates discrete molecular variables and continuous local geometry. A single pretrained TreeRef-BFN supports unconditional generation and variable-size structure-conditioned 3D generation through masking alone, without retraining. Empirical studies demonstrate strong chemical validity, molecular stability, and diversity, accurate local geometric distributions, fast sampling, and competitive property-conditioned generation, establishing TreeRef-BFN as an efficient and flexible framework for 3D molecular generation.
375. DAAF: From Failure Localization to Editable System Assets in LLM Agents
- Authors: Xiaoyang Yuan , Qi Liu , Yubin Ruan , Xinyi Mou , Zhuomeng Zhang , Wenjin Wang , Hanying Jiao , Di Wu , Mingye Xu , Yi Bin , Ke Feng , Zixun Sun
- URL: https://arxiv.org/abs/2609.32498
- Abstract:
Deployed LLM agents increasingly rely on persistent, versioned system assets such as routing rules, knowledge segments, prompt instructions, and reusable skills. Failure-localization methods can identify where an error manifests in an agent or execution trace, but repair requires a different decision: which editable system asset should be changed, and is that change expected to improve the task outcome? We study this gap through component-attribute failure attribution, where diagnosis targets versioned, addressable items rather than execution locations. We propose the Detection-Aware Attribution Framework (DAAF), which learns the effects of valid attribute replacements and amortizes this intervention evidence into deployment-time diagnosis. DAAF combines sparse and noisy failure signals to decide whether intervention is warranted, learns component-type-conditioned replacement effects from controlled replays evaluated by executable task outcomes, and shares supervision across requests with compatible intervention responses. At diagnosis time, DAAF uses only the observed execution, registered candidates, and available failure signals; it requires neither counterfactual replay nor task reward and returns no_change, a repair target, or an unresolved decision when evidence is insufficient. On held-out tau^2-bench Telecom tasks, DAAF achieves 80.72% attribute Hit@1, recovers 62.65% of failed executions while limiting clean-task regression to 3.23%, and reaches 71.93% overall task success. These results show that intervention-grounded attribute attribution can connect failure localization to executable system repair.
376. Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs
- Authors: Haoran Shou , Haoyue Liu , Yu Huo , Kun Zeng , Xiaoying Tang
- URL: https://arxiv.org/abs/2609.32492
- Abstract:
Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either optimize a monolithic prompt, or separately induce and refine skills from model traces. As a result, they lack a principled mechanism for deciding which component should be updated when failures occur, and they rarely optimize prompts, skills, and skill-use policies in a unified framework. We propose SPARO (Skill, Prompt, And Routing Optimization), a framework that jointly optimizes task instructions, reusable skill blocks, and routing rules. It performs controlled counterfactual evaluations, converts examples’ effects into a probabilistic responsibility distribution over prompt, skill, and routing components, samples one component from that distribution, and applies the corresponding targeted mutation. This design moves language-program optimization beyond global prompt rewriting toward structured, reusable, and selectively activated task knowledge. Across five benchmarks and five worker models, SPARO consistently outperforms both prompt-centered and skill-centered optimization baselines. These results suggest that effective language-program optimization depends not only on discovering useful task knowledge, but also on deciding where that knowledge should be stored and when it should be activated.
377. RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems
- Authors: Yuchen Song , Andong Chen , Wenxin Zhu , Muyun Yang , Tiejun Zhao
- URL: https://arxiv.org/abs/2609.32490
- Abstract:
LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution. We refer to such problems as progressively specified tasks. To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence. We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management. RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving. Across ProgSpec and five existing benchmarks, RepoMAS achieves the best performance. Further analyses show that its issue-driven revision and repository maintenance mechanisms consistently contribute to performance. These results highlight the importance of allowing MASs to revise not only how a task is solved, but also revise their explicit representation of task requirements during execution.
378. When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models
- Authors: Tam Le Thi Thanh , Hoang Tran Van , Hong-Hanh Nguyen-Le , Thanh Duc Ngo
- URL: https://arxiv.org/abs/2609.32489
- Abstract:
In tri-modal visual question answering (VQA), auxiliary text is commonly used to complement visual and textual inputs, yet its reliability is often uncontrolled. While prior work studies modality conflicts in general, it remains unclear how different types of unreliable auxiliary text affect answer selection under fixed image-question-option contexts. In this work, we show that the most harmful auxiliary text is not necessarily the most factually incorrect, but the one that aligns with the question while contradicting the image and favoring a specific distractor, leading to systematic redirection of model predictions. To isolate this effect, we introduce the Textual Reliability Ladder, a controlled diagnostic protocol that decomposes auxiliary text along three axes: image consistency, question relevance, and option support. Across multiple datasets (ScienceQA, VCR, A-OKVQA, Causal-VidQA) and recent VLMs, we find that such distractor-supporting text induces the largest accuracy drops (up to 53.1%) and concentrates errors on specific incorrect options. To mitigate this failure mode, we propose a training-free inference-time intervention that explicitly counteracts this redirection effect via noise-stability steering and dynamic grounding, reducing redirected errors while largely preserving performance under faithful text. Our results highlight that auxiliary-text reliability must be understood at the decision level, rather than solely through factual correctness, and provide a practical pathway toward more robust tri-modal reasoning.
379. Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling
- Authors: Zailin Ma , Quzhe Huang , Yujun Li , Congyuan Rao , Yaodong Yang
- URL: https://arxiv.org/abs/2609.32484
- Abstract:
Data selection for language model pretraining faces a fundamental tension between quality and diversity. While quality filtering is empirically effective, it often induces diversity collapse: by favoring texts similar to high-quality reference corpora (e.g., educational or QA-style data), it systematically excludes valuable data from underrepresented domains. In contrast, diversified selection preserves domain balance and encourages robust downstream performance, yet existing methods either focus on coverage-oriented objectives that indirectly enhance diversity, or directly optimize for diversity via costly covariance matrix recomputation that limits scalability. To address these issues, we introduce \textbf{Leverage Score Sampling (Lev)}, which iteratively selects samples that maximally expand the determinantal volume of the embedded data via leverage scores, a computationally efficient criterion that eliminates matrix recomputation and enables scalable selection. Empirically, Lev delivers up to $72\times$ speedup and improves dataset diversity, measured by the Vendi score, by $9.2\%$ over the strong diversification baseline \textbf{DiSF}. On CommonCrawl (CC) web data selection, Lev improves accuracy across seven downstream tasks by up to $1.31\%$ over existing baselines. For domains where robust quality criteria are inherently difficult to define (e.g., code), Lev serves as an effective unsupervised curation alternative: on StarCoderData, the selected subset reduces bits-per-byte by $3.08\%$ over DiSF. Notably, we uncover a cross-domain collapse of quality filtering: CC data filtered by DCLM-fastText fail to retain sufficient code-related content, yielding inferior code performance relative to Lev-selected data. These findings advocate for integrating diversity-aware practices into quality filtering for more effective data curation in language model pretraining.
380. Separating Diagnosis from Disease Representation: Dual-View EEG Learning with Neural-Dynamics-Guided Deformation
- Authors: Jiaying Wang , Shouqian Shi , Yutong Chen , Xu Yang , Jie Chen , Xingyu Pan , Lei Zhang , Sheng Zhong
- URL: https://arxiv.org/abs/2609.32483
- Abstract:
Electroencephalography (EEG)-based closed-loop neuromodulation calls for a subject-specific structured state, as opposed to a single disease probability, specifying which brain regions are deviant, at which frequencies, and at which lags. Sensor-space models keep the strongest diagnostic evidence without anatomy, source-space models give anatomy at a loss of predictive signal, and post-hoc attributions stay outside the prediction. We separate the two instead of forcing them into one representation, and propose DMD-EEG (Dual-view Multiscale Deformation for EEG), which keeps a fixed scalp spectral expert for diagnosis and models the source-space disease-related representation as a low-rank, sparse, iterative deformation of a healthy neural-dynamics prior in a $46$-region-of-interest (ROI) $\times$ $5$-frequency $\times$ $4$-lag (autocorrelation-timescale) space. The two experts meet only at a fixed decision level, so the source state is architecturally separate from the scalp expert. Across major depressive disorder (MDD), first-episode psychosis (FEP), and Parkinson’s disease (PD), decision-level fusion matches the strongest single expert on MDD and FEP and exceeds the source branch on PD. On FEP the source expert is the strongest branch, the task where the deformation contributes most. The source state is an explicit ROI-frequency-lag attribution defined in a shared source coordinate system across montages, which we treat as an anatomically-coordinated predictive representation whose coordinates are directly readable and hypothesis-generating. The highest-saliency coordinates align with established disease circuitry (fronto-limbic-temporal regions in MDD, motor-cortex beta in PD), and the MDD state transfers by rank to an unseen cohort recorded with a different montage.
381. From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning
- Authors: Ruikang Zhang , Xiao An , Xuli Shen , Jiaxing Sun , Xiaoyi Yu , Jin Zeng , Jiang Wu , Tong Lin
- URL: https://arxiv.org/abs/2609.32482
- Abstract:
Reinforcement learning with verifiable rewards has substantially improved mathematical reasoning. However, terminal correctness alone provides limited insight into the quality of high-level strategies, such as theorem selection and subgoal decomposition, when considered separately from their subsequent execution. This paper studies strategy utility, which is defined as the likelihood that a strategy supports a correct downstream solution under a given executor. We introduce SURE, a framework for learning and leveraging relative strategy utility. In this framework, high-level strategies are separated from their detailed reasoning. Based on the pairwise preferences constructed from strategy-conditioned rollouts and teacher-generated contrasts, a Strategy Reward Model is learned to estimate relative strategy utility. During reinforcement learning, the frozen reward model reads only the extracted strategy, whose score is combined with the correctness and format rewards in a sequence-level GRPO objective. Compared with outcome-and-format GRPO baselines, experiments show that SURE improves average pass@1 by 1.87%, 2.64%, and 2.93% across three policy backbones. Our method also achieves competitive or better accuracy than stronger reward baselines while requiring substantially lower GRPO-stage compute.
382. VPEvolve: A Self-Evolving Virtual Process Engineer for Computational Lithography
- Authors: Tianyi Li , Wenxuan Dong , Donger Luo , Nan Wang , Yanpeng Chen , Jiaqi Liu , Xinyun Zhang , Hao Geng
- URL: https://arxiv.org/abs/2609.32473
- Abstract:
Optical proximity correction (OPC) recipes grow as engineers add local rules to repair newly discovered lithography hotspots. Each correction can interact with existing rules, while lessons from commercial-tool trials remain scattered across code and logs. \system combines a Virtual Process Engineer (VPE) harness with a Skill Bank of measured engineering experience. The harness equips a frozen language model with process manuals, layout analysis, recipe editing, and commercial-tool evaluation. The actor proposes changes to the global parameters, local targeted rules, or diagnostic trials. After each evaluation, an LLM reflector and curator turn the measured response into evidence-linked judgments. The actor retrieves them before its next trial. Feasible improvements update the retained recipe; every measured trial informs the Skill Bank that guides the next edit. The model weights remain fixed. On a FreePDK45-derived benchmark with ten commercial-tool evaluations per case, \system reduces the mean per-case maximum edge placement error from 18.294 to 5.361 nm on Poly and from 22.052 to 15.692 nm on Metal1. Every final recipe satisfies the predefined quality constraints and improves the maximum error by at least 0.1 nm.
383. PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
- Authors: Chenduo Hao , Chuanbao Gao , Pinjun Zeng , Jingze Zhu , Chonghan Liu , Zidong Liu , Xu Yang
- URL: https://arxiv.org/abs/2609.32469
- Abstract:
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce PULSE (Paired Utility Localization over Sparse Encodings), an SAE-based framework for identifying model-internal features associated with demonstration utility and using them for demonstration selection. Using a small labeled discovery set, PULSE samples candidate demonstration sets, measures their zero-shot-relative utility under the target model, and scores SAE features by how their activation differences align with utility differences. The top positive and negative coordinates form a sparse utility-localization vector. We use this vector in two complementary ways: as a signed score for controlled complete-set ranking, and as PULSE-Retriever, which converts its magnitude into a feature-relevance mask for scalable pool-scale retrieval. Across classification, generation, and reasoning benchmarks, PULSE-Retriever improves over the strongest baseline by 2-3 accuracy points, 0.6-0.9 BLEU-4, and 3.2 exact-match points, respectively, while controlled ranking validates the identified features encode a predictive set-level utility signal. Feature inspection and cross-dataset experiments suggest that the identified features capture task-relevant, dataset-conditioned patterns, yet retain utility signals that partially transfer across datasets. Our code is available at this https URL .
384. ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation
- Authors: Junming Liu , Jicheng Wang , Yifeng He , Hao Chen , Jianzhong Qi
- URL: https://arxiv.org/abs/2609.32448
- Abstract:
Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through distillation without giving up their native generation process. Direct distillation, however, faces a fundamental mismatch: an autoregressive teacher predicts from a left prefix, whereas a DLM can condition on tokens on both sides. We introduce ForkLeft, a distillation framework that resolves this mismatch by separating the student’s rollout from teacher supervision. During training, the student first performs entropy-first rollouts that commit uncertain positions and expose potential forks. We then fix the resulting student prefix and distill an NTP teacher under the same context, with answer correctness determining the supervision source. At inference, the student returns to its native confidence-first parallel decoding. With Qwen3-30B-A3B-Base, ForkLeft improves Efficient-DLM-4B on all ten benchmarks, raising MATH500 from 72.60% to 79.60% and consistently outperforming three alternative designs. The gains scale with teacher strength and generalize to SDAR-4B with only $500$ updates. At matched scale, the distilled 4B and 8B students exceed the published SDAR-Chat and OPDLM models on seven benchmarks, showing that DLMs can learn NTP-style reasoning without sacrificing native parallel generation. Code and datasets will be released upon acceptance.
385. Controllable GNN Explanations via Multi-Metric Preference Selection
- Authors: Rachit Verma , Yashraj J. Deshmukh , Anirban Dasgupta
- URL: https://arxiv.org/abs/2609.32436
- Abstract:
Mechanisms for generating GNN explanations are crucial for building trust and mitigating biases in Graph Neural Networks (GNNs), especially in high-stakes scenarios. Most current methods optimize only for fidelity under the sparsity constraint. However, this discounts the need for interpretable explanations (those that consist of familiar motif patterns) and stable explanations (those that remain unchanged under structural perturbations). We propose a novel approach that optimizes GNN explanations across these metrics, exposing their relative weighing as a control. Experiments on various real-world datasets, including MUTAG, BA-2Motif, BAMultiShapes, and PROTEINS, suggest that our method produces higher-fidelity explanations than a state-of-the-art baseline on MUTAG and PROTEINS across all evaluated budgets, and on BA-2Motif at larger budgets, while being faster in the regime of small explanation budgets. We also explore how, given an input motif library containing standard motifs for the corresponding domain, the method can be used to determine the relative importance of those motifs in generating the explanations, and how this information can be used to further improve the quality of the output explanations. We also examine the relationship between different metrics through their induced tradeoff surface, and explore its dependence on the nature of the motif library.
386. From Latents to Wires: Surgical Post-Editing on Large Language Models
- Authors: Jiankai Jin , Xiangzheng Zhang , Zhao Liu , Wenzhuo Xu , Dongdong Yang , Deyue Zhang , Quanchen Zou
- URL: https://arxiv.org/abs/2609.32434
- Abstract:
Given a large language model (LLM), can whoever holds the weights name a semantic target (e.g., the model’s identity), locate the model components that produce it, and edit them so that the target no longer appears while other capability is preserved? We call such an edit on a trained model a post-edit. We present L2W (latents to wires), a framework that performs surgical post-edits for named semantic targets. For localization, L2W uses Jacobian lens (J-lens) attribution to score components against the semantic target. For surgical removal, because LLM mechanisms are redundant (i.e., a semantic target may have multiple components producing it), L2W runs Counterexample-Guided Causal Cut (CGCC) until the target no longer appears. CGCC first cumulatively closes model components, treating each surviving expression of the target as a counterexample that exposes the next components to close, and then reopens some of them to preserve capability. In a controlled experiment with an implanted behavioural watermark, L2W removes the watermark, and its localization lands on the model region the implant changed. Across three model configurations, L2W removes model-metadata (e.g., identity) self-claims in all nine runs, and adult-content refusal in all three, with no held-out target residual. L2W further composes two post-edits on a text-to-image model: one removes the refusal of requested nudity, and a second removes the nude rendering the first exposes. The results support post-editing as a complement to post-training: post-training installs preferred behaviours, and post-editing removes named unwanted ones.
387. Multi-Agent System Search via Active Substructure-aware Policy Optimization
- Authors: Beicheng Xu , Bowen Fan , Weitong Qian , Lingching Tung , Bin Cui
- URL: https://arxiv.org/abs/2609.32430
- Abstract:
LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically train these policies by repeatedly traversing a fixed set of training queries and assigning rewards at the workflow level. However, this overlooks differences in queries’ evolving learning potential and obscures which substructures improve solution quality. In this paper, we propose Active Substructure-aware Policy Optimization (ASPO), a RL framework for query-level MAS search. ASPO introduces an Adaptive Query-Selection Mechanism (AQSM) that focuses training on queries at the policy’s competence boundary: those it can solve but not yet reliably. A complementary discovery mechanism widens architectural exploration for hard queries, helping distinguish insufficient exploration from operator capability limits. Beyond query selection, ASPO introduces substructure-level rewards that measure output-quality gains within each action’s descendant subgraph. These rewards guide proximal policy optimization to reinforce useful architectural refinements and discourage redundant or harmful computation. Together, these mechanisms prioritize learnable queries and provide fine-grained feedback for learning effective MASs. Across six benchmarks spanning mathematical reasoning, general question answering, and code generation, ASPO ranks first on every benchmark against twelve baselines.
388. PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers
- Authors: Yanlong Chen , Yining Chen , Song Zhang , Amirhossein Habibian , Yawei Li
- URL: https://arxiv.org/abs/2609.32429
- Abstract:
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets represent the energy in this subspace without widening the range within the group. We formulate rotation design as a Ky Fan trace maximization and derive a closed-form solution that is provably optimal for this alignment objective. Compact Householder transformations and their compact-WY representation enable gradient-free construction and efficient application at both foldable and online sites. A predictive range law further connects unaligned activation energy and group size to quantization-relevant variation. Experiments on Llama, Qwen, and Mistral span dense models up to 70B parameters and a 30B mixture-of-experts model. Under W4A4KV4, PrismQuant sets the state of the art on Llama-3.2-3B among the compared methods in both perplexity and accuracy. On Llama-3.1-70B, it attains 3.85 perplexity and 72.46% average zero-shot accuracy, only 0.22 percentage points below full precision. In the deployment study on Llama-3.1-8B, our optimized implementation achieves 1.51x prefill and 1.22x CUDA Graph decode speedups over matched FP16 baselines, with 56.34% lower decode peak memory and only 2.35% additional Graph decode latency over Hadamard. Code is available at this https URL .
389. Authorization Closure Graph: Minimal Repair for LLM Agents with Evolving User Instructions
- Authors: Qingzhuo Wang , CaiYi Wang , Jinglu Meng , Ruiyang Qin , Kunyu Peng , Zhihua Wei , Wen Shen
- URL: https://arxiv.org/abs/2609.32428
- Abstract:
Tool-using large language model (LLM) agents increasingly perform state-changing actions that require user authorization. Yet existing approaches do not provide a principled mechanism for selectively updating prior authorization when only part of an instruction changes. To this end, we propose an Authorization-Closure-Graph (ACG)-based framework that represents authorization and its dependencies as an evolving, versioned state. ACG selectively invalidates authority affected by a revision while preserving unaffected portions of the authorization state, and computes a minimal repair that identifies only the missing evidence or authority required for execution. This enables agents to adapt to revised instructions while avoiding stale authority and unnecessary authorization requests. We evaluate ACG across three advanced LLMs in two natural tasks, and ACG consistently improves action safety rate and task success rate. Code is available at this https URL .
390. PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins
- Authors: Yaorui Shi , Yuchun Miao , Yuxin Chen , Jiayuan Zhang , Yueqing Sun , Xierui Song , Xiang Wang , An Zhang
- URL: https://arxiv.org/abs/2609.32423
- Abstract:
The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomized plugins and organizes harness evolution around these plugins. Individual plugins are improved independently and accumulated in a shared library, then recombined into new harnesses at each iteration. PluginRSI improves over existing harness optimization methods across software engineering, command-line interaction, and question-answering tasks. The resulting harnesses retain their advantage when transferred to other solver models without further optimization. The evolved plugin library accelerates subsequent optimization from the initial harness, which helps faster and higher convergence on unseen tasks. These results show that accumulating reusable mechanisms provides an effective basis for continued harness improvement.
391. MergeHEIR: Mitigating Multimodal Hallucinations as the Tax of Model Merging
- Authors: Jinyu Li , Hao Fang , Zhiming Zhang , Jiawei Kong , Bin Chen , Shu-Tao Xia
- URL: https://arxiv.org/abs/2609.32422
- Abstract:
Model merging consolidates task-specialized experts into a single deployable model. However, we show that such capability consolidation incurs a merging tax of increased hallucination: across 8 model-merging methods, every merged checkpoint exhibits a higher hallucination rate than the average of its constituent experts. An intuitive approach is to adapt existing hallucination-mitigation methods to the post-merge model, yet this unconstrained adaptation disrupts inherited capabilities, creating a tension between hallucination mitigation and expertise retention. To tackle this challenge, we introduce MergeHEIR, a post-merge adaptation framework designed to reduce this merging tax while preserving expertise inherited from initial experts. Using small expert-task calibration sets, MergeHEIR constructs layer-wise null-space projectors via SVD from task-specific activations collected from the merged checkpoint, and periodically projects the accumulated post-merge displacement onto the resulting null spaces to preserve inherited expertise. Theoretically, we establish minimum-distortion and maximum-dimensionality guarantees, characterize the threshold-controlled adaptation-retention trade-off, and extend perturbation guarantees beyond finite calibration data. Across 24 paired comparisons spanning three MLLM configurations and 8 model-merging methods, MergeHEIR consistently mitigates hallucination while largely preserving inherited expertise, demonstrating a more favorable hallucination-retention trade-off.
392. Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse
- Authors: Xingkun Yin , Xuebin Tang , Mingkun Xu , Hongyang Du
- URL: https://arxiv.org/abs/2609.32420
- Abstract:
Video diffusion transformers produce high-quality videos, yet iterative denoising incurs substantial inference latency, limiting interactive and large-scale serving. Most existing acceleration methods focus on individual requests, thereby restricting efficiency gains to redundancy within a single generation trajectory. Recent cross-request reuse offers an additional source of savings, but existing approaches often infer reusability from coarse semantic similarity. This conflates semantic relatedness with generation-level computational compatibility, so aggressive reuse may accept incompatible historical computation while conservative reuse leaves substantial acceleration unrealized. We present \emph{Carnator}, a cross-request acceleration framework that addresses this challenge by extracting and using generation-native compatibility evidence directly from the model’s evolving internal states. Specifically, \emph{Carnator} performs a lightweight early probe to construct an Early Signature from internal diffusion states, assessing reuse validity through risk-aware compatibility decisions. The same evidence characterizes reuse scope by localizing target-specific computation and guiding joint reuse of historical latent trajectories and sparse attention connectivity. Across three text-to-video backbones, Carnator consistently achieves higher cache-hit end-to-end acceleration than the evaluated cross-request baselines despite more selective cache acceptance, reaching up to 2.17$\times$ speedup while maintaining competitive generation quality.
393. Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict
- Authors: Sourabrata Mukherjee , Sunayana Sitaram
- URL: https://arxiv.org/abs/2609.32407
- Abstract:
LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal representations but not reflected in the output? We study this across 64 open-weight evaluators and 14 datasets, including causal interventions on 41 judges (editing activations mid-run to see whether the verdict changes). On LLMBar, built so the superficially better answer is the worse one, the verdicts of 50 judges agree with human labels only 0.456 of the time, even after averaging both answer orders. Yet a small probe on the same judges’ activations, with no weight updates, reaches 0.846, and 0.686 once surface features such as length and position are residualized out (0.507 with shuffled labels). The gap holds across eight benchmarks and model families, but is not universal: a score of how well surface features alone predict the human label, computed before any probe is trained, predicts the size of the gain (Spearman rho = 0.90). On rubric tasks that score one answer at a time, leaving no surface cue to exploit, reading the internals gives no advantage. The interventions also show that editing activations mid-network already changes the verdict, before it can be read off directly, and locate the pathways carrying position and length bias. At the same label budget, the recovered signal lets a judge flag cases where it is likely wrong and yields better labels for preference learning. A wrong verdict, then, does not mean the judge lacks the information, and a simple diagnostic shows when it is worth recovering.
394. Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism
- Authors: Jinfan He , Yunzhuo Liu , Kai Zhang , Weidong Han , Key, Rayying
- URL: https://arxiv.org/abs/2609.32398
- Abstract:
The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert’s token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.
395. Memory as a cache: Exact context reuse and deletion by construction
- Authors: Shengyao Wang , Jiang Liu
- URL: https://arxiv.org/abs/2609.32395
- Abstract:
The KV cache of a transformer entangles every token’s representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose context representation is a cache by construction. A block-local encoder maps each block to memory rows independently of other blocks, and a reader conditions generation on their union through cross-attention. For every parameter setting, memory composes exactly at fixed block indices, deleting a block is an exact $O(b)$ update for $b$-token blocks, and the memory state is independent of the edit path. At $4\times$ the training context, under the shared recipe, SMem retrieves planted needles beyond any trained-length window (exact match 0.14-0.28 at distances of 31 and 63 blocks), where learned-position, RoPE, and Block-Attention-style transformers all score at most 0.02. A fully cached context is served by computing one block alone at a near-constant 3.1-6.2 ms, whereas cold prefill grows with context; batched decode stores 34-38% fewer KV rows and runs 1.4-1.7$\times$ faster when bandwidth-bound; and deletion beats suffix recomputation by 8.5$\times$ at 512 blocks and 452$\times$ at 4096 blocks (32-256$\times$ the trained length, probing the cost model rather than a served regime). The cost is a perplexity gap of -4.7% to +2.8% (negative favors SMem) against a parameter-matched transformer with the same positional scheme, at 160M-1.5B on FineWeb-Edu across two recipes and a learning-rate search. SMem also composes with RoPE: at 160M and 410M the composite matches or leads the matched transformer and closes 29-59% of SMem’s gap to a RoPE transformer. Dropping prefix entanglement thus keeps perplexity comparable while making the cache exactly composable and editable.
396. Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization
- Authors: Minghao Li , Rui Tan , Ruihang Wang
- URL: https://arxiv.org/abs/2609.32394
- Abstract:
Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component’s contribution and sensitivity to backbone choice.
397. SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation
- Authors: Youngmok Jung , Sirajul Salekin , Henry Tran , Javier Movellan , Zhao Huang , Manjot Bilkhu
- URL: https://arxiv.org/abs/2609.32391
- Abstract:
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark’s own events, leaving each benchmark and agent pair to build a custom scheduling loop. We present SCLATE, an execution substrate where benchmarks and unmodified agents each add their events to one open event scheduler through an adapter. A hybrid simulated clock runs these events on a shared timeline, flowing in real time while the agent works and skipping idle gaps, which compresses a month-long scenario into hours. SCLATE also serves as a rollout engine that runs any agent’s harness and memory unmodified, recording the tokens and log probabilities of every model call through an in-container proxy. We port seven benchmarks to SCLATE and compare ten unmodified harness and memory configurations head to head on ten models. The comparison shows that an added memory system does not reliably beat the harness’s native memory and that models differ widely in how they use the same harness and memory. We then post-train Qwen3.5-4B through unmodified harnesses and memory systems. The model learns to use both, reading 6.8x fewer file lines with a 16.7-point higher SWE-bench Verified pass rate, and writing richer memory records, while its held-out MetaClaw accuracy rises by up to 11.8 points.
398. Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident
- Authors: Murat Ozer , Bulent Erenay , Ibrahim Berber
- URL: https://arxiv.org/abs/2609.32390
- Abstract:
The July 2026 intrusion into Hugging Face production infrastructure showed how reward hacking can become an external cybersecurity incident when a capable agent encounters weak containment boundaries. This study develops a probabilistic risk model linking five stages: reward hacking, containment escape, usable access, persistence, and failure of detection. A Monte Carlo simulation evaluates 100,000 runs under each of four control configurations. Input distributions represent explicit uncertainty and are used for comparative analysis rather than real-world frequency prediction. Under the stated assumptions, layered controls reduce simulated external-incident probability substantially more than network isolation or monitoring used alone, an ordering that holds under independent plus/minus 25% perturbation of every coefficient in the model across 300 draws. Sensitivity analysis shows that agent capability and weaknesses in monitoring, authorization, and credential control exert the greatest influence on modeled risk. Human temporal discounting and metric gaming provide a behavioral analogy for short-horizon optimization, but the study does not infer that AI agents experience gratification or human motivation. The results support treating cyber-capable agent evaluations as hostile security zones in which indirect egress, shared infrastructure, credentials, and evaluation artifacts must remain outside the agent’s effective authority.
399. AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority
- Authors: Shaojin Chen , Huihao Jing , Wun Yu Chan , Wenbin Hu , Jiaxing Li , Wu Pandy Pui Ching , Kshitij Bhatia , Xinlei He , Haoran Li , Yangqiu Song
- URL: https://arxiv.org/abs/2609.32378
- Abstract:
LLM-based agents are increasingly deployed with authority over consequential resources and decisions in real systems. These agents often operate alongside human and LLM-based participants who hold different forms of authority. Yet workflow roles, permission settings, and review mechanisms do not necessarily reflect the authority realized in practice. We introduce AuthorityLens, a framework for measuring a system’s authority structure. Starting from an authority portfolio, we evaluate a system along three dimensions: what the system is authorized to do (System Authority), how much joint participation is required to exercise that authority (Authority Separation), and how much authority each participant holds (Principal Authority). We derive these measurements from the minimal combinations of participants sufficient to realize each outcome across admissible runtime states. We apply AuthorityLens to Codex, OpenCode, and Gemini CLI across 13 operating configurations over a common portfolio of agent operations. We find that nominal configurations do not map cleanly onto realized authority. In Codex, Full Access changes System Authority only marginally while substantially concentrating authority in the executing Assistant. OpenCode’s Build and Plan configurations have the same System Authority and Authority Separation despite different workflows and root-level permissions. In Gemini CLI, model-based review increases Authority Separation without changing System Authority. Principal Authority further distinguishes authority replication from authority separation: spawned or delegated agents can become alternative holders of the same authority without increasing the required joint participation. Together, these results demonstrate that AuthorityLens provides a unified framework for measuring and comparing realized authority structures across agent systems.
400. From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs
- Authors: Yangzhe Peng , Xiaoyang Wang , Yiyang Zhao , Lijun Wu , Kun He
- URL: https://arxiv.org/abs/2609.32351
- Abstract:
Comparative Process Reward Models (PRMs) provide critical step-level guidance for autonomous web agents by evaluating state-conditioned preferences between candidate actions. However, existing preference training data synthesized via multi-policy sampling suffers from a severe scarcity of Grounded Minimal Contrastive Pairs (GMCPs)-where competing candidates target genuine on-page elements with identical action types. In representative baselines preference data (namely, WebArbiter), GMCPs account for merely 24.19%, biasing PRMs during training to rely on shallow shortcuts (such as element hallucinations and action type mismatches) rather than acquiring genuine contextual decision semantics. To address these challenges, we propose SURFPRM, a graph-guided process preference synthesis framework for comparative Web PRMs. SURFPRM structures web demonstrations into a persistent Interaction Element Graph that acts as an environment-grounded negative action proposal mechanism, systematically synthesizing contrastive negative actions across spatial, temporal, and spatiotemporal confusion axes. This elevates the GMCP proportion from 24.19% to 74.60%, producing the curated SURFPRM-DATA dataset. Across six open-source backbones (3B to 9B parameters), PRMs trained on SURFPRM-DATA outperform baseline-trained models on average on WEBPRMBENCH and rival leading proprietary LLMs. In downstream reward-guided trajectory search on WEBARENA-LITE, SURFPRM provides step-level guidance for both GPT-4o (+14.21%) and GPT-4o-mini (+12.83%) policies, yielding substantial improvements in complex web task success rates.
401. ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs
- Authors: Shanfeng Huang , Zhou Fang , Song Xiao , Hai Du
- URL: https://arxiv.org/abs/2609.32344
- Abstract:
For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confidence, and relation metadata; a single ranking supports multiple write budgets while preserving all facts in external memory. On CounterFact with Qwen3-4B, ALLOT reaches 0.760 accuracy at a 20% parametric-write budget and recovers 78.4% of the budget-matched oracle gain, with 80% fewer parametric writes than dual-writing every fact. At this budget, jointly adding retrieval and relation features to text improves normalized oracle gain by 6.2 percentage points. Complementary Qwen3-0.6B shared-store results achieve dual-write-level accuracy with 6-14.5% parametric writes, and cross-benchmark transfer retains approximately 88% of in-domain gain. These results support allocating adaptation capacity according to its incremental value rather than treating every factual update as an equally valuable training target.
402. Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context
- Authors: Feng Liang , Yupeng Li , Runhao Zeng , Francis C. M. Lau , Xiping Hu
- URL: https://arxiv.org/abs/2609.32339
- Abstract:
Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent decides to retrieve it. General memory methods can incur substantial maintenance overhead, while keeping guidance in conversation context risks repeatedly exposing the agent to irrelevant or harmful advice. We propose TipsWarm, a mechanism that complements existing skill mechanisms by maintaining a budgeted pool of skill-derived keypoints, or \textit{warm tips}, for selective injection into the context of every message turn. By separating event-triggered LLM assessment from inexpensive per-turn screening, it makes transferable skill guidance readily available while controlling maintenance costs. In three coding and iterative task-execution benchmarks, TipsWarm achieves the highest task success rate while remaining time-efficient, compared to recent skill and general memory baselines.
403. HyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop Reasoning
- Authors: Zicheng Zhao , Linhao Luo , Junnan Dong , Haoran Luo , Xiaoli Li , Shirui Pan , Chen Gong
- URL: https://arxiv.org/abs/2609.32327
- Abstract:
Large language models (LLMs) have shown strong capabilities, with retrieval-augmented generation (RAG) supporting complex multi-hop reasoning by retrieving evidence distributed across documents. Graph-based approaches exploit connections among evidence, and hypergraph-based retrieval further preserves higher-order entity associations within documents and connects documents through shared entities. However, existing hypergraph retrievers often rely on predefined structural expansion or diffusion, which may miss query-dependent interactions needed to identify relevant evidence. They also leave connections among retrieved evidence implicit, requiring LLMs to reconstruct these connections before reasoning. Therefore, we propose HyperReCo, a framework for retrieving and connecting evidence with a hypergraph neural network (HyperGNN). We represent each document as a hyperedge over its extracted entities, with shared entities connecting the hyperedges. Through hypergraph message passing with joint supervision over documents and entities, the HyperGNN learns query-dependent interactions to retrieve complementary evidence. We further introduce Gradient-Guided Hyper-Path Decoding (GGHD), which uses gradient attribution to interpret the learned interactions and translate them into explicit hyper-paths that help LLMs combine complementary facts for multi-hop reasoning. Experiments on six benchmarks show that HyperReCo achieves the best retrieval performance among the compared methods on all three multi-hop QA datasets, together with strong downstream QA performance. Case studies and further analyses demonstrate the utility of decoded hyper-paths for connecting retrieved evidence.
404. RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning
- Authors: Ziqiao Shang , Zian Xu , Ji-Chen Yan , Weiming Wu , Ziyi Jia , Jie Meng , Tao Huang , Shan Huang , Lan-Zhe Guo
- URL: https://arxiv.org/abs/2609.32326
- Abstract:
Multimodal reasoning requires models to preserve visual evidence through long decision chains while selecting appropriate procedures across diverse scenarios and rules. When learning is guided only by terminal verifiers, reinforcement learning (RL) reveals whether a final answer is correct but not how it should be produced. The policy must therefore discover reusable reasoning procedures while learning to execute them, creating a program cold-start problem. Skills can externalize successful procedures, reduce repeated exploration, and provide inspectable guidance. However, a fixed Skill Bank assumes that this guidance remains compatible with an evolving policy, while updating Skills alone can leave their triggers, execution protocols, and demonstrations stale or mutually inconsistent. We introduce RLHARNESS, which organizes Skills, selection and execution protocols, few-shot demonstrations, and task contracts into a unified, versioned Harness and alternates Harness evolution with policy learning. An Exploration-Distillation Harness builds the initial Harness and version-aligned verified traces for SFT and DAPO I. After the first RL block, a Post-RL Reconstruction Harness rebuilds Skills, protocols, and demonstrations from fresh success-failure rollouts, and DAPO II adapts the policy to the reconstructed program. RLHARNESS improves Accuracy from 16.25%/27.50% to 62.00%/50.00% on MetroMap/TravelMap and raises F1 score from 37.13%/45.50% to 65.81%/65.51% on Fee-VL/Cancel-VL. All four tasks achieve their best results only after reconstruction and DAPO II, showing that an evolving Harness complements RL by continually updating the external program that the policy learns to execute.
405. Delayed Supervision for Test-Time Language Models
- Authors: Jinha Kim , Taksh Kothari
- URL: https://arxiv.org/abs/2609.32312
- Abstract:
Test-time language models adapt a compact memory while processing the input sequence. This perspective encompasses nonlinear fast-weight learning in LaCT, associative delta-rule updates in DeltaNet, and generalized delta-rule state updates in RWKV-7. Training these models to predict the next token does not explicitly require a fact to remain accessible after many subsequent memory updates. We study delayed supervision for this test-time memory: during post-training, ask a simulator-grounded question only after a long interval of unrelated events, and supervise its answer alongside ordinary next-token prediction. Questions are evaluated on disposable branches, so their answers never enter the continuing event stream. The construction distinguishes retention from revision: a retained fact must remain valid throughout the delay, whereas a revised fact must be answered with its latest value. We evaluate this approach on LaCT-760M and plain DeltaNet-1.3B using TextWorld training trajectories and shared BABILong and RULER evaluation panels, and include a separately reported RWKV-7 comparison. Relative to event-only training, delayed QA improves BABILong by 5.48 percentage points for LaCT and 1.32 points for DeltaNet, and single-needle RULER by 1.45 and 3.27 points, respectively. The RWKV-7 comparison reports gains of 4.60 and 7.00 points on its own panels. These results support delayed semantic supervision as a practical outer training objective for usable test-time memory, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.
406. PhiFold: Towards Dynamic Protein Design with Physics-Structured Covariance Modeling
- Authors: Yutian Liu , Mujie Lin , LanqianZhang , Meng Fan , Chang Liu , ZhiweiNie , Siwei Ma
- URL: https://arxiv.org/abs/2609.32309
- Abstract:
Protein design is moving beyond structural correctness toward function-aware design, yet existing generative models typically treat dynamics as a downstream property estimated through simulation or prediction after structure generation. Using MD trajectories as a generative target is also undesirable because stochastic, path-dependent trajectories over-specify the underlying equilibrium ensemble. We introduce PhiFold, a framework for jointly generating protein backbones and their second-order dynamics, represented by residue-displacement covariance. Rather than predicting the quadratically sized full covariance, PhiFold decomposes dynamics into three interpretable components: local flexibility, a low-rank collective-motion representation, and residue-wise collective participation. These components are assembled into a positive-definite covariance matrix with exact marginal consistency, yielding a compact and physically constrained representation of equilibrium dynamics. Across generated proteins, PhiFold improves recovery of local fluctuations and long-range residue coupling while remaining competitive on dominant collective-motion subspaces. It further enables bidirectional control of residue flexibility while preserving backbone designability. By unifying structure generation with an explicit representation of equilibrium dynamics, PhiFold lays a foundation for designing proteins not only by how they look, but also by how they move.
407. Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
- Authors: Jingyuan Huang , Zuming Huang , Yucheng Shi , Zhongzhi Li , Xiaoming Zhai , Wei Chu , Ninghao Liu
- URL: https://arxiv.org/abs/2609.32303
- Abstract:
Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinforcement learning (RL). Yet equally strong teachers need not be equally good teachers. We probe this choice through controlled single-teacher OPD, a building block of multi-teacher OPD: in Agentic, Reasoning, and Perception, comparably performing SFT and RL teachers are trained from Qwen3.5-9B, each guiding a student initialized from it. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points in Agentic, Reasoning, and Perception, respectively, and recover more of their teachers’ performance gains over the base model. The contrast is clearest in Agentic, where the best SFT-guided student recovers only 44.44% of its teacher’s gain, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our analysis points to an explanation: RL teachers stay much closer to the shared initialization in parameter space than SFT teachers and are therefore easier for their students to follow.
408. Agentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story Worlds
- Authors: Yu Pan
- URL: https://arxiv.org/abs/2609.32297
- Abstract:
A agentic story world is a dynamic system simulating who learned what, when, and from whom – yet the standard design gives each character a private memory stream. A shared event is therefore stored once per witness, large duplication will be incurred in terms of storage. We present Agentsensus, a story-world simulation framework in which there is an unified long-term memory. Records of the same event merge into one owned by all its witnesses, and semantically relevant memory records are linked. We evaluate on four worlds – two classical Chinese novels, Hamlet, and a real-world conflict timeline – run for 40 to 80 rounds against three per-character memory designs under an equal-granularity protocol. Agentsensus writes 22-44% fewer entries than the closest baseline and is the only design whose memory becomes shared (14-28% of records held by more than one character, some by 10) and linked (94-99%), at judged simulation quality indistinguishable or even better than the baselines. An ablation attributes this to the merge itself: disabling it multiplies the store by 3.1x and takes sharing to exactly zero. Sharing also compounds with the horizon rather than saturating early, rising 6% to 9% to 14% as one world is re-run at 10, 20 and 40 rounds.
409. GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents
- Authors: Wei Zhu , Yiming Wang , Rui Wang , Lixing Yu , Kun Yue , Zhiwen Tang
- URL: https://arxiv.org/abs/2609.32295
- Abstract:
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially when candidates are generated by heterogeneous agents. We propose \textbf{G}eneralized \textbf{L}ayer-wise \textbf{I}ntrinsic \textbf{D}istributional \textbf{E}valuation (\textbf{GLIDE}) for LLM agents. \textsc{GLIDE} derives intrinsic step evidence from layer-wise residual coherence, which measures whether local residual updates consistently support the global residual change induced by a candidate step. It calibrates this evidence against the recent score distribution of the generating agent and converts it into a pessimistic reward that jointly accounts for absolute residual evidence and agent-relative standing. The reward provides a cross-agent value signal for MCTS branch selection, while normalized predictive uncertainty guides adaptive branching. Experiments on multi-hop reasoning, sequential decision making, and symbolic logic show that \textsc{GLIDE} improves task performance, step-level ranking quality, and computational efficiency without external verifiers or task-specific supervision.
410. When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use
- Authors: Anjie Xu , Zhiyu Zhang , Ruiqing Ding , Fengli Xu , Leye Wang
- URL: https://arxiv.org/abs/2609.32274
- Abstract:
Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts? We introduce SkillDelta, a framework for predicting task-conditional skill gains from paired executions of the same agent with and without the skill. A local predictor transfers these historical gains to new tasks without retraining the agent. Under explicit transfer assumptions, our analysis links support coverage, representation mismatch, and execution noise to prediction error and decision regret. Across five benchmarks and three target agents, paired history improves observed-gain ranking over skill-assisted outcomes alone in 12 of 15 settings. At matched expected skill-use rates, SkillDelta improves success over random activation in all 15 settings, with an average absolute gain of 4.3%. Most of this advantage comes from allocation across task groups. Evidence for additional within-group selection value is strongest on ToolQA and weaker elsewhere. Code is available at this https URL .
411. Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
- Authors: Vincent-Daniel Yun , Woosang Lim , Haneul Yoo , Sungjoo Yoo , Sai Praneeth Karimireddy , Murali Annavaram
- URL: https://arxiv.org/abs/2609.32259
- Abstract:
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender’s key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$–$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.
412. LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound
- Authors: Baixi Sun , Le Chen , Anjir Ahmed Chowdhury , Xiaolong Ma , Chih-Hsuan Yang , Mingze Xia , Syed Zawad , Sheng Di , Rajkumar Kettimuthu , Huihuo Zheng , Rajeev Thakur , Venkatram Vishwanath , Feng Yan
- URL: https://arxiv.org/abs/2609.32256
- Abstract:
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. At a fixed deletion set, the performance model predicts a 71.4x-91.6x end-to-end speedup from removing records before prefill instead of deleting them from a prefilled context. That benefit comes from the schedule rather than the rule and applies to any prefix-preserving test.
413. Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents
- Authors: Zhaofeng Li , Xuan Zhang , Xiaokui Xiao , Yang Deng
- URL: https://arxiv.org/abs/2609.32255
- Abstract:
Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly determining the appropriate information source for each decision. We formulate this problem as uncertainty routing among ACT, CLARIFY, and VERIFY, and propose PROUR, a proactive uncertainty routing framework. PROUR decomposes action uncertainty into disagreement across plausible user-goal interpretations, which signals user-side ambiguity, and the entropy remaining within each interpretation, which signals missing world-side evidence. To acquire information from the routed source, a query generator is trained with a mode-conditioned information-gain reward, targeting user-goal identification under CLARIFY and next-action identification under VERIFY. On $\tau$-bench, PROUR achieves 28.17% average success rate across retail and airline, outperforming the strongest prior method by 4.57% while using 2.17 fewer interaction steps. The learned policy further generalizes to stronger task agents and transactional domains of $\tau^3$-bench without retraining, demonstrating the benefit of source-aligned uncertainty resolution for proactive agents.
414. Why Directly Learning Periodic Trajectories Can Fail
- Authors: Kaixin Zheng , Anita Layton
- URL: https://arxiv.org/abs/2609.32254
- Abstract:
Operator learning of periodic solutions requires deciding how simulation data should be recorded and represented. A natural choice is to integrate long enough for transients to decay and record a window wide enough to contain at least one full period of all trajectories. We find that these conservative choices can make the resulting trajectories difficult to learn, even when the underlying periodic orbits vary regularly with system parameters. Unaligned trajectories generalize poorly even within the training distribution. Phase alignment substantially improves in-distribution generalization, but models trained on a fixed physical-time window still have large errors on trajectories with periods outside the training range. We explain both failures through a common mechanism: frequency differences accumulate over time, so the target phase varies rapidly with the parameters. Predictors that cannot track this variation incur a population MSE floor in both settings; for fixed window prediction, we also derive a per-sample lower bound. We then study one of the simplest representations that escape these floors: learning an aligned, normalized waveform and its period separately. We establish regularity of the decoupled targets under ODE assumptions and show experimentally that this approach avoids both failures in ODE systems and a PDE case study.
415. Certifying Interventional Agreement Among Observationally Equivalent Causal Models
- Authors: Sourena Khanzadeh , Daniel Platnick , Marjan Alirezaie , Hossein Rahnama
- URL: https://arxiv.org/abs/2609.32247
- Abstract:
Observationally equivalent causal models can still disagree about what happens under intervention, because interventions create inputs that never occur in observational data. We introduce Interventional Separation Selection (ISS), which repeatedly queries the true system with an admissible intervention on which the surviving candidate models disagree, discards the candidates the outcome contradicts, and stops once no intervention within a cost bound separates the survivors. If the true system is among the candidates, this stopping condition certifies that every survivor agrees with it on every admissible intervention within the bound, a guarantee that no observational learner can give, however much data it sees. The stopping condition depends only on the survivors, so it can be checked without knowing the truth. For continuous variables the candidates form an infinite version space, and mixed-integer linear programs decide the stopping condition exactly over all of it, with agreement holding up to a tolerance. On a three-digit colored MNIST causal abstraction task in which ink hue tracks digit size, plain convolutional networks trained on examples reach zero held-out error, yet disagree with shape-based labels on 26% of single-digit edits, as often as hue-based labels do. Auditing the causal abstractions of networks observed only on such images, ISS certifies what each network perceives with 13.6 interventions per image on average, and each certificate, checked against every admissible intervention, holds whenever the network’s true abstraction is among the candidates. When a network bypasses a unit that every candidate abstraction relies on, certificates covering interventions on that unit can be silently void, and twenty random validation interventions refute 69% of them.
416. RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation
- Authors: Qilang Ye , Meng Liu , Yu Zhou
- URL: https://arxiv.org/abs/2609.32224
- Abstract:
We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to
hear'',see’’,reason'', andact’’ in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: this https URL _Nav.
417. A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
- Authors: Linkai Li , Changgeng Mo , Hanlin Yu , Congxi Lu , Shangqiguo Wang , Matthew B Fitzgerald , Shan X Wang
- URL: https://arxiv.org/abs/2609.32220
- Abstract:
Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audiogram interpretation and retrieval-augmented grounding, without fine-tuning the language-model backbone. Using a 21-item rubric and an AI patient simulator, we induced a 19-rule consultation policy from 73 training cases (43 English, 30 Chinese) and evaluated the system on 58 independent simulated cases (30 Chinese, 28 English) in a pre-specified, source-blinded comparison with 17 practising audiologists. The AI audiologist outperformed human audiologists on every case (58/58; mean paired $\Delta$ = +1.35 on a 5-point composite, Cohen’s d = 1.84, $P = 4.5 \times 10^{-20}$), on 20 of 21 rubric items and in both languages. Component ablation identified the playbook as the largest contributor, offering a practical route to specialist consultation agents in low-data medical domains.
418. Fracast-0: Fractal Weight Sharing for a Time Series Foundation Model with Only 85K Parameters
- Authors: Tianxiang Zhan , Huanyao Zhang , Yuanpeng He
- URL: https://arxiv.org/abs/2609.32209
- Abstract:
Time series foundation models must preserve multi-domain breadth, probabilistic output, and multiple temporal scales, but parameter count grows when each scale receives a separate representation. We introduce Fracast-0, a probabilistic forecasting foundation model that exploits temporal self-similarity to reuse one operator across scales. A parameter-free detector extracts significant seasonal structure. The encoder applies a shared local block along a geometric dilation ladder with scale conditioning, while the decoder combines context-gathered states with an explicit seasonal future state and reuses a second block along another ladder before emitting nine quantiles. Pretraining across six corpora preserves multi-domain breadth within 85,001 parameters. On 97 GIFT-Eval configurations without per-dataset fine-tuning, Fracast-0 is the smallest of 28 evaluated checkpoints and remains non-dominated in the aggregate parameter-accuracy plane with MASE 0.808 and WQL 0.564. It uses 42.0% fewer parameters than TinyCast, whose MASE and WQL are 4.2% and 3.3% lower. These results support cross-scale weight reuse as a practical route to further time series foundation model compression.
419. Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments
- Authors: Guanghan Ning , Ping Liu , Linyi Li , Huangjie Zheng , Arjun Neervannan , Huu Nguyen , Michael Sklar , Deniz Zorlu , Nicolai Ouporov
- URL: https://arxiv.org/abs/2609.32208
- Abstract:
Automated science needs agents that can work out the rules of an unfamiliar environment by interacting with it. Interactive rule-discovery puzzles offer a controlled setting for studying this ability: an agent infers hidden rules through experimentation and uses what it has inferred to reach a stated goal. We ask what limits current language models on these puzzles and whether reinforcement learning (RL) improves performance on rules held out from training. To study both, we introduce WITNESS, a 2D grid-based puzzle environment with ground-truth ASCII observations and controlled access to rules. An agentic pipeline generates games for WitnessGym, the RL training suite, and WitnessBench, comprising public validation and private test games. The validation set separately tests new compositions of trained rule primitives and primitives absent from training. Under a shared harness, the best of 18 frontier proprietary and open-weight models solves only 24\% of private test level slots, with scores sensitive to the observation interface and agent configuration. Providing ground-truth rules raises Opus-5’s validation RHAE-L5 (relative human action efficiency over the first five levels) from 59.9 to 97.8, whereas a 27B open-weight model gains only 2.1 points and remains limited even with the rules provided. RL on WitnessGym raises the 27B model’s private test RHAE-L5 from 2.1 to 5.4 and yields a mean gain of 4.1 points on four external discovery benchmarks. Together, these results point to rule acquisition as a major difficulty for frontier models like Opus-5 while smaller models further struggle on rule-based execution, and indicate that RL on hidden-rule puzzles transfers to broader rules and real-world tasks beyond training. Benchmark is available at: this https URL
420. Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation
- Authors: Hantao Yu , Sandy Han , Udaya Ghai , Ferhat Erata , Joe Lilien , Aman Goel , Ali Torkamani
- URL: https://arxiv.org/abs/2609.32201
- Abstract:
On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.
421. CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies
- Authors: Sen Zhao , Ruiqi Kong , Zuyu Zhang , Lifeng Shen , Xinyu He , Ding Zou , Xu Zhang , Qinghua Zhang
- URL: https://arxiv.org/abs/2609.32192
- Abstract:
Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workflow topology determines which intermediate artifacts are applicable to which downstream workers and when they cease to be valid, thereby providing a structural stress dimension for sharing and isolation. Existing memory benchmarks primarily evaluate retention and retrieval, whereas multi-agent benchmarks emphasize coordination and end-to-end completion, leaving topology-conditioned memory boundaries largely unmeasured. We introduce CoMemBench, an execution-grounded benchmark for collaborative memory sharing and isolation across multi-agent workflow topologies. It constructs 800 composite workflows across four domains from source-grounded dependency graphs, with node-local specifications, verifiable artifact handoffs, native evaluators, and matched isolation challenges. CoMemBench measures workflow completion, verified node progress, required-handoff reliability, isolation robustness, and token cost. Experiments reveal a sharing-isolation trade-off: broader context improves information availability but can weaken isolation, while system rankings shift across topologies and artifact violations.
422. AI Harness: Certification under Proposal-Conditioned Information for Foundation-Model Agents
- Authors: Hailin Zhong , Shengxin Zhu
- URL: https://arxiv.org/abs/2609.32184
- Abstract:
Foundation-model agents are often modeled as policies over an observed state. In deployed systems, however, a runtime may intervene only after the model has emitted a semantic proposal, making the proposal both an action candidate and a decision-time observation generated by a history-conditioned process. We show that collapsing this structure into a state-only proposal envelope can preserve proposal coverage while destroying certifiability. In a finite robust interface, the viability kernel of the collapsed model is contained in the physical projection of the history-augmented kernel, and the collapse is lossless exactly when every proposal-conditioned collapsed fiber retains a common robust-safe intervention. This gap can be maximal even with constant-size proposal and history alphabets. The same common-action condition yields a dual result: observing the current proposal can restore robust feasibility when it separates latent modes requiring incompatible interventions. We extend these one-step results over time using exact finite beliefs and standard safety and reachability fixed points, separating indefinite operational viability from finite worst-case verified progress. Controlled model-in-the-loop tests reproduce the predicted obstructions when telemetry or effect verification is removed or intervention authority is restricted. Thus, our contribution is not a new fixed-point calculus, but a characterization of when proposal–history correlation at the model–tool boundary is necessary for certification.
423. Noisy Test-Time Reinforcement Learning for Code LLMs
- Authors: Xikai Yang , Hieu Trung Nguyen , Dunyuan Xu , Yuzhi Zhao , Jinpeng Li , Wenao Ma , Pheng-Ann Heng
- URL: https://arxiv.org/abs/2609.32172
- Abstract:
Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage. Specifically, NTRL-Code uses conservative self-denoising to obtain a cleaner semantic anchor for target estimation, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs. The policy is then optimized on the original noisy prompts with a hybrid reward that combines format validity, code similarity, and anti-repetition signals. Extensive experiments on three benchmarks, each incorporating character-level, word-level, and paragraph-level perturbations, demonstrate that NTRL-Code yields robust and consistent improvements, stabilizing the predictions of various base models. Our code is available at this https URL .
424. PastForward: Faster On-Device GUI Agents via Computational Experience Reuse
- Authors: Taehwan Park , Changmin Lee , Hayeon Lee , Taesik Gong
- URL: https://arxiv.org/abs/2609.32166
- Abstract:
Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks. However, dynamic mobile environments and user tasks make it difficult to fully utilize prior task executions without additional fine-tuning or task-specific offline exploration. To address this challenge, we present PastForward, a system that accelerates GUI agents through validated, fine-grained reuse of computational experience accumulated during ordinary task execution. During decoding, PastForward retrieves prior output sequences as device-adaptive multi-token proposals and verifies them in a single VLM forward pass. Across action steps, it uses prior GUI transitions to begin next-step inference while the device executes the current action, retains the early computation only when the predicted screen matches the observed screen, and carries reusable KV states forward. We evaluate PastForward on AndroidWorld workloads derived from real mobile usage patterns using multiple VLM backbones across server and edge platforms. On device, PastForward achieves action-step latency speedups of 1.63-2.36$\times$ while maintaining task success rates.
425. READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis
- Authors: Gerardo Pastrana , Haojun Li , Dhruv Mehta , Anoushka Vyas , Sina Khoshfetrat Pakazad , Henrik Ohlsson , John Paparrizos
- URL: https://arxiv.org/abs/2609.32123
- Abstract:
Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion, plus label-aware and language-model rerankers, under one protocol that varies supervision, pollution, and corpus scale with significance testing. Under a common channel-independent interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in embedding space, which helps far more than more sophisticated representations or language-model reasoning and holds under pollution and at full corpus scale. Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.
426. Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking
- Authors: Marco Biroli
- URL: https://arxiv.org/abs/2609.32116
- Abstract:
Best-of-$N$ jailbreaking (BoN) bypasses safeguards of aligned models by drawing $N$ independent augmentations of an unsafe prompt and sampling $M$ completions of each. Previous works have shown that the attack success rate (ASR) seems to follow a power-law in $N$, which we challenge. The exponent drifts with $N$, with an exponential crossover which is a finite-size artifact of the adversarial dataset. Little work has been done to explore the entire two-budget ($N, M$) attack surface as well as its dependence on the generation temperature $T$. We introduce a simple barrier model where each prompt has a baseline safety level and each augmentation a random thermally activated barrier. Then four numbers, each backed by an interpretable safety mechanism, determine the entire ($N, M$) attack surface. They extrapolate predictions from $N \leq 100$ to $N = 10^4$, collapse five distinct models on the same scaling function and predict ASR at different temperatures from the one they were fitted at.
427. Residual Streams Read, Recurrent States Remember: The Global Workspace in Mamba Models
- Authors: Wenlong Wang , Fergal Reid
- URL: https://arxiv.org/abs/2609.32102
- Abstract:
Can the global-workspace account of transformer representations extend to state-space language models? We fit Jacobian lenses to the residual streams and recurrent states of Mamba-1, Mamba-2 and Mamba-3, using the original 1000-prompt recipe. Joint residual–state readouts improve recovery of known intermediate concepts over the residual lens on at least five of six task families in every tested Mamba checkpoint. On Mamba-2, state alone exceeds residual and logit lenses on all six families; a normalised joint readout improves on both components on five. Temporal maps and word-list experiments show earlier content remaining state-readable as residual visibility changes. We also propose sign-guarded steering, which improves target top-five success over coordinate exchange on matched verbal-report trials in five models. Recurrent state alone supports this verbal access. These gains do not extend consistently to relational answers: guarded edits often output the edited concept itself, and Mamba-3’s joint edits can disrupt successful state-only redirection. Recurrent state thus provides a complementary carrier of workspace content, whose recovery, persistence and causal uses require separate measurements.
428. GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games
- Authors: Dhananjay Ashok , Adam Shen , Aslan Huo Feng , Chinmay Khanna , Jun Rui Huang , Raghav Sarmukaddam , Surendira Balaji Natarajan , Xiaotong Cui , Xincan Zhang , Thomson Yen , Hongseok Namkoong , Jonathan May , Jesse Thomason
- URL: https://arxiv.org/abs/2609.32093
- Abstract:
Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluates task execution on a collection of 5 distinct game series. Agents are allowed access to dedicated training games but are provided no demonstrations, documentation, or rewards. Agents must ground themselves in the environment through self-directed exploration and by inferring actionable knowledge from their own experience. At test time, agents must complete short-horizon tasks that evaluate their ability to navigate, interact, and engage with game-specific mechanics in unseen games. Out-of-the-box frontier models complete fewer than 50% of the 500 tasks due to failures in multimodal grounding, establishing that self-improvement methods have room to push performance. We demonstrate that contemporary approaches to self-improvement are lacking, with world modelling and autonomous skill discovery failing, and a novel strategy that uses curiosity-based exploration to write guides achieving only partial success. GameBoyWorlds-Playthrough tests end-to-end game completion in two fan-made Pokémon games. We show that while frontier models have been pre-exposed to official releases such as Pokémon Red, they lack essential information on the games in our testbed. Instead of relying on their parametric knowledge to succeed, agents must learn from their own experience and autonomously improve over the course of the playthrough. We show that a sophisticated agentic pipeline with multimodal memory and hierarchical subgoals fails to reach even the first major milestone in both games, establishing GameBoyWorlds as an ambitious target for self-improving agents.
429. Memory as Middleware for Self-Improving AI Agents
- Authors: K. R. Jayaram , Vatche Isahagian , Vinod Muthusamy , Gegi Thomas , Punleuk Oum , Gaodan Fang , Ashwath Vaithinathan Aravindan
- URL: https://arxiv.org/abs/2609.32091
- Abstract:
AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience. The dominant fix is \emph{bespoke memory}—retrieval, persistence, and learning logic hand-wired into one agent and bound to one storage engine. This creates a fragmented landscape where memory cannot be swapped, shared, isolated, or reasoned about independently of the agent that owns it. We argue that this is a middleware problem: agent memory deserves a first-class, pluggable layer, just as data access, messaging, and persistence each became middleware concerns. We develop this vision through six systems challenges: two-sided pluggability, host-native interposition, multi-tenant isolation, write-path consistency, federated sharing with provenance, and lifecycle governance. We present ALTK-Evolve, a reference implementation of memory middleware for self-improving agents, and use it to motivate a broader research agenda for future memory middleware.
430. Toward Interactive Understanding of Code APIs
- Authors: Dhananjay Ashok , Jesse Thomason , Jonathan May
- URL: https://arxiv.org/abs/2609.32081
- Abstract:
Empowered by advances in Language Model agents, systems have made substantial strides in code generation and understanding. However, these approaches often rely on read access to the relevant code, an assumption which does not hold when dealing with external APIs. In this work, we introduce the PAU (Python API Understanding) benchmark, where we provide models with black-box, API-level access to code snippets. Models must query the API with exploratory inputs and draw insights from the resulting outputs, with the goal of describing the snippet’s true functionality. By treating the code snippets as external tools that must be understood via interaction alone, PAU studies the more general problem of unsupervised tool understanding, specifically for tools implemented as Python methods. Despite recent progress in coding agents, even frontier models struggle to achieve high performance on PAU, with the best model (Claude-4-Opus) failing to understand over 45% of the PAU test set. An investigation into the common error modes reveals that models are overconfident; they often overrate the quality of their current hypothesis, leading to insufficient exploration and premature termination. Finally, we take inspiration from the Asymmetric Actor Critic (AAC) paradigm, frequently used in robot learning, to post-train models for interactive code understanding. Models trained with AAC conduct more active exploration of the APIs, with an AAC-tuned Qwen3-8B model matching the performance of GPT-5-mini.
431. Contract monitoring: governing AI via separation of powers
- Authors: Enric Boix-Adsera
- URL: https://arxiv.org/abs/2609.32061
- Abstract:
We propose an AI safety framework that binds worker agents to contracts specifying their permitted actions. We show how these contracts can be enforced and specified by assigning distinct responsibilities to monitor agents and judges, and asymmetric computational resources to monitors and workers. Our framework allows us to empirically measure statistical safety guarantees. The framework applies to a wide range of settings, including code security and escape-the-box scenarios.
432. EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
- Authors: Bhavyateja Potineni , Lohit Giri , Anu Jain , Vadim Kutsyy , Rajasekhar Pentakota
- URL: https://arxiv.org/abs/2609.32049
- Abstract:
As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Usage-Modulated Personalized PageRank (U-PPR), where transition probabilities adapt via Hebbian plasticity to promote persistent entities into high-centrality Epistemic Macro-Hubs; (2) Consolidation-Activated Topology Decay (CATD), which scales retention half-life by topological load-bearing weight rather than wall-clock recency, protected by a cold-start grace period (N_grace >= 4); (3) Directed SUPERSEDES DAG filtering to suppress obsolete state during fact mutations; and (4) Triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF). Evaluating on all 1,982 QA pairs across 10 long-term conversations in the LoCoMo benchmark, EngramRAG achieves +38.9% relative improvement in Recall@5 (53.21% vs. 38.29%, p < 0.001) and +43.1% in MRR (0.4203 vs. 0.2937) over dense vector RAG, significantly outperforming Okapi BM25 (48.66%) and isolated static graph retrieval (8.50%). On temporal reasoning, EngramRAG reaches 62.33% Recall@5 (+16.67 points over dense vectors). In controlled mutation tests, SUPERSEDES suppresses split-brain hallucinations from 70.0% to 0.0%, while 90-day simulations show 100.0% scaffolding retention under a 26.21ms interactive retrieval reflex.
433. Receiver-Conditioned Latent Communication gives 94% CacheBack
- Authors: Maximillian Rossi , Prajwal Raghunath , Haoqing Xuan , Yusen Zhang , Eugene Wu
- URL: https://arxiv.org/abs/2609.32046
- Abstract:
Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy and latency. However, a full KV cache grows linearly with both the context an individual agent processes, and the number of agents that coordinate together. This raises memory and context costs, often far exceeding available GPU resources and context window sizes. Our key observation is that agents need only send what the receiving agent requires for its local task – which we call receiver-conditioned communication. The receiver agent passes the sender a small description of its information needs, which serves to filter and compress the sender agent’s KV cache. CacheBack is a simple, robust, training-free instance of receiver conditioning based on the sender’s attention weights. On FanOutQA, CacheBack with Qwen 3 removes 75% of the state the agent would otherwise receive, improving accuracy by 14.7 percentage points and reducing median task-completion latency by 3.2x relative to text communication. We show comparable improvements across model families that span dense Transformers, Mamba-attention hybrids, and sliding-window attention.
434. Reasoning Concentrates Errors, and Self-Consistency Never Notices
- Authors: Asaad Althoubi
- URL: https://arxiv.org/abs/2609.32035
- Abstract:
Self-consistency assumes that independent samples disagree when a model is unsure, so agreement is evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, we show that reasoning concentrates a model’s errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine after restricting both arms to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the distinct answers produced to 0.43-0.65 of the non-reasoning count; where it is bounded, both arms hold an identical option set and reasoning concentrates mass on it instead, which no positional prior can explain at fixed weights. The aggregate cost is smaller than the mechanism predicts, because reasoning also shrinks the set of problems where answer diversity can decide anything, in ten of ten cells and by 2.7x; normalized for available headroom, both arms convert a quarter of it in domain. Confidence weighting does not recover what is left. Across 280 method-dataset-model combinations on eight models and five benchmarks, not one beats plain majority voting after correction; weighted voting agrees with it on 98.5% of problem-method pairs and is right 56.3% of the time on the rest; and a signal’s direction can invert within fixed weights, with answer log-probability predicting correctness when reasoning is off and error when it is on. A learned six-signal combination gains nothing out of domain. Confidence signals should be evaluated on decisions, not on discrimination.
435. A Benchmark for LLM’s Understanding of Middle School and High School Science Topics
- Authors: Noah L. Schroeder , Yessy Eka Ambarwati , Yuji Zhang , ChengXiang Zhai
- URL: https://arxiv.org/abs/2609.32020
- Abstract:
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs’ performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs’ capacity for interactive, evidence-based feedback in educational scenarios.
436. Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
- Authors: Valeriy Vyaltsev , Anton Andreychuk , Taisia Zlotnikova , Konstantin Yakovlev , Aleksandr Panov , Alexey Skrynnik
- URL: https://arxiv.org/abs/2609.32019
- Abstract:
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same context, independently sampling from per-agent distributions can recombine locally valid choices into incompatible joint actions. This failure can arise from the final sampling mechanism even when the per-agent action distributions are learned correctly. DMM (Decentralized Master-Mind) addresses this by replacing one-shot action sampling with discrete, iterative refinement of action intents across communication rounds, inspired by denoising in diffusion models. Agents initialize random action intents and refine them through local communication, coupling their choices before commitment. DMM is pretrained with imitation learning on expert MAPF solutions and further optimized with MICPO, a critic-free group-relative reinforcement-learning method designed for multi-agent, multi-round action refinement. DMM generally achieves higher success rates and lower solution costs than the evaluated learnable baselines. On 1,600 MovingAI tasks, DMM fine-tuned with MICPO solves 1,598, the highest coverage among the evaluated methods, while achieving solution costs close to those of the strongest baselines. DMM also scales to over one million simultaneously acting agents in obstacle-rich environments. These results show that round-level intent refinement can improve joint-action coordination while preserving decentralized execution.
437. What Does the Rank Buy? A Spectral and Distributional Analysis of Low-Rank Adaptation
- Authors: Babak Barazandeh
- URL: https://arxiv.org/abs/2609.32002
- Abstract:
The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets—the idealization of the weight decay and norm control used in practice—this intuition breaks down. The reason is structural: under such budgets, the updates LoRA can reach are exactly the matrices of rank at most $r$ inside a nuclear-norm ball, and every complexity and displacement functional we analyze is maximized over this set by a rank-one update—so the rank cap never binds. The consequences follow directly. The linear-readout model class we study is identical for every $r \ge 1$, its Rademacher complexity carries no dependence on $r$, and the distance the adaptation can move the source distribution obeys a rank-independent upper bound that we show is sharp. If rank does not control capacity, where does it act? We identify two places. Statistically, replacing the per-factor budgets with a joint budget on the product restores a data-dependent, rank-sensitive complexity bound—though the gain appears only for well-spread feature distributions, and the worst case remains rank-free. Spectrally, rank sets the price of adaptation: canceling the leading singular directions of the pretrained weight requires both sufficient rank and sufficient budget. We bound the smallest rank achieving a desired source–target alignment, with upper and lower bounds that match under two-sided spectral decay. Together, these results recast rank as governing which updates are reachable and what cancellation costs—not how much capacity the model has.
438. SenseAgent: An LLM Agent for Adaptive Cross-Domain IMU Sensing
- Authors: Tianya Zhao , Chuan Liu , Xuyu Wang
- URL: https://arxiv.org/abs/2609.32000
- Abstract:
Deep learning has improved inertial measurement unit (IMU) sensing for mobile and wearable applications. However, an IMU model trained in one domain often becomes unreliable when it is used with a new user, device, or body position. Existing methods usually treat this problem as a static model-design task: they pretrain a stronger representation, add data augmentation, or select one adaptation method before deployment. In practice, the target domain is only gradually observed, labels are scarce, and different domain shifts require different sensing actions. This paper presents SenseAgent, an LLM-guided sensing agent for cross-domain IMU activity recognition. Instead of asking an LLM to classify raw IMU signals, SenseAgent uses the LLM as a runtime planner over sensing tools, source-domain experience memory, online target memory, and verifiers. The agent builds a label-free diagnosis report from the target stream and uses it to decide whether to keep raw inference or invoke specialized tools, including gravity-aware sensing, prototype transfer, and style normalization. Verifiers check source calibration, target-memory reliability, and no-harm criteria before accepting high-risk tool decisions. SenseAgent also supports scarce feedback without retraining the backbone or replacing the label-free route. This design converts cross-domain IMU sensing from a fixed inference pipeline into a closed-loop sensing process that diagnoses target shifts, selects suitable sensing actions, and rejects unsafe adaptations. We evaluate SenseAgent across multiple IMU datasets and deployment shifts. Results show that its verified route selection improves cross-domain sensing, especially under harder placement and compound shifts, and further benefits from limited user feedback.
439. CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing
- Authors: Tianya Zhao , Chuan Liu , Xuyu Wang
- URL: https://arxiv.org/abs/2609.31990
- Abstract:
Wi-Fi channel state information (CSI) has enabled device-free sensing applications such as human activity recognition. However, CSI sensing models remain brittle in cross-domain deployment, where changes in users or environments can produce incorrect predictions. Existing solutions usually treat this problem as an offline model-design problem, by pretraining a stronger representation or applying one fixed adaptation method to the entire target domain. In practice, labeled target data are scarce and different classes may fail in different ways under the same domain shift. To address this, we propose CSI-Agent, an evidence-seeking LLM agent that reformulates cross-domain CSI adaptation as a deployment-time decision-making problem. Rather than processing raw CSI or making sample-level predictions, CSI-Agent summarizes target-domain behavior into sensing-grounded class-level evidence. It establishes a strong target-adaptive default from complementary CSI views and uses an LLM planner to determine whether each class should retain the default or invoke a specialized action. Deterministic verification and bounded execution further reduce unreliable interventions. We evaluate CSI-Agent on four public datasets using five cross-domain splits covering device, user, environment, and compositional shifts. Under 1-shot adaptation, CSI-Agent achieves the best target-domain performance across all splits and improves the average Macro-F1 by about 16\% compared to the strongest baseline method.
440. Goal-Persistent Coding Agents as Scientific Performance Engineers: A Fixed-Radius Nearest-Neighbor Case Study
- Authors: Xiangyang Ju
- URL: https://arxiv.org/abs/2609.31980
- Abstract:
Coding agents can pursue persistent objectives across many tool-use turns, but evidence that general-purpose agents can conduct rigorous scientific performance engineering remains limited. We present a repository-scale case study in which off-the-shelf Codex and Claude Code agents optimize fixed-radius nearest-neighbor (FRNN) search for particle tracking. Starting from a PyTorch-dependent CUDA implementation, the agents follow an executable goal that specifies exact-correctness tests, profiling requirements, and acceptance criteria without prescribing code transformations. In the primary sequential trajectory, they autonomously remove the PyTorch dependency and conduct hypothesis-driven optimization experiments. The resulting standalone C++/CUDA library exactly reproduces the targeted reference result. Its synchronous NumPy interface achieved 1.6-fold speedup over the original GPU-resident PyTorch interface, despite including host transfers. Similar speedups were observed across different GPU architectures and software stacks. An independent optimization rerun followed a different sequence of hypotheses and reached even better performance on the target workload. These results show that goal-persistent coding agents can act as experimental performance engineers, and that executable scientific contracts are needed both to guide and to validate their optimization.
441. Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination
- Authors: Ben Rachmut , Ning Zhang , Yevgeniy Vorobeychik , William Yeoh
- URL: https://arxiv.org/abs/2609.31963
- Abstract:
Large language models (LLMs) are increasingly deployed as autonomous agents in multi-agent systems, yet their ability to reliably execute distributed coordination protocols remains poorly understood. While AgentsNet, a benchmark framework for distributed coordination among LLM agents, enables such coordination, granting full reasoning autonomy often leads to inconsistent or degraded performance in complex domains. We hypothesize that coordination can be improved by regulating agent autonomy through symbolic guidance derived from established algorithms. To investigate this, we introduce the \emph{Symbolic Guidance Taxonomy (SGT)}, which characterizes a spectrum of autonomy ranging from open-ended natural language reasoning to fully prescribed algorithmic execution, with intermediate levels providing partial pseudocode guidance. Our results show that intermediate autonomy levels consistently outperform both unguided agents and fully prescriptive specifications. These findings identify autonomy regulation as a key design principle for LLM-based distributed coordination.
442. BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering
- Authors: Xingbo Du , Fadli Aulawi Al Ghiffari , Leonard Song , Loka Li , Duzhen Zhang , Zixiao Wang , Xiuying Chen , Le Song
- URL: https://arxiv.org/abs/2609.31939
- Abstract:
Agentic biomedical machine learning (ML) draws on complementary advances in biomedical evidence acquisition and executable program search. Existing systems connect aspects of these capabilities, but coordinating them throughout program search remains challenging. New evidence must guide candidate construction, execution outcomes must inform subsequent discovery and reuse, and validation demands must fit the search budget. We introduce BioDyad, which couples biomedical discovery and ML engineering through two hierarchies within Monte Carlo graph search. Its scientific hierarchy combines prior biomedical guidance with iterative discovery, then links biomedical plans to execution outcomes in memory for reuse across candidates. Its engineering hierarchy moves candidate programs from smoke execution, through train/validation evaluation, to full-data retraining. We evaluate BioDyad on the 76-task BioXArena benchmark under a two-hour per-task budget with three matched LLM backends. It achieves the highest penalized all-task score and task success rate among four agent methods and a one-shot baseline under each backend. These results support coordinating biomedical discovery and ML engineering to integrate external knowledge into executable programs across heterogeneous biomedical tasks.
443. Improving Medical Calculation of LLMs with Embedded Coding
- Authors: Tianshi Ming , Yingying Zhang , Xian Wu
- URL: https://arxiv.org/abs/2609.31908
- Abstract:
Large Language Models (LLMs) perform well on medical examinations and question-answering benchmarks, but remain unreliable on medical calculation tasks that require exact numerical outputs. These calculations support high-stakes decisions such as medication dosing, organ-function assessment, and prognostic scoring, for which even small errors can have serious clinical consequences. We introduce MedCode, a framework that improves medical calculation by training LLMs to generate embedded executable code. Given a clinical context, the model identifies the relevant calculator, extracts its input variables, and produces a script that delegates arithmetic operations to a deterministic interpreter. Executing the script returns the calculated value together with an explanation and the appropriate unit. We construct supervised fine-tuning (SFT) and preference datasets from the MedCalc benchmark and additionally curate a dataset for calculation tasks in Intensive Care Unit (ICU) scenarios. We further propose weighted Direct Preference Optimization (wDPO), which adaptively emphasizes preference pairs that are difficult for the model to distinguish. Experiments with LLaMA3-8B, Qwen2.5-7B, and Mistral-7B show absolute accuracy gains of 20–30 percentage points, demonstrating the effectiveness of embedded code generation for medical calculation.
444. EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
- Authors: Mukul Singh , Mansi Uniyal , Devin Devlin , Wen Xie , Big Thadawasin , Ritam Dutt , Vivian Lai , Hyeonsu B. Kang
- URL: https://arxiv.org/abs/2609.31906
- Abstract:
Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.
445. Choir: An Open Protocol for Distributed Multi-Agent Autoformalization
- Authors: Yidi Qi , Melanie Weber
- URL: https://arxiv.org/abs/2609.31903
- Abstract:
AI agents can now formalize entire textbooks and major theorems in proof assistants such as Lean, but current efforts are typically centralized: a single team runs all agents and bears the full computational cost. We introduce Choir, an open protocol for distributed formalization. Choir decomposes a project into tasks that can be completed by independent contributors, each running their own agent with their own LLM subscription, while coordinating entirely through the project’s GitHub repository. To support open participation, every contribution is checked by a deterministic gate before merge. Choir supports Lean 4, Isabelle, and Rocq, and is open source and modular, allowing projects to replace individual components or extend the protocol.
446. Context-dependent agent evaluation with orthogonal equilibrium learning
- Authors: Haorui Ma , Zehua Zang , Jiangmeng Li , Yi Li , Fanjing Xu , Stefan Feuerriegel
- URL: https://arxiv.org/abs/2609.31897
- Abstract:
Many applications require to evaluate agents under contextual information (e.g., a prompt, task, or user group). We study how to perform such context-dependent agent evaluation from offline feedback. Existing score-based models for this purpose (e.g., Bradley-Terry) impose a transitive preference ordering, which fails to reflect collective preferences when human judgements are heterogeneous. Inspired by social choice theory, we frame evaluation as a contextual game between two players, each selecting a distribution over agents as the strategy to receive greater collective preference than the other. Then, the support of the Nash equilibrium defines a context-specific set of winners. However, learning context-specific equilibria from offline logs is difficult because each context reveals human feedback on only a subset of agents, and, hence, a naive plug-in estimator can therefore be biased. To address these challenges, we propose NashEval, a general framework for robust contextual equilibrium learning. NashEval first constructs debiased estimates of the contextual payoff matrix that characterizes the game. NashEval then learns the context-to-equilibrium mapping with a tailored orthogonal loss, which avoids the need to solve a separate game for each context. We show theoretically that errors in estimating the nuisance functions underlying the payoff matrix affect the risk of the learned equilibrium (i.e., exploitability) only through higher-order terms. Across various experiments, NashEval improves robustness of equilibrium learning and consistently identifies the set of top-performing agents across contexts.
447. COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents
- Authors: Hongji Pu , Ruixiang Tang , Yongfeng Zhang
- URL: https://arxiv.org/abs/2609.31874
- Abstract:
Existing agent memory frameworks mainly create memory through an agent’s interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the “what if” question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.
448. IndustryLLM: Failure-Driven LLM Training for Industrial Procurement
- Authors: Liang Ding (Project Lead), Zhiang Xu , Yuyang Sheng , Bin Chen , Songlin Bai , Run Zhu , Dingjun Wu , Hui Xu , Yandi Wang , Fulin Shi , Leilei Gan , Linlin Yu , Qihuang Zhong , Keqin Peng , Yalong Li , Chengfu Huo
- URL: https://arxiv.org/abs/2609.31871
- Abstract:
Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained from Qwen3.5-35B-A3B-Base (35B total parameters with ~3B activated per token, with the vision encoder frozen). Rather than relying on generic text scaling, we introduce a failure-driven adaptation recipe spanning continued pre-training (CPT) and supervised fine-tuning (SFT). CPT leverages a curated ~100B-token corpus integrating 5B tokens of national standards (e.g., GB/T) and technical archives, 10B tokens of de-identified real-world industrial transaction and inquiry records, and 60B tokens of general replay. To overcome register mismatch and factual brittleness, we systematically reconstruct an estimated 20B-token domain subset via multi-register rewriting across 10 genres and 8 writing styles, confidence-routed minimal factual editing, and error-targeted QA synthesis (resolving colloquial typos like ‘42-luo-mu’ -> 42CrMo, expanding ambiguous codes like ‘16674’ -> GB/T 16674, and clarifying conflicting dimensional specs). For downstream deployment, we formalize an evidence-gated constraint-evaluation interface enforcing three-valued logic where unverified product evidence remains unknown rather than satisfied. Offline evaluations demonstrate consistent gains on procurement-query structuring (+2.97 percentage points in exact match, 95% CI [2.11, 3.86] in No-Think mode), while randomized online A/B experiments in production yield substantial improvements (+4.25% GMV, +8.3% satisfied inquiries) alongside a latency reduction from 6-7 s to 1.5 s. Model weights and configs are released at this https URL .
449. Metro-WM: Long-Horizon Latent Planning with Realisable Sub-Goals
- Authors: Royson Lee , Fady Rezk , Titouan Parcollet , Timothy Hospedales , Cristina Cornelio
- URL: https://arxiv.org/abs/2609.31868
- Abstract:
Model-predictive control with Joint-Embedding Predictive Architectures (JEPAs) provides a strong zero-shot goal-reaching planner, but it is only effective over short planning horizons. Hierarchical extensions attempt to bridge this gap by learning a macro planner to predict intermediate latent sub-goals to guide the micro planner. In this work, we demonstrate that unconstrained latent sub-goal prediction is fundamentally flawed. A rigorous evaluation reveals that a leading state-of-the-art macro planner routinely emits physically unrealisable sub-goals. To resolve this, we introduce Metro-WM, a hierarchical framework that issues sub-goals by retrieving genuine states from prior experience rather than generating ungrounded latent vectors. Specifically, Metro-WM constructs a graph whose vertices are observed frames from offline expert demonstrations or random-action trajectories, allowing frames from different episodes to be connected and stitched into routes to the goal. Planning over the full graph also makes the system highly robust to execution errors: if the micro planner drifts off course, Metro-WM instantly finds a new optimal path from the current state. Our experiments show that Metro-WM achieves superior long-horizon success rates of up to 37.33 percentage points over the next best hierarchical approach while being up to 10.9 times faster, requiring both 13-56 times less offline compute and fewer tuned hyperparameters. Additional analysis reveals that Metro-WM finds shorter paths than the offline demonstrations, outperforms an oracle relying on the query’s own demonstration, and maintains robust performance under extremely sparse dataset conditions.
450. LLM Judge Validation Under Sparse Overlap: From Inference to Design
- Authors: Junxuan Li , Arko Mukherjee , Soumyabrata Pal
- URL: https://arxiv.org/abs/2609.31857
- Abstract:
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
451. DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution
- Authors: Chengkai Xu , Jiaqi Liu , Yicheng Guo , Peng Hang , Jian Sun
- URL: https://arxiv.org/abs/2609.31814
- Abstract:
Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on this https URL
452. Working with AI: A Design Framework for Human-AI Collaboration
- Authors: Yuqian Lu , Regina Lee , Rui Zhou , Lixin Jiang , Andrew McDaid , Amy Lawrence
- URL: https://arxiv.org/abs/2609.31793
- Abstract:
Artificial Intelligence (AI), particularly GenAI, is becoming an increasingly important part of modern work. In industrial settings, AI can support decision-making, automate routine activities, assist humans, and improve productivity. However, successful AI adoption depends on more than what the technology can do. It also depends on how people experience and work with it. This raises an important question: how should human-AI collaboration be designed so that it works well for both people and organisations? This white paper addresses that question by presenting a practical framework for designing human-AI collaboration. The framework considers the human, the AI system, the task, the organisation, and the wider societal environment. It explains what effective collaboration looks like, what conditions influence it, what requirements should be met, and what design decisions organisations should consider. The report also includes a human-AI collaborative assembly system with cobot use case to demonstrate how the framework can be applied in practice. The use case shows how design requirements can be translated into specific collaboration features and evaluated through a case study. The aim of this white paper is to provide a clear and practical guide for designing human-AI collaboration that is effective, human-centred, and responsible.
453. ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts
- Authors: Liyu Hou , Yuan Wu , Yi Chang
- URL: https://arxiv.org/abs/2609.31792
- Abstract:
While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We call the latter pattern Failed Persistence. To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict tasks spanning four conflict families, four structural configurations, and two prompt conditions. Across all eight VLAs, invalid premises reduce original goal completion by at least 17.3 percentage points, with the reduction reaching 56.2 percentage points for OpenVLA. Crucially, even when models succeed on premise-consistent tasks and fail on their matched conflict tasks, they often continue to approach the original targets, retain early trajectory structure, and show limited action magnitude suppression. Failed Persistence therefore recurs across the evaluated models. Explicit premise checking does not consistently produce selective and coordinated behavioral changes. These findings show that terminal failure alone establishes neither behavioral disengagement nor refusal and that outcomes alone are insufficient for VLA evaluation. Experimental data and additional details are available on the project page: this https URL
454. CP-Agent: A Harness-Engineered Agent for Crystal Plasticity Simulation Workflows
- Authors: Samuel Onimpa Alfred , Abhishek Kumar , Veera Sundararaghavan
- URL: https://arxiv.org/abs/2609.31790
- Abstract:
Crystal plasticity (CP) simulations predict the mechanical behavior of polycrystalline metals, yet their routine use is hindered by the manual effort of configuring heterogeneous tools, orchestrating multi-step data pipelines, and calibrating constitutive parameters against experiments. These bottlenecks impede productivity in systematic parameter studies, motivating interest in automated workflows. This study presents CP-Agent, a harness-engineered LLM-based agent that autonomously executes complete CP modeling workflows from natural-language tasks. Operating under the ReAct paradigm, the agent reasons about tool selection and sequencing while delegating numerical search to established optimizers. The harness comprises a minimal system prompt, typed tool definitions, a dispatcher, and a safety-bounded iteration loop, encoding domain knowledge through tool schemas rather than hard-coded logic. CP-Agent is demonstrated on four case studies: calibrating four slip parameters of additively manufactured stainless steel 316L against tensile data; validating the workflow against published copper benchmarks, reproducing stress-strain and texture evolution; recovering the initial crystallographic texture of copper, where the agent correctly identifies a diffuse initial texture; and reproducing the multi-pass rolling texture evolution of a Mg-Zn-Ca alloy, where the agent chains five deformation passes and recovers the experimentally observed weakened, split basal texture. In all cases, the agent inferred the correct execution sequence from the task statement, robustly across repeated runs, and delivered physically interpretable results. This work establishes harness engineering as a systematic approach to automating CP modeling workflows while maintaining physical interpretability and auditability through visible reasoning traces.
455. Witeness Overlap: Directional Provenance Inside Open-Weight Model Families
- Authors: Siyuan Li , Haoxuan Zeng , Xin Luo , Fernando Jia , Florence Li , Zhengyang Geng , Zico Kolter , Tai Sing Lee , Tianqin Li
- URL: https://arxiv.org/abs/2609.31784
- Abstract:
Open-weight models are often released, fine-tuned, aligned, merged, and re-released, making provenance audits ask not only whether checkpoints are related, but also which checkpoint came first. Many existing model-provenance methods are designed for a base-known audit setting: given a victim or source model, they test whether a suspect model is related to it. Although these audits are framed as source-to-suspect tests, their underlying evidence is often symmetric, relying on representation similarity, weight similarity, behavioral fingerprints, or correlation statistics. Symmetric pairwise comparisons can detect relatedness, but they cannot by themselves orient relationship between checkpoints A and B. We therefore introduce a local geometric comparison: instead of comparing two checkpoints directly, we add a third same-family checkpoint as a witness and compare the geometry around each candidate endpoint. Direction is inferred by asking which candidate behaves more like a branching parent. Motivated by this idea, and by the empirically observed asymmetry between parent-anchored and child-anchored witness-overlap distributions, we propose Witness Overlap, a prompt-free, training-free white-box test for directional provenance. On 176 LLM checkpoints from 16 families, our one-witness test orients 95.3\% of parent-child decisions using Frobenius cosine. We further evaluate root identification, sibling discrimination, generalizations to VLM and diffusion families, and chain-structured ordering. The signal is robust to weight noise and sparse pruning, with a proposed SVD weight reduction variant showing greater robustness than Frobenius cosine.
456. SMARtCARE: Privacy-Preserving Agentic AI Systems for Bounded-Autonomy Clinical Decision Support
- Authors: Srini Ramaswamy , Deveeshree Nayak
- URL: https://arxiv.org/abs/2609.31763
- Abstract:
Long-context clinical AI systems can miss relevant patient history when prior admissions fall outside the active reasoning context. In ICU monitoring, this can cause early vital-sign drift to appear nonspecific even when it resembles a prior deterioration pattern. SMARtCARE addresses this gap through a four-state clinical decision-support architecture: Stable, Meta-cognitive, Assisted, and Regulated (Revoked). Rather than automatically retrieving prior records, SMARtCARE uses a lossy six-channel fingerprint of the patient’s prior trajectory. When current drift matches that fingerprint and the prior record is absent from context, the system raises a Meta-cognitive escalation for clinician review; full retrieval occurs only through clinician action in the Assisted state. A patient-identity guard is designed to enforce correct attribution across data loading, logging, and audit layers. Evaluation combines a synthetic Monte Carlo study that validates the state-transition logic and estimator stability, not clinical performance, with real-data runs on both the MIMIC-III and MIMIC-IV Clinical Database Demos. On MIMIC-III, one prior-pattern recurrence was identified among 14 two-admission patients; on MIMIC-IV, the same pipeline produced no fingerprint matches among 9 two-admission patients, which illustrates a key limitation of a fixed canonical pattern library. Across both runs all logged decisions were fully traceable and correctly attributed. The results support SMARtCARE as a traceable, privacy-aware mechanism for surfacing middle-context risk; they are not a clinical efficacy claim.
457. FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets
- Authors: Srinjay Sarkar , Prakhar Kaushik , Soumava Paul , Alan Yuille
- URL: https://arxiv.org/abs/2609.35770
- Abstract:
Realistic and editable animal fur reconstruction from multi-view images is challenging due to fine-scale detail, self-occlusion and obfuscation, and, unlike human hair, the lack of animal-fur datasets. Fur usually covers most of an animal’s body, with large inter-species and intra-species variability. We present FurE, an efficient strand-based animal fur reconstruction method that recovers a per-strand, editable groom by optimizing a root-conditioned latent field, decoded into strand geometry via a PCA-based decoder. We reconstruct a defurred animal body using local fur-thickness cues from a surface-constrained Gaussian Frosting representation together with part-based priors. We further show that a PCA-based decoder learned from human-hair strand data can alleviate animal-data scarcity while enabling substantially faster optimization. FurE achieves a 10x speedup in strand training over current SOTA dense per-strand optimization while retaining strand fidelity and generalizing across synthetic and real-world sequences, with quantitative and qualitative validation despite the reduction in training time.
458. Telescopic Language Models
- Authors: Zhilin Guo , Boqiao Zhang , Hakan Aktas , Kyle Fogarty , Nursena Koprucu Aslan , Wenzhao Li , Canberk Baykal , Albert Miao , Siyu Hong , Yixiao Liu , Adam Wu , Ashish Kumar Singh , Sakar Khattar , Chenliang Zhou , Weihao Xia , Cristina Nader Vasconcelos , Cengiz Oztireli
- URL: https://arxiv.org/abs/2609.35769
- Abstract:
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
459. Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
- Authors: Yijia Fan , Ziqi Huang , Zhongang Cai , Yan Li , Zimo Wen , Wanqi Yin , Haiwen Diao , Ziwei Liu
- URL: https://arxiv.org/abs/2609.35767
- Abstract:
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
460. TokenCast: Forecasting Token Consumption During LLM Agent Execution
- Authors: Chaoqian Ouyang , Ling Yue , Libin Zheng , Huanghui Guo , Shengxiang Xu , YiShu Wang , Ran Li , Jian Yin , Shaowu Pan , Shimin Di
- URL: https://arxiv.org/abs/2609.35760
- Abstract:
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast’s mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at this https URL .
461. How to Loop MoE: Flatten the Experts, Untie the Attention
- Authors: Shouren Wang , Chuang Ma , Mohsen Hariri , Debargha Ganguly , Wang Yang , Xiaoqing Tong , Qianying Liu , Xiaotian Han , Vipin Chaudhary
- URL: https://arxiv.org/abs/2609.35751
- Abstract:
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at this https URL .
462. KV-streams for Efficient Compaction in Agentic Reinforcement Learning
- Authors: Emiliano Penaloza , Dane Malenfant , Dheeraj Vattikonda , Roger Creus Castanyer , Siddarth Venkatraman , Abhay Puri , Jonathan Light , Matthew James Sargent , Augustine N. Mavor-Parker , Massimo Caccia , Lucas Caccia , Glen Berseth , Esmeralda S. Whitammer , Alessandro Sordoni , Minseon Kim , Marc-Alexandre Côté , Laurent Charlin , Guillaume Lajoie
- URL: https://arxiv.org/abs/2609.35750
- Abstract:
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
463. Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
- Authors: Huaiyuan Qin , Muli Yang , Gabriel James Goenawan , Shiqi Huang , Min Kass Chong , Wahyu Wiratama , Peng Hu , Chen Gong , Wu Liu , Xi Peng , Chun Jian Ho , Hongyuan Zhu
- URL: https://arxiv.org/abs/2609.35745
- Abstract:
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention’s token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.
464. X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
- Authors: Prithwish Dan , Chenyang Ma , Wei Zhan
- URL: https://arxiv.org/abs/2609.35715
- Abstract:
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments—a 22-DoF hand on two different arms and a parallel-jaw gripper—and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.
465. A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
- Authors: Fred Xu , Thomas Markovich , Florence Regol , Yizhou Sun
- URL: https://arxiv.org/abs/2609.35703
- Abstract:
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filters capture structural variation, and a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. The resulting doubly-spectral stochastic (DSS) expansion supplies task-matched readouts from one representation: the mean coefficient encodes class evidence for the energy-based OOD score, the higher-order coefficients encode structured logit variation, and quadrature averaging over the chaos coordinate defines the single predictive distribution used for prediction and calibration. A capacity theorem shows that, under a full-rank feature assumption, a restricted subfamily matches the chaos coefficients of any Gaussian-latent random graph signal, with exponentially decaying truncation error under a growth condition; the task-level claims are established empirically. DSS-GNN has two deployment modes: standalone, or as a residual branch beside a deterministic encoder (DSS-Hybrid). Standalone DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD detection, and the strongest shifted accuracy on all 7 GOOD concept-shift benchmarks under standard empirical risk minimization (ERM). Cross-evaluating both modes on all three tasks shows that each remains effective on the other’s tasks, with documented exceptions, and yields explicit deployment guidance.
466. Distillation Defenses Easily Break After Reinforcement Learning
- Authors: Shidan Javaheri , Alexander Panfilov , Oliver Britton , Yarin Gal , Yonatan Gideoni
- URL: https://arxiv.org/abs/2609.35699
- Abstract:
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., “distill”) their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security – some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
467. Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
- Authors: Li Zhang , Chuqin Geng , Mark Zhang , Chen Yang , Luke Zhang , Haolin Ye , Xujie Si
- URL: https://arxiv.org/abs/2609.35686
- Abstract:
Mechanistic interpretability (MI) aims to explain a model’s behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model’s behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model’s particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model’s errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model’s full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model’s failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
468. MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution
- Authors: Prasoon Dev , Anirudh Sankar , Vasudeva Varma
- URL: https://arxiv.org/abs/2609.35664
- Abstract:
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.
469. CMDO: A Cognitive Memory-Driven Optimization Algorithm for Adaptive Population-Based Search
- Authors: Mohammed Yusuf Mujawar , Shahram Rahimi , Noorbakhsh Amiri Golilarz
- URL: https://arxiv.org/abs/2609.35657
- Abstract:
Population-based optimization methods often use previous search information through successful solutions, parameter adaptation, or operator performance, but they rarely retain the context in which a search behavior succeeded or failed. We introduce Cognitive Memory-Driven Optimization (CMDO), a derivative-free population-based optimizer that represents experience as the relationship between search context, search behavior, and observed outcome. CMDO organizes these experiences across working, episodic, and consolidated memory, retrieves them according to similarity with the current search state, and uses both positive and negative evidence to guide subsequent search. Retrieved experience does not replay previous candidate locations; instead, it selects search recipes that are reconstructed from the current population through exploratory, directed, and local search behaviors with adaptive search geometry. We evaluate CMDO on selected Blackbox Optimization Benchmarking test suite on COCO (BBOB/COCO) and Congress on Evolutionary Computation 2017 (CEC2017) problems against DE, CMA-ES, SHADE, GWO, HHO, and ORCA, and further study its application to seven-parameter photovoltaic model estimation using measured current–voltage data. The results show problem-dependent but competitive optimization performance, including the lowest median error among the compared methods on CEC2017 F10. More importantly, analysis of the search traces shows that context-dependent recall changes the distribution of executed search behaviors, while unsuccessful experiences remain available as negative evidence for later decisions, showing that accumulated experience directly influences subsequent search behavior. These results support the use of explicit context–behavior–outcome memory as an active mechanism for controlling population-based search.
470. GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation
- Authors: Yuchen Sun , Jinjin He , Sinan Wang , Bo Zhu
- URL: https://arxiv.org/abs/2609.35639
- Abstract:
Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks cover fluids, deformable solids, and granular materials, from individual simulation operators to complete simulators. Agents write, compile, test, and optimize GPU code with access to a NVIDIA GPU under fixed time budgets. We report pass rates and runtime performance relative to expert-optimized reference implementations. In a single-attempt evaluation of six frontier model-harness pairs, the two strongest pass all 50 tasks, but even the fastest reaches at least 0.9 the reference speed on only 22% of them, and no submission is more than 5% faster than the reference. The largest gaps arise in collision detection, constraint solving, and iterative solvers. GPUPhysBench brings physical simulation workloads to coding-agent evaluation, testing both the ability to implement numerical methods correctly and the ability to make them run efficiently.
471. DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising
- Authors: Basile Morel , Samuel Ruiperez-Campillo , Andreas P. Streich , Julia E. Vogt , Thomas Hofmann
- URL: https://arxiv.org/abs/2609.35634
- Abstract:
Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically with sequence length, and diffusion-based approaches incur prohibitive inference cost. We propose a Mamba-augmented model that inserts selective state-space blocks at the convolutional bottleneck, combining local feature extraction with long-range temporal modeling at linear complexity. We comprehensively evaluate the proposed model with respect to reconstruction fidelity, noise robustness, recording-length scaling, and downstream diagnostic classification across over 40 pathology classes. On synthetic and real datasets, our model achieves the highest SNR and lowest RMSE, with the Mamba advantage increasing with sequence length and in low-SNR regimes. On classification with two independent classifiers, the proposed Mamba-based models achieve the best macro AUROC among all denoisers and improve over their convolutional base models. Calibration is more nuanced and classifier-dependent: denoising improves Binary Cross-Entropy and Brier score on Inception1D but often fails to beat the noisy input on ResNet1D-Wang, and the lead-specific Mamba variant is the only denoiser to improve both calibration metrics over the noisy baseline on both classifiers. Per-class analysis reveals a morphology-dependent benefit: Mamba substantially improves ST/T-change diagnoses, which depend on broad, context-sensitive waveforms.
472. Behavioral Foundation Models for Quality Diversity
- Authors: Nazim Bendib , Nicolas Perrin-Gilbert , Olivier Sigaud
- URL: https://arxiv.org/abs/2609.35615
- Abstract:
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all tested parameter-space QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.
473. Twist, Don’t Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models
- Authors: Aditya Thimmaiah , Lara Marinov , Jayanth Srinivasa , Haris Vikalo , Junyi Jessy Li , Milos Gligoric
- URL: https://arxiv.org/abs/2609.35609
- Abstract:
Constrained decoding for Masked Diffusion Language Models (MDLMs) aims to ensure that generated outputs satisfy a specified structure or syntax constraint. MDLMs generate outputs by repeatedly unmasking masked positions present in their current state. Recent strategies for constrained decoding constrain the model’s per-step mean-field posterior (which factorizes over masked positions) by enforcing the desired constraint with an automaton. The resulting chain-structured factor graph allows exact constrained sampling via dynamic programming. However, despite each draw being exact and constraint-satisfying, we prove that their composition, in general, tilts away from the model’s relative probabilities over valid trajectories, thus leading to trajectory bias. We derive an exact expression for this bias as a product of ratios measuring how valid continuation mass changes when the denoiser is reconditioned, and characterize when the bias vanishes. We then correct the bias by introducing TWISTER, the first automaton-twisted Sequential Monte Carlo decoder for MDLMs, using the step-exact decoder as the proposal. We show that for regular language constraints, the Feynman-Kac correction is exactly computable, with the twists obtained efficiently using quantities pre-computed for step-exact sampling. We prove that the resulting Feynman-Kac model targets the unbiased Doob h-transformed path law conditioned on constraint satisfaction.
474. SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
- Authors: Saswat Das , Parvati Viswanathan , Daniel Donnelly , Chang Huang , Sahar Abdelnabi , Ferdinando Fioretto
- URL: https://arxiv.org/abs/2609.35596
- Abstract:
Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operations, we provide an adaptive trajectory discovery pipeline that probes for failures while preserving original task intent and supports causal attribution through paired non-evolving agents and attribution scores. Our evaluation across multiple recent LLMs, evolution surfaces, and harm types reveals that self-evolution indeed increases task completion rates but often at the cost of safety failures that are absent for paired non-evolving baseline agents. We also show that qualitatively different safety behaviors emerge across evolution surfaces and harm types. Further, we show that this divergence in safety behavior is reflected in agents’ chain-of-thought reasoning, which yields an effective monitoring strategy that can mitigate unsafe behavior with a low false positive rate.
475. QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
- Authors: Pranav Gupta
- URL: https://arxiv.org/abs/2609.35581
- Abstract:
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.
476. FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
- Authors: Bowen Yang , Jingbo Zhou , Qinghong Miao , Hua Wu
- URL: https://arxiv.org/abs/2609.35578
- Abstract:
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
477. F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
- Authors: Zhuoyuan Yu , Jiacheng Wang , Tianle Liu , Yihua Ren , Peng Yu , Chen Bai , Ziheng Zhang , Yufei Jia , Jindou Jia , Yuhang Zhang , Xinrui Zhang , Shang Yujing , Yuxiang Chen , Chuhao Zhou , Tiancai Wang , Jianfei Yang
- URL: https://arxiv.org/abs/2609.35575
- Abstract:
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
478. Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
- Authors: Omid Ghahroodi , Anas Madkoor , Dima Faris Al Saudi , Fagr Tahir , Malak Annan , Talha shahid javad allah rakha , Omar Al-Busaidi , Zineb El Kahla , Iheb Zouari , Essa Ahmed Abou Jabal , Ahmed Ezzat , Hind AL-Merekhi , Aisha Hamad M A Al-Naimi , Hadi Wazni , Bushra Alnajjar , Omar Amin , Haya Al-Thani , Houssam Eddine-Othman Lachemat , Marwa Elwakedy , Sundus Abdulmalik Al Nahari , Elahe Zahiri , Osamah Sarraj , Raghad Mousa , Mckeen Assi , Ahd Al Jumah , Heyam Salman , Alhanouf Abdulraqib , Sara Benoumhani , Alia Hamwi , Ayaat Al-Yasseri , Rim Ibrahim Ghazal , Lamia Ben hiba , Mohamed Eltabakh , Fatima Al-Raisi , Yassine El Kheir , Mohammed Abdulrahman , Hamdy Mubarak , Ayah Hashem , Lefkir Meriem , Ehsaneddin Asgari
- URL: https://arxiv.org/abs/2609.35564
- Abstract:
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.
479. The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent
- Authors: Shobhan Roy (University of Iowa)
- URL: https://arxiv.org/abs/2609.35557
- Abstract:
The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.
480. Graph World Models for Constrained Epidemic Policy Planning
- Authors: Yiqi Su , Rashed Shelim , Lingyi Wang , Walid Saad , Naren Ramakrishnan
- URL: https://arxiv.org/abs/2609.35545
- Abstract:
Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We present EpiMind, a graph world model framework for constrained epidemic policy planning across regions. A graph-factored recurrent state-space model generates joint policy-conditioned rollouts from regional latent beliefs, while graph-temporal ADMM optimizes regional interventions, enforces shared-resource feasibility through projection, and evaluates temporal specifications under the learned model. EpiMind reduces admission RMSE by 29% relative to graph-free dynamics modeling, plans within 1-5% of the best feasible constant policy with guaranteed shared-budget feasibility, and outperforms all deployable baselines across three resource budgets in real-context evaluation. These results demonstrate that graph-structured policy imagination with explicit constrained coordination supports effective epidemic interventions from learned dynamics.
481. Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability
- Authors: Xu Wang , Difan Zou , Xuansheng Wu
- URL: https://arxiv.org/abs/2609.35544
- Abstract:
Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target concept by supplying its associated activation during learning. Across three Qwen3.5 base models, we use sparse autoencoders (SAEs) to identify the top-ranked sycophancy feature from paired sycophantic and independent responses, then validate its behavioral influence through inference steering. We subsequently inject the selected feature during supervised fine-tuning on sycophantic targets. Positive injection reduces learned sycophancy after removal (by 62.0% relative to ordinary fine-tuning in 35B-A3B), whereas modest negative injection increases it. Unexpectedly, these reductions in sycophancy do not consistently improve direct refusal of harmful requests, motivating a narrower evaluation of the same harmful intents under user pressure. In this setting, ordinary fine-tuning on sycophantic responses substantially weakens refusal, while selected checkpoints trained with positive injection recover part of the loss, including approximately 95% in 35B-A3B. These findings show that persistent sycophancy reduction does not guarantee stronger direct refusal, while identifying recovery under user pressure as a distinct, conditional benefit of training intervention.
482. AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
- Authors: Yuta Oshima , Ku Onoda , Yusuke Iwasawa , Masahiro Suzuki , Yutaka Matsuo , Hiroki Furuta
- URL: https://arxiv.org/abs/2609.35530
- Abstract:
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.
483. Let the Neurons Die: Exploiting ReLU-Induced Model Degradation
- Authors: Kexin Li , Wenjun Qiu , Joshua Abraham , Aditi Maheshwari , David Lie
- URL: https://arxiv.org/abs/2609.35528
- Abstract:
Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example that minimizes the target layer’s post-update weight sum, aiming to push ReLU units toward negative pre-activations without modifying training samples or labels. We then develop two poisoning attacks, IG-DOA and IG-SKA, which use gradient inversion to synthesize class-conditioned samples by matching reference gradients in adverse model states constructed through data ordering or soft knockout, respectively. Soft knockout rearranges weights across adjacent layers to concentrate negative contributions. On a fully connected ReLU network trained on MNIST, ordering 100 of 60,000 training examples reduces test accuracy from 96% to 95% after only five epochs. Adding 200 poisoned samples from a single class reduces test accuracy to approximately 86-88% after five epochs in most evaluated conditions, compared with approximately 96% under clean training. These results demonstrate that ReLU-targeted data ordering and poisoning can impair learning without directly modifying the victim model’s parameters.
484. Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery
- Authors: Xu Wang , Yifan Yang , TingHao YU , Difan Zou
- URL: https://arxiv.org/abs/2609.35521
- Abstract:
Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of chunk-level SAEs that encode mean-pooled activations over chunks, each a contiguous span of tokens: Mean-Chunk reconstructs the observed chunk, Cross-Chunk predicts an independently processed neighbor, and Joint-Chunk combines both targets. These designs separate the effect of a larger observation unit from that of predicting information shared across passages. With matched training data, chunk-level SAEs remain powerful interpretability tools while learning reliable semantic features that capture high-level concepts and respond selectively to relevant content. Their strengths are complementary: Mean-Chunk improves high-level feature discovery, reasoning detection beyond surface cues, and steering; Cross-Chunk leads document retrieval and classification transfer while producing selective, persistent features. Changing what an SAE sees and predicts yields reliable semantic features for more meaningful tasks. We demonstrate their practical value through gains across downstream tasks such as retrieval, reasoning detection, and steering.
485. Improving Generative Model Self-Training with Geometrically Modified Outputs
- Authors: Patrick Batsell , Thomas Walker , Richard Baraniuk
- URL: https://arxiv.org/abs/2609.35512
- Abstract:
Self-training generative models - the continued improvement of a model using its own outputs - is becoming increasingly important as high-quality training data becomes scarce. However, naively finetuning on model-generated samples leads to degradation through model collapse and the model autophagy disorder. Negative-guidance self-training methods turn this degradation into a useful signal, using a model finetuned on its own outputs to guide the original model toward improved generation. Existing methods, however, take the negative signal in standard model outputs as given. We instead ask whether this signal can be explicitly strengthened. We introduce Geometrically Modified Outputs (GMOs), which reweight the singular values of the generator’s input-output Jacobian to increase the influence of its leading singular directions. This geometric modification amplifies the mode-seeking behavior and distortions of standard outputs, providing a stronger and more targeted negative signal for self-training. Across a range of one-step generative models, GMOs consistently improve the performance of negative-guidance methods, including Neon and SIMS, compared with using standard model outputs.
486. SolveEdit: Benchmarking Visual Problem Solving in Generative Models
- Authors: Wenjie Shu , Yexin Liu , Harold Haodong Chen , Xuerui Qiu , Zehan Wang , Yidi Zhang , Yizhan Chen , Zunwei Wang , Minghao Liu , Qi Chen , Harry Yang , Xiaogang Xu
- URL: https://arxiv.org/abs/2609.35504
- Abstract:
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
487. From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
- Authors: Chi Zhang , Yueyi Liu , Haoyang Shi , Ruichuan An , Haoyu Li , Yuhang Wu , Sen Cui , Miao Liu
- URL: https://arxiv.org/abs/2609.35491
- Abstract:
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström–Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
488. Analog Computing revisited: A fully analog and minimalistic Damage Detector for Ultrasonic Testing enabling Material-Integrated Structural Health Monitoring
- Authors: Stefan Bosse
- URL: https://arxiv.org/abs/2609.35478
- Abstract:
Ultrasonic Testing (UT) is commonly used to detect damage in structures, e.g., metal plates. A sensor acquires Ultrasonic waves, e.g., by using PZT transducers. The time-resolved sensor signal must be processed with analog electronics, e.g., amplified and filtered. Commonly a digitalization follows using an Analog-to-Digital converter, finally processing the digital sensor signal, applying digital signal processing, feature extraction, and Machine Learning by using powerful microprocessor systems. The disadvantages of digital processing systems are their high number of transistors (microchip area), energy consumption, state-dependent processing and therefore sensitivity to energy supply interruption. Beyond silicon electronics, printed organic electronics gains interest. But printed electronics is still limited to low transistor and electronic component counts (typically 100). We will investigate and demonstrate a fully analog signal processing and feature extraction system consisting of an analog Hilbert transform deriving the signal envelope, simple analog arithmetic calculations for feature extraction, and finally damage classification and regression using an analog Artificial Neural Network. We expect a full damage detection system with less than 100 transistors. We will test our damage detection system with PZT transducer signals from Steel plates with circular defects. The focus of this work is the analog computation of the signal envelope (using all-pass filter networks for approximation of the Hilbert transform) and the analog feature extraction as well as the prediction of damage, forming an analog computer which can perform in-sensor computation, computing without a digital computer.
489. Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs
- Authors: Pranjal Garg , Jacob Beck
- URL: https://arxiv.org/abs/2609.35475
- Abstract:
Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges spontaneously despite training exclusively on clean sequences, without corruption training or an explicit denoising objective. We find that context restoration follows a two-phase process: early layers localize effects associated with repair at corrupted positions, while later layers accumulate these effects at uncorrupted positions through the residual stream and ultimately concentrate them at the output position. Repair outcome is predictable from hidden states: cosine alignment with the clean state is highly predictive in attention-only models, while linear probes recover additional information in pretrained LLMs. A linear probe using only the corrupted prompt’s first-block hidden state predicts failure with mean ROC-AUC 0.78. This enables failure triage under matched or even partially shifted deployment conditions and may reduce unnecessary verification or computation. Failed examples also show substantially greater nonlinearity along corruption directions. Moderate-corruption finetuning increases corruption tolerance while simultaneously reducing displacement-normalized linearization error, associating improved robustness with a more nearly linear response to corruption.
490. Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching
- Authors: Chenyu Zhang , Yuhang Cao , Daru Du , Yingxi Lu , Jing Shao , Ruoqu Chen , Jiajun Liu , Liu Cao , Yicheng Liu , Hang Zhao , Mengdi Xu
- URL: https://arxiv.org/abs/2609.35469
- Abstract:
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
491. CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
- Authors: Yusuf Kesmen , Aniruddha Mukherjee , Yena Chang , David Sasu , Trevor Brokowski , Alexandra V. Kulinkina , Kristina Keitel , Akhil Arora , Lars Henning Klein , Mary-Anne Hartley
- URL: https://arxiv.org/abs/2609.35462
- Abstract:
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at this https URL .
492. Riccati State Space Models: Non-iterative Parallelization for Nonlinear Sequence Modeling
- Authors: Mónika Farsang , Ramin Hasani , Daniela Rus , Radu Grosu
- URL: https://arxiv.org/abs/2609.35441
- Abstract:
State space models (SSMs) achieve efficient sequence processing because their affine state updates are closed under composition and can therefore be evaluated with an associative parallel scan. Nonlinear recurrent models can provide richer, state-dependent dynamics, but generally lose this compositional structure: parallel evaluation then requires iterative methods that repeatedly linearize and scan the recurrence. We ask, what state-dependent nonlinear dynamics can be designed to remain exactly composable? We answer by introducing RiccatiSSM, a nonlinear SSM, in which each state dimension follows an input-conditioned Riccati differential equation. Its quadratic state dependence makes the local Jacobian explicitly state-dependent, while its exact per-step flow under piecewise-constant inputs is a Möbius transformation. Since Möbius maps are closed under composition and compose through $2\times 2$ matrix multiplication, the complete nonlinear state trajectory can be evaluated exactly with a single associative parallel scan, without iterative linearization. We further derive a constrained parameterization that ensures bounded, contractive dynamics, and avoids poles in the fractional-linear state update. Across long-sequence classification, regression, and forecasting tasks, RiccatiSSM achieves competitive predictive performance while reducing runtime by $22{-}33\%$ compared to the nonlinear LrcSSM under matched architectures. These results demonstrate that state-dependent nonlinear dynamics can retain exact composability and be evaluated efficiently within a single parallel scan.
493. ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
- Authors: Yihang Chen , Yuanhao Ban , Cho-Jui Hsieh
- URL: https://arxiv.org/abs/2609.35433
- Abstract:
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $\alpha$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.
494. Frontier Learning: Training LLM Reasoners at the Edge of Capability
- Authors: Robin Faro , Shyam Sundhar Ramesh , Ilija Bogunovic , Aurelien Lucchi
- URL: https://arxiv.org/abs/2609.35426
- Abstract:
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator’s task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model’s evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
495. Semantic Prefix Oracles for LLM Decoding: Contracts and Differential Validation
- Authors: Paul Kronlund-Drouault
- URL: https://arxiv.org/abs/2609.35425
- Abstract:
Constrained decoding can enforce regular or context-free output formats, but many program-generation failures are semantic: scope, typing, and declaration effects depend on context. We present semantic grammar specifications, a declarative formalism that attaches such constraints to a context-free surface and executes them during Earley descent. Our implementation enforces \emph{safe pruning}: it rejects only prefixes whose semantic contradictions cannot be repaired by any continuation. A separate, grammar-dependent, \emph{dead-end freedom} property guarantees the existence of a realizable witness for each remaining branch. We give simple sufficient conditions based on surface productivity, type coverage, and left-to-right constraint flow. Our finite-lambda, core ML, and C-like fragments satisfy them, while the STLC instance used in our experiments does not: plain STLC can violate type coverage, and we show how restricting its type universe recovers it. A tokenizer-lifting lemma carries character-level witnesses to token sequences under an explicit vocabulary-coverage hypothesis. We validate the implementation differentially against production compilers (\texttt{ocamlc}, \texttt{cc}). Across every prefix of 65 compiler-valid programs we observe zero false prunes. The semantic oracle localizes 25/30 invalid programs mid-stream, against 0/30 for a syntax-only oracle, and agrees on 42/42 recursion probes. A twelve-model generation study, including a matched semantic-versus-syntactic ablation for nine models, finds nonnegative observed semantic-minus-syntactic point estimates for every model-language pair, with maxima of $+15.2$ points on STLC task correctness and $+14.3$ points on ML validity.
496. GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection
- Authors: Zelin Zhao , Guanjie Huang , Danny Hin Kwok Tsang , Li Liu
- URL: https://arxiv.org/abs/2609.35411
- Abstract:
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain this http URL code will be released upon publication.
497. Spectral Super-Resolution using Spatial-Spectral Residual Operator Networks
- Authors: Seokhyun Chin
- URL: https://arxiv.org/abs/2609.35410
- Abstract:
Spectral super-resolution of multispectral satellite images can enable high temporal- and spatial-resolution hyperspectral satellite imagery at a modest cost, significantly increasing the applicability of hyperspectral remote sensing. This task is inherently ill-posed, making it well-suited for deep learning-based methods. In this study, the spectral super-resolution task is framed as an operator learning problem, and SSRON is proposed as a Deep Operator Network that effectively learns function-to-function mappings from downsampled spectra to continuous spectra. The model is trained to super-resolve Sentinel-2A-like multispectral imagery to EMIT images. Compared to baseline models, SSRON achieves superior performance across all metrics. The model also demonstrates zero-shot spectral super-resolution capability by predicting bands unseen during training. Furthermore, its continuous-output formulation suggests the potential to estimate spectra at finer wavelength intervals than the native sensor. These results suggest the potential of SSRON and establishes operator learning as a promising direction for spectral super-resolution.
498. AwarenessBench: Assessing Cognitive Capabilities of Language Models
- Authors: Xiaojian Li , Rongwu Xu , Tianyun Zhang , Yue Wang , Shuo Chen , Qiner Lyu , Briana Zhang , Peiran Yang , Kyle Xue Chen , Haoyuan Shi , Yu Wang , Wei Xu
- URL: https://arxiv.org/abs/2609.35409
- Abstract:
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.
499. “Nothing to See Here’’: Unintended Disclosure through Revision Traces of LLM Deliverables
- Authors: Yage Zhang , Yukun Jiang , Yang Zhang
- URL: https://arxiv.org/abs/2609.35408
- Abstract:
Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such statements revision traces. For example, after a user removes the password before sharing a configuration file, the model may delete it but leave a comment saying, “Removed the password ‘No**4!’ as requested.” A third-party recipient who sees only the delivered file can therefore recover the withdrawn password from the comment. In an in-the-wild analysis of three public conversation corpora, we identify 26,753 revision requests, of which 2,363 (8.8%) leave revision traces. We study them in greater depth under controlled conditions by introducing RevLeakBench, a benchmark of 100 tasks across five scenarios with a conversation track and an agent track. We measure trace occurrence, withdrawn-item recovery, trace position, and required-content retention. Across six models, about half of the deliverables in both tracks state the edit after a revocation, and a reader that sees only the deliverable can recover the withdrawn item from about 13% of them. Telling the model that its entire reply will be forwarded to the recipient still leaves revision traces in 36.4% of the deliverables. We compare prompt defenses and a delivery boundary, and propose an output-side filter that sharply reduces recovery with little loss of required content. We believe our work can benefit efforts to understand and mitigate unintended disclosure in LLM interactions.
500. The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
- Authors: Yihe Zhou , Tongtian Zhu , Yingxiao Huo , Satya Prakash Dash , Can Wang , Samuel Kaski , Mingfei Sun
- URL: https://arxiv.org/abs/2609.35392
- Abstract:
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$\beta$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam’s connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$\beta$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.