Last 7 Days (October 03 – October 09, 2026)
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Primary: ETH Zurich (Inferred from "eth-lre" in GitHub URL and Swiss AI Initiative funding)
All Institutions: ETH Zurich, Swiss National Supercomputing Centre (CSCS)
The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
The paper addresses a critical flaw in existing RL-based tutoring methods: the "telling" problem, where models learn to simply provide answers to maximize immediate student success metrics. The authors propose a novel reward structure based on three conditions: (1) Near-transfer testing (student solves a variant problem, not the exact one), (2) Masked post-test (tutor's utterances are hidden during the test, forcing reliance on student-generated notes/reasoning), and (3) Binary reward gates (strict penalties for factual errors or solution handover). This design effectively decouples "helpfulness" from "teaching," forcing the model to elicit student reasoning rather than provide it. The use of a frozen LLM student with genuine errors (rather than prompted confusion) is a strong methodological choice that ensures the training signal reflects real diagnostic and adaptive challenges.
The experimental setup is rigorous and comprehensive. The authors train models across multiple sizes (4B to 27B) and architectures (Qwen3 variants), demonstrating the robustness of the recipe. The ablation study is particularly strong, isolating the contribution of each condition (transfer, masking, gates) and showing that the gates specifically reduce handover while the masked transfer test improves out-of-domain generalization. The evaluation uses independent benchmarks (MathTutorBench, TutorMoments) and judges (Gemini) distinct from the training setup, mitigating overfitting concerns. The results show that Eduardo-27B matches or exceeds frontier models (Gemini 3.1 Pro, Claude Opus 4.8) in pedagogical quality while using significantly fewer thinking tokens, a crucial efficiency gain for interactive applications.
High. The authors open-source the training environment, the 8,671-problem near-transfer dataset, and the trained models. Detailed hyperparameters, prompts, and compute costs are provided in the appendices. The use of standard RL algorithms (DPPO/GRPO) and open-source base models further enhances reproducibility.
The primary limitation is the reliance on a single frozen LLM student (Llama-3.1-8B-Instruct) for training, which may lead to overfitting to that specific model's failure modes. The paper acknowledges this and suggests future work with diverse student ensembles. Additionally, the evaluation is limited to mathematics, and the "affective gap" (lack of socio-emotional support) is noted as a consequence of optimizing purely for cognitive gain. The ablation study is limited to one seed and one model size (4B), which slightly weakens the statistical confidence in the ablation results.
This work has significant implications for the development of AI tutors and, more broadly, for any domain where an agent must build user capabilities rather than just provide answers. The "masked near-transfer" reward design is a generalizable principle that could be applied to other educational or collaborative tasks. The efficiency gains (fewer thinking tokens) make high-quality tutoring more feasible in real-time interactive settings. The open-sourcing of the dataset and models will likely accelerate research in AI for Education. The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Primary: University at Buffalo (inferred from Empire AI Consortium and NSF grants)
All Institutions: University at Buffalo, Empire AI Consortium, Inc, Simons Foundation, Secunda Family Foundation, Modal Labs
LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
The paper proposes LayerRoPE, a method that reinterprets the growth of hidden state norms in deep Transformers not as a pathology to be suppressed, but as an emergent depth-positional encoding. The core insight is that the RMSNorm weights ($\gamma$) across layers exhibit systematic magnitude growth and directional rotation, effectively encoding the layer index. LayerRoPE makes this explicit by replacing the $L$ independent per-layer $\gamma$ vectors with a single shared vector modulated by depth-conditioned scalars (magnitude and rotation) derived from a RoPE-like mechanism along the depth axis. This reduces parameter count and introduces a structured, learnable depth schedule. The methodology is sound, leveraging the analogy between positional encoding in sequence space and depth encoding in network space. The ablation studies confirm that both magnitude and rotation components are necessary and complementary.
The experimental evaluation is extensive and rigorous. The authors train a model ladder from 58M to 1.3B parameters on C4, comparing LayerRoPE against Pre-Norm, Post-Norm, Peri-Norm, and Layer-Norm Scaling. LayerRoPE consistently outperforms baselines, reaching Pre-Norm's 1.3B loss with 3.4x less compute. Crucially, the paper demonstrates superior depth scaling, with LayerRoPE being the only method to show monotonic loss improvement up to 512 layers, whereas baselines diverge or plateau. The paper also shows improved learning rate robustness (wider basin, delayed divergence) and successful transfer to looped latent models (Parcae) and Vision Transformers (ViT) without tuning. The analysis of the learned depth schedule reveals that LayerRoPE widens the residual stream rather than damping it, which is a counter-intuitive and significant finding.
The paper provides detailed hyperparameters, training recipes, and compute accounting. It specifies the initialization of depth slopes, rotation base frequency, and optimizer settings. The use of standard datasets (C4, ImageNet) and architectures (LLaMA-style, ViT) enhances reproducibility. However, the specific code implementation is not linked in the text provided, which is a minor gap for immediate reproduction, though the description is sufficiently detailed for re-implementation.
The primary limitation is the scale of the experiments. While 1.3B parameters and 512 layers are significant, the method's behavior at frontier scales (100B+) is extrapolated or tested at reduced budgets (6.7B). The paper relies on C4 for language modeling, which may not fully capture the benefits on more diverse or multilingual data. Additionally, the "widening" of the residual stream requires careful monitoring to ensure it doesn't lead to numerical instability in mixed-precision training, although the paper notes stability improvements.
This paper has high potential impact by challenging a fundamental assumption in Transformer architecture design: that activation norm growth is a defect. By providing a principled way to exploit this growth for depth encoding, it offers a path to training deeper, more stable, and more compute-efficient models. The transferability to looped models and ViTs suggests broad applicability across modalities. The reduction in learning rate sensitivity is particularly valuable for practitioners, as it simplifies hyperparameter tuning for deep networks. LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology
LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
The paper proposes LASER, a framework for offline RL that combines support-constrained latent space policies with explicit entropy regularization. The core methodological contribution is the use of "adjoint matching" to optimize the entropy-regularized policy objective in latent space without backpropagation through time (BPTT). This is a significant technical advancement over previous flow-based offline RL methods (like ReFORM or DSRL) which either relied on implicit regularization or computationally expensive BPTT. The derivation of the optimal Q-tilted distribution and its implementation via stochastic optimal control (SOC) is mathematically sound and provides a clear path for stable training. The separation of support constraints (handled by the flow decoder) from policy extraction stability (handled by entropy regularization) is a well-motivated architectural choice.
The experiments are conducted on the OGBench benchmark, covering 40 tasks across locomotion and manipulation. The results show that LASER achieves state-of-the-art performance, particularly on noisy datasets where support constraints are critical. A key strength is the demonstration of robustness: LASER uses a fixed set of hyperparameters across all tasks, outperforming baselines that require task-specific tuning. The ablation studies effectively isolate the contribution of entropy regularization and the efficiency of adjoint matching vs. BPTT (showing a 3.6x speedup). The comparison against recent flow-based baselines (IFQL, FQL, QAM) is appropriate and rigorous.
The paper provides a clear algorithmic summary and references a project website. The use of standard benchmarks (OGBench) and fixed hyperparameters enhances reproducibility. However, the specific implementation details of the adjoint matching solver and the flow matching training are deferred to the appendix, which is standard but requires careful reading. The code availability is implied by the project URL.
The method relies on the quality of the behavior cloning (BC) decoder to enforce support constraints; if the decoder is imperfect, OOD actions can still occur. The computational cost of solving ODEs during training remains high, though mitigated by adjoint matching. The paper acknowledges that the value function learning is relatively simple compared to more advanced pessimistic methods, suggesting room for improvement in the critic component.
This work advances the state of the art in safe offline RL, which is critical for deploying RL agents in real-world robotic systems where online interaction is costly or dangerous. The technique of using adjoint matching for entropy-regularized flow policies could be adapted to other generative model-based RL settings, potentially influencing how expressive policies are trained in broader RL research. LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets. Zatom-2 features a multiscale Transformer architecture coupled with conditional flow matching that supports force conditioning and foundational pretraining tasks such as generation, structure prediction, and prediction of molecular and material energies and forces. Empirically, Zatom-2 achieves better molecular distribution fidelity than Zatom-1 and achieves strong performance on existing molecule and material generation benchmarks. Zatom-2 demonstrates the ability to control sample generation across low- and high-force regimes, and enhances protein generation in a low-data setting through joint generative-predictive pretraining and transfer learning, increasing protein backbone designability in a length extrapolation setting from 67.8% without pretraining to 74.8% after finetuning on 2,000 protein domains.
Primary: University of California, Berkeley
All Institutions: University of California, Berkeley, Lawrence Berkeley National Laboratory, University of Cambridge, University of Oxford, University of Washington, University of Oxford (Inferred from authors: Mahoney, Erichson, Jacobson, Blau, Perez, Anand, Abrudan, Panescu, Li, Knowles, Morehead, Cretu)
Zatom-2 introduces a unified atomistic generative model that leverages multitask pretraining on large-scale molecular and material datasets to enable cross-domain transfer to protein generation. By combining a novel atom1 tokenization scheme with conditional flow matching and physics-grounded auxiliary objectives (energy/force prediction), the model demonstrates superior distribution fidelity and, crucially, the ability to transfer geometric representations from small molecules to larger biomolecules, achieving significant gains in protein designability in low-data and out-of-distribution settings.
The paper introduces Zatom-2, a unified atomistic generative model that bridges the gap between small molecule/material generation and protein structure prediction. The core methodological contribution is the "atom1" tokenization scheme for molecules/materials (analogous to "atom14" for proteins), which allows a single Transformer backbone to handle variable atom counts without separate diffusion tracks for atom identity. The model employs conditional flow matching for coordinate generation and incorporates a novel "force conditioning" mechanism that uses the mean force norm as a global descriptor to control the relaxation state of generated structures. A key innovation is the multitask pretraining curriculum that jointly optimizes for generation, structure prediction, and energy/force prediction (MLIP objectives). The authors hypothesize and demonstrate that the physics-grounded supervision from energy/force prediction shapes the latent representations in a way that improves generative fidelity and transferability, particularly in low-data regimes.
The experimental evaluation is extensive, covering pretraining on ~5 million structures from OMol25 and OMat24. The paper demonstrates that Zatom-2 outperforms its predecessor (Zatom-1) in distribution fidelity and achieves competitive results on standard benchmarks like QM9, MP20, and GEOM-Drugs. A significant portion of the evaluation focuses on the "force consistency" metric, showing the model can reliably distinguish between high-force (rattled) and low-force (relaxed) regimes. The most compelling results are in the transfer learning section, where Zatom-2 pretrained on small molecules/materials is fine-tuned on a small protein dataset (SCOPe-2k). The model shows significant improvements in designability and novelty, especially in out-of-distribution (longer sequence) settings, compared to training from scratch or using protein-specific models like RFdiffusion3. The scaling study confirms that both data coverage and model capacity yield complementary gains.
The paper provides a strong reproducibility statement, including a link to the open-source code repository (https://github.com/Zatom-AI/nucleus) which contains documentation, data loading, and training/inference materials. The appendix details the model architecture, hyperparameters, and compute resources used (NVIDIA A100 GPUs). The authors also provide an AI Use Statement, disclosing the use of LLMs for code and writing assistance, which adds to the transparency of the research process.
The model does not enforce exact rotational equivariance, relying instead on data augmentation and pair-biased attention, which may limit its performance in tasks requiring strict symmetry preservation compared to equivariant diffusion models. The MLIP (energy/force) prediction accuracy is lower than specialized potentials, though it serves as an auxiliary task. The protein generation experiments are limited to a small dataset (2k domains) and do not yet match the performance of state-of-the-art protein-specific models in in-distribution settings, though it shows promise in extrapolation. The paper acknowledges that precise force matching within regimes remains challenging.
This work has significant potential to accelerate discovery in chemistry, materials science, and biology by providing a unified framework for atomistic modeling. The ability to transfer knowledge from data-rich chemical domains to data-scarce biological contexts could reduce the computational cost of training specialized models. The force conditioning mechanism offers a new tool for controlling the physical state of generated structures, which is valuable for simulating non-equilibrium processes. The open-sourcing of the code and the use of large public datasets (OMol25, OMat24) will facilitate further research in this area. Zatom-2 introduces a unified atomistic generative model that leverages multitask pretraining on large-scale molecular and material datasets to enable cross-domain transfer to protein generation. By combining a novel atom1 tokenization scheme with conditional flow matching and physics-grounded auxiliary objectives (energy/force prediction), the model demonstrates superior distribution fidelity and, crucially, the ability to transfer geometric representations from small molecules to larger biomolecules, achieving significant gains in protein designability in low-data and out-of-distribution settings.
Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.
Primary: University of Helsinki
All Institutions: University of Helsinki, Aalto University, University of Oxford
PrivTab introduces a novel architectural integration of differential privacy into tabular foundation models, enabling provably private classification with negligible utility loss and orders-of-magnitude speedup over traditional private training methods. By embedding a differentially private cross-attention mechanism that compresses sensitive data into a fixed-size private summary, the model leverages the post-processing property of DP to ensure all downstream predictions are private, while noise-aware pretraining allows the model to learn how to predict effectively through the privacy noise. This approach resolves the tension between the high utility of foundation models and the rigorous guarantees of differential privacy, offering a practical, auditable, and efficient solution for sensitive tabular data applications.
The paper introduces PrivTab, a tabular foundation model that integrates differential privacy (DP) directly into the architecture via a "Differentially Private Multi-Head Cross-Attention" (DP-MHCA) layer. This is a significant architectural shift from traditional DP-SGD, which applies noise during iterative optimization. By compressing the context dataset into a fixed-size, noisy summary in a single forward pass, PrivTab leverages the post-processing property of DP to ensure that all subsequent predictions are private without additional privacy cost. The methodology is rigorous, featuring a formal proof of the privacy guarantee (verified in Lean 4), a noise-aware pretraining objective that teaches the model to predict through noisy summaries, and a clear separation between the privacy-critical path (the summary module) and the prediction path. The use of Gaussian DP (GDP) and adaptive composition is handled correctly, and the sensitivity analysis is sound.
The evaluation is comprehensive, utilizing the TabArena benchmark with 33 datasets across six privacy regimes. PrivTab demonstrates superior AUC and Log Loss compared to DP-Linear Regression and DP-MLP baselines, particularly in moderate-to-strong privacy settings where traditional methods suffer significant utility loss. The computational efficiency is a standout result, with fitting times reduced by ~10,000x (milliseconds vs. minutes). The membership inference attack analysis convincingly demonstrates that non-private tabular foundation models (TabPFN, TabICL) are highly vulnerable, whereas PrivTab maintains strong privacy. The inclusion of a private preprocessing case study adds practical relevance, showing how to handle feature scaling within the privacy budget.
High. The authors provide a public GitHub repository containing the code, Lean 4 formal verification scripts, and empirical auditing tools. The pretraining details, including the simulator used (TabICL simulator) and hyperparameters, are described in detail. The use of standard benchmarks (TabArena) and clear evaluation metrics (AUC, Log Loss, Elo) facilitates comparison with other works.
The model is currently limited to classification tasks, not regression. The privacy model assumes one user per row, which may not hold in all real-world scenarios (e.g., longitudinal data). The advantage diminishes on very large datasets or under very weak privacy constraints, where traditional DP-SGD methods can leverage more data. The reliance on public descriptions for clipping bounds in preprocessing is a practical constraint that may not always be met.
This work has the potential to significantly lower the barrier to entry for differentially private machine learning. By removing the need for dataset-specific private training and hyperparameter tuning, it makes DP accessible to domain experts who are not privacy specialists. The architectural insight—that privacy can be embedded into the foundation model's in-context learning mechanism—could inspire similar approaches in other modalities (text, image) where privacy is a concern. The formal verification aspect sets a new standard for trust in private ML implementations. PrivTab introduces a novel architectural integration of differential privacy into tabular foundation models, enabling provably private classification with negligible utility loss and orders-of-magnitude speedup over traditional private training methods. By embedding a differentially private cross-attention mechanism that compresses sensitive data into a fixed-size private summary, the model leverages the post-processing property of DP to ensure all downstream predictions are private, while noise-aware pretraining allows the model to learn how to predict effectively through the privacy noise. This approach resolves the tension between the high utility of foundation models and the rigorous guarantees of differential privacy, offering a practical, auditable, and efficient solution for sensitive tabular data applications.
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
Primary: Columbia University
All Institutions: Columbia University
The paper introduces a decoupled exploration-distillation framework for RLVR that effectively boosts reasoning performance and diversity without degrading general capabilities. By separating the aggressive exploration policy from the stable optimization policy, it resolves the tension between novelty and correctness, offering a robust and scalable improvement over standard RLVR methods.
The paper proposes Exploration-Distillation (ExpDis), a framework that decouples exploration from optimization in Reinforcement Learning with Verifiable Rewards (RLVR). The core insight is that applying novelty bonuses directly to the policy being optimized leads to degradation of general capabilities because the verifiable reward only supervises a narrow slice of the model's knowledge. By training separate "explorer" policies with aggressive novelty incentives (using Random Network Distillation) and distilling only their correct, high-quality trajectories into a "student" policy via SFT followed by standard RLVR, the method isolates the benefits of exploration from the risks of instability. The methodology is sound, building on established concepts like intrinsic motivation and expert iteration, but the specific application to LLM reasoning with the decoupled architecture is a novel and effective contribution.
The experiments are rigorous and well-controlled. The authors compare against strong baselines (DAPO, GRPO, Dr. GRPO) and a direct novelty-bonus variant of DAPO. They demonstrate that ExpDis outperforms DAPO at the same wall-clock budget across multiple model families (Qwen3, Ministral) and benchmarks (AIME, MATH, etc.). Crucially, they show that while direct novelty bonuses degrade general knowledge benchmarks (MMLU, GPQA), ExpDis preserves or improves these capabilities while boosting reasoning performance. The analysis of pass@$k$ scaling and diversity metrics provides strong evidence that the method genuinely expands the model's solution space rather than just sharpening existing modes.
The paper provides high reproducibility. Code and checkpoints are released on GitHub and HuggingFace. Hyperparameters, including the novelty bonus weight schedule and filtering criteria, are detailed in the appendix. The use of standard datasets (DAPO-Math-17K) and public benchmarks further aids reproducibility.
The method requires training multiple models (explorers + student), which increases memory and compute overhead compared to single-model RLVR, although the paper argues this is offset by the ability to use aggressive exploration. The current evaluation is limited to mathematical reasoning; it is unclear if the benefits generalize to other domains like code or open-ended generation without similar verifiable rewards. The optimal number of explorers and rounds may be sensitive to the specific task and model scale.
This work has significant implications for the scaling of LLM reasoning. As models become proficient at known strategies, the need for discovering new strategies grows. ExpDis provides a safe mechanism to do so, potentially enabling the discovery of novel reasoning paths that standard RLVR would miss due to entropy collapse. It offers a practical path for practitioners to enhance model diversity and capability without the risk of catastrophic forgetting or quality degradation. The paper introduces a decoupled exploration-distillation framework for RLVR that effectively boosts reasoning performance and diversity without degrading general capabilities. By separating the aggressive exploration policy from the stable optimization policy, it resolves the tension between novelty and correctness, offering a robust and scalable improvement over standard RLVR methods.
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Primary: University at Buffalo (inferred from Empire AI Consortium and NSF grants)
All Institutions: University at Buffalo, Empire AI Consortium, Inc, Simons Foundation, Secunda Family Foundation, Modal Labs
LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
The paper proposes LayerRoPE, a method that reinterprets the growth of hidden state norms in deep Transformers not as a pathology to be suppressed, but as an emergent depth-positional encoding. The core insight is that the RMSNorm weights ($\gamma$) across layers exhibit systematic magnitude growth and directional rotation, effectively encoding the layer index. LayerRoPE makes this explicit by replacing the $L$ independent per-layer $\gamma$ vectors with a single shared vector modulated by depth-conditioned scalars (magnitude and rotation) derived from a RoPE-like mechanism along the depth axis. This reduces parameter count and introduces a structured, learnable depth schedule. The methodology is sound, leveraging the analogy between positional encoding in sequence space and depth encoding in network space. The ablation studies confirm that both magnitude and rotation components are necessary and complementary.
The experimental evaluation is extensive and rigorous. The authors train a model ladder from 58M to 1.3B parameters on C4, comparing LayerRoPE against Pre-Norm, Post-Norm, Peri-Norm, and Layer-Norm Scaling. LayerRoPE consistently outperforms baselines, reaching Pre-Norm's 1.3B loss with 3.4x less compute. Crucially, the paper demonstrates superior depth scaling, with LayerRoPE being the only method to show monotonic loss improvement up to 512 layers, whereas baselines diverge or plateau. The paper also shows improved learning rate robustness (wider basin, delayed divergence) and successful transfer to looped latent models (Parcae) and Vision Transformers (ViT) without tuning. The analysis of the learned depth schedule reveals that LayerRoPE widens the residual stream rather than damping it, which is a counter-intuitive and significant finding.
The paper provides detailed hyperparameters, training recipes, and compute accounting. It specifies the initialization of depth slopes, rotation base frequency, and optimizer settings. The use of standard datasets (C4, ImageNet) and architectures (LLaMA-style, ViT) enhances reproducibility. However, the specific code implementation is not linked in the text provided, which is a minor gap for immediate reproduction, though the description is sufficiently detailed for re-implementation.
The primary limitation is the scale of the experiments. While 1.3B parameters and 512 layers are significant, the method's behavior at frontier scales (100B+) is extrapolated or tested at reduced budgets (6.7B). The paper relies on C4 for language modeling, which may not fully capture the benefits on more diverse or multilingual data. Additionally, the "widening" of the residual stream requires careful monitoring to ensure it doesn't lead to numerical instability in mixed-precision training, although the paper notes stability improvements.
This paper has high potential impact by challenging a fundamental assumption in Transformer architecture design: that activation norm growth is a defect. By providing a principled way to exploit this growth for depth encoding, it offers a path to training deeper, more stable, and more compute-efficient models. The transferability to looped models and ViTs suggests broad applicability across modalities. The reduction in learning rate sensitivity is particularly valuable for practitioners, as it simplifies hyperparameter tuning for deep networks. LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology
LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
The paper proposes LASER, a framework for offline RL that combines support-constrained latent space policies with explicit entropy regularization. The core methodological contribution is the use of "adjoint matching" to optimize the entropy-regularized policy objective in latent space without backpropagation through time (BPTT). This is a significant technical advancement over previous flow-based offline RL methods (like ReFORM or DSRL) which either relied on implicit regularization or computationally expensive BPTT. The derivation of the optimal Q-tilted distribution and its implementation via stochastic optimal control (SOC) is mathematically sound and provides a clear path for stable training. The separation of support constraints (handled by the flow decoder) from policy extraction stability (handled by entropy regularization) is a well-motivated architectural choice.
The experiments are conducted on the OGBench benchmark, covering 40 tasks across locomotion and manipulation. The results show that LASER achieves state-of-the-art performance, particularly on noisy datasets where support constraints are critical. A key strength is the demonstration of robustness: LASER uses a fixed set of hyperparameters across all tasks, outperforming baselines that require task-specific tuning. The ablation studies effectively isolate the contribution of entropy regularization and the efficiency of adjoint matching vs. BPTT (showing a 3.6x speedup). The comparison against recent flow-based baselines (IFQL, FQL, QAM) is appropriate and rigorous.
The paper provides a clear algorithmic summary and references a project website. The use of standard benchmarks (OGBench) and fixed hyperparameters enhances reproducibility. However, the specific implementation details of the adjoint matching solver and the flow matching training are deferred to the appendix, which is standard but requires careful reading. The code availability is implied by the project URL.
The method relies on the quality of the behavior cloning (BC) decoder to enforce support constraints; if the decoder is imperfect, OOD actions can still occur. The computational cost of solving ODEs during training remains high, though mitigated by adjoint matching. The paper acknowledges that the value function learning is relatively simple compared to more advanced pessimistic methods, suggesting room for improvement in the critic component.
This work advances the state of the art in safe offline RL, which is critical for deploying RL agents in real-world robotic systems where online interaction is costly or dangerous. The technique of using adjoint matching for entropy-regularized flow policies could be adapted to other generative model-based RL settings, potentially influencing how expressive policies are trained in broader RL research. LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.
Primary: Tsinghua University
All Institutions: Tsinghua University
The paper presents a rigorous and effective framework for evolving RL algorithms using LLMs, demonstrating significant performance gains and discovering interpretable, transferable mechanisms. Its combination of structured search (PCE) and efficient evaluation (PPE) addresses key bottlenecks in program evolution for RL, offering a practical and impactful contribution to the field of automated machine learning.
The paper introduces RLDiscover, a framework for the self-evolution of model-free deep RL algorithms using LLMs. The core methodological contribution is the combination of Progressive Co-Evolution (PCE) and Progressive Probabilistic Evaluation (PPE). PCE addresses the difficulty of joint search over coupled algorithmic components by first using a UCB-inspired bandit to identify high-value components for isolated mutation, then transitioning to joint co-evolution. PPE addresses the high cost and noise of RL evaluation by implementing a multi-fidelity funnel (L0: 50k steps, L1: 200k steps, L2: 1M steps) that allocates compute efficiently. The decomposition of RL algorithms into five specific components (state encoding, value target, critic loss, policy loss, exploration bonus) is a structured and logical choice that enables targeted LLM mutations. The methodology is sound, leveraging established concepts (bandits, multi-fidelity optimization) in a novel application context.
The experiments are extensive, covering three major RL families (SAC, PPO, DQN) across 26 task-algorithm pairs from four benchmark suites (DMControl, Gymnasium, Meta-World, MetaDrive). The results show substantial improvements, with median gains of 32-84% over baselines. A particularly strong aspect is the "SOTA-init" experiment, where evolution starts from strong open-source implementations (CleanRL, denisyarats/pytorch_sac) and still yields improvements, validating that the gains are not merely due to fixing a weak baseline. The ablation studies are rigorous, comparing the full method against variants removing PCE or PPE components, and analyzing the trade-off between search breadth and evaluation fidelity. The cross-task transfer analysis provides evidence that the evolved programs contain generalizable mechanisms rather than just task-specific hacks.
The paper provides detailed descriptions of the component interfaces, the PCE algorithm, and the PPE promotion rules. It explicitly states that code and data will be released upon acceptance. The "AI use statement" clarifies that the LLM is the mutation operator, not just a writing aid, and details the audit process for generated programs. The reproducibility statement is thorough, specifying training budgets, seeds, and evaluation protocols. The inclusion of per-seed results and the discussion of variance (e.g., the bimodal outcome of MountainCar) enhances transparency.
The authors acknowledge that PPO and DQN results are suboptimal due to budget constraints carried over from SAC. The five-component decomposition is only active at the L0 screening stage, limiting structured exploration at higher fidelities. L0 screening at 50k steps may discard strategies with slow warm-up. The LLM identity is withheld for anonymity, which is standard for submission but limits immediate reproducibility of the exact mutation behavior until publication.
This paper contributes to the growing field of LLM-driven algorithm discovery. By demonstrating that LLMs can evolve interpretable, reusable RL mechanisms (like robust critic losses and adaptive Q-blending) that transfer across tasks, it suggests a path toward automated RL algorithm design. The findings on "convergent discovery" of specific mechanisms provide insights into what makes RL algorithms robust, potentially guiding human designers. The efficiency gains from PPE (15x reduction in compute) make this approach more practical for broader adoption. The paper presents a rigorous and effective framework for evolving RL algorithms using LLMs, demonstrating significant performance gains and discovering interpretable, transferable mechanisms. Its combination of structured search (PCE) and efficient evaluation (PPE) addresses key bottlenecks in program evolution for RL, offering a practical and impactful contribution to the field of automated machine learning.
Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.
Primary: Technical University of Denmark (DTU)
All Institutions: Technical University of Denmark (DTU), Novo Nordisk Foundation
The paper introduces a novel framework for molecular discovery that leverages information-dense stochastic synthesis to achieve exponential speedups in finding rare active molecules. By rigorously combining Bayesian experimental design with combinatorial chemistry, it provides a theoretically grounded and empirically promising solution to the sparsity problem in lab-in-the-loop systems, though its practical impact awaits wet-lab validation.
The paper proposes "Lab-in-the-loop learning with Information-Dense Synthesis" (LIDS), a framework that shifts molecular discovery from testing individual molecules to testing complex mixtures (pools). The core innovation is the application of Bayesian experimental design, specifically maximizing Expected Information Gain (EIG), to the design of these mixtures. The authors provide a rigorous theoretical foundation, proving that under specific conditions (e.g., discrete activity, no noise), this approach can reduce the number of experiments from $O(d)$ to $O(1)$ or $O(\log d)$. They introduce a tractable sparse function class (a Bayesian neural network with exponential nonlinearity) to model the molecule-activity map, which allows for closed-form computation of the inner product between the activity map and the mixture distribution. The method uses variational inference to approximate the posterior and optimize the synthesis parameters. The theoretical contribution is strong, particularly the proof that standard acquisition functions like UCB and EI fail to exploit stochastic synthesis, whereas EIG does.
The experiments are conducted entirely in simulation. The authors evaluate LIDS on two types of oracles: (1) a synthetic sparse DNA sequence oracle and (2) a protein language model (ProGen2) based oracle for peptide design. The results show significant speedups, with LIDS finding optimal molecules in 2-5 experiments compared to 100+ for baselines like Thompson Sampling and Gaussian Process Bayesian Optimization. The ablation studies effectively demonstrate that the gains come from the stochastic synthesis design rather than the specific surrogate model. However, the lack of wet-lab validation is a significant weakness, as the assumptions of linearity in the assay and the ability to synthesize arbitrary mixtures with high fidelity are not tested in a real-world setting.
The paper provides detailed algorithmic descriptions and hyperparameters for the simulations. It specifies the use of NumPyro for probabilistic programming and provides details on the variational distribution architecture and training procedures. However, the code is not explicitly linked in the provided text, and the reliance on specific protein language models and synthetic oracles makes direct reproduction of the exact results dependent on access to these models and the specific simulation setup.
The primary limitation is the absence of experimental validation in a physical laboratory. The method assumes that mixtures can be synthesized and tested with high fidelity and that the assay response is linear with respect to the mixture composition, which may not hold for all biological assays. Additionally, the computational cost of LIDS is significantly higher than individual synthesis methods, requiring substantial GPU resources for EIG optimization and posterior updates. The method also relies on known noise levels, which may not always be available in practice.
If validated in the wet lab, this approach could revolutionize automated molecular discovery by drastically reducing the number of experiments needed to find rare active molecules. It offers a new paradigm for scaling laboratory feedback by increasing the information density of each experiment rather than just the number of experiments. This could have significant implications for drug discovery, materials science, and other fields where searching large combinatorial spaces is a bottleneck. The paper introduces a novel framework for molecular discovery that leverages information-dense stochastic synthesis to achieve exponential speedups in finding rare active molecules. By rigorously combining Bayesian experimental design with combinatorial chemistry, it provides a theoretically grounded and empirically promising solution to the sparsity problem in lab-in-the-loop systems, though its practical impact awaits wet-lab validation.
A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents' private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent's contribution scales with its squared centrality in a network reweighted by agents' preferences across activities. This decomposition motivates Poll, a polling algorithm in which the planner samples one agent per round, walks briefly through the agent's neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, Poll uses significantly fewer operations than exact computation and other distributed methods, requiring up to three orders of magnitude less communication on a real-world network with over 300,000 agents; statistically, its query complexity scales with topology and preference heterogeneity rather than explicitly with population size; and economically, it converges to welfare-maximizing prices while admitting behavior-specific implementations that induce truthful reports and detect adversarial deviations.
Primary: MIT
All Institutions: MIT
The paper introduces a scalable polling algorithm for network intervention that leverages centrality-based welfare decomposition to achieve computational, statistical, and economic efficiency. By combining stochastic approximation with mechanism design techniques like decoy signals, it offers a robust solution to the entangled challenges of learning, strategy, and computation in large-scale networked systems, representing a significant advance in algorithmic game theory and distributed optimization.
The paper proposes a novel framework for network intervention where a planner sets non-discriminatory prices in a multi-activity network game with heterogeneous private technologies. The core theoretical contribution is a centrality-based decomposition of the welfare kernel, showing that optimal prices depend on agents' squared centralities in a network reweighted by their preferences. This insight motivates "Poll," a stochastic approximation algorithm that samples one agent per round, performs local random walks, and updates prices based on local reports. The methodology is rigorous, combining game theory, network science, and stochastic optimization. The introduction of "decoy signals" to induce truthful reporting via Knightian uncertainty is a particularly creative and theoretically sound mechanism design contribution that avoids the need for monetary payments.
The experiments are extensive, covering synthetic networks (hierarchical, ER, regular) and a large real-world dataset (DBLP coauthorship network with ~317k agents). The paper demonstrates significant computational efficiency, claiming up to three orders of magnitude less communication than distributed baselines like Federated Learning and Consensus. It also validates the learning efficiency by showing query complexity scales with heterogeneity rather than population size, and the economic efficiency by demonstrating robustness to strategic manipulation under various ambiguity attitudes. The results are convincing and align well with the theoretical predictions.
The paper provides detailed descriptions of the algorithm, the network models, and the experimental setup. However, no code repository or specific implementation details (e.g., hyperparameter sweep ranges, exact random seeds) are provided in the text. The reliance on specific graph structures and preference distributions makes full reproduction dependent on the authors' released code, which is not linked.
The model assumes linear-quadratic utilities, which may not capture all real-world strategic behaviors. The incentive guarantee relies on agents being ambiguity-averse (max-min preferences), which is a strong behavioral assumption. The "decoy" mechanism requires the planner to generate and send multiple signals, which adds communication overhead, though still less than full aggregation. The analysis assumes the planner does not know the network topology, but the algorithm's performance depends on the network's spectral properties, which might be hard to estimate in dynamic networks.
This work has significant implications for platform design, public policy (subsidies, tariffs), and distributed AI systems. It provides a scalable, privacy-preserving, and incentive-compatible method for coordinating large-scale networks of strategic agents. The insights into centrality and preference heterogeneity could inform future designs of federated learning systems and market mechanisms. The paper introduces a scalable polling algorithm for network intervention that leverages centrality-based welfare decomposition to achieve computational, statistical, and economic efficiency. By combining stochastic approximation with mechanism design techniques like decoy signals, it offers a robust solution to the entangled challenges of learning, strategy, and computation in large-scale networked systems, representing a significant advance in algorithmic game theory and distributed optimization.
AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain "reduced-concurrency" versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.
Primary: Microsoft Research
All Institutions: University of California, Riverside, Math, Inc., Microsoft Research
The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
The paper proposes RESOLVE, a three-stage pipeline for validating GPU kernels: (1) binary instrumentation to test for nondeterminism/races, (2) agent-based rewriting to create "reduced-concurrency" versions that are bitwise equivalent to the originals, and (3) formal verification of the reduced versions using F*/Pulse. The core novelty lies in the decoupling of concurrency correctness (handled via testing) from functional correctness (handled via proof), and the use of LLM agents to bridge the gap between complex production kernels and the limited semantic coverage of formal verifiers. The approach is clever but relies heavily on the agent's ability to produce correct reductions, which is a non-trivial assumption. The use of bitwise equivalence as an oracle for the reduction step is a strong design choice that avoids numerical tolerance issues during the intermediate phase.
The evaluation covers KernelBench, fused GEMMs in CUTLASS/Triton/Gluon, and mega-kernels. The finding of four previously unreported issues (including two bugs) in state-of-the-art frameworks is a significant empirical result, demonstrating the tool's practical utility. The comparison against existing tolerance-based tests shows that RESOLVE catches errors that standard testing misses. However, the evaluation is somewhat limited in scale (only three mega-kernels) and lacks a detailed analysis of the cost (time/compute) of the agent-based reduction and proof steps. The claim that agents can repair the issues with minimal performance impact is supported but would benefit from more extensive benchmarking.
The paper describes the pipeline clearly, but the reliance on "agents" to perform rewrites and proofs introduces variability. Without a fixed prompt strategy or a deterministic agent framework, exact reproduction of the results may be difficult. The use of NVBit and F*/Pulse is standard, but the specific agent configurations are not detailed enough for full reproduction. The code availability is not explicitly stated in the provided text, which is a gap for a systems paper.
The primary limitation is the dependence on LLM agents for the reduction and proof steps. If the agent fails to produce a valid reduction or proof, the pipeline stalls. The paper does not deeply analyze the failure modes of the agent or the success rate of the reduction step across a larger corpus. Additionally, the formal verification step is limited to the subset of CUDA/Kuiper supported by F*/Pulse, meaning kernels with exotic hardware features may still require manual intervention or may not be verifiable. The performance overhead of the validation pipeline itself is not thoroughly quantified.
This work has high potential impact on the reliability of AI-generated code, particularly in high-stakes domains like autonomous driving or financial modeling where GPU kernel correctness is critical. It provides a framework for integrating formal methods into the agentic coding loop, which is a growing area of interest. The discovery of bugs in production frameworks like CUTLASS and Triton highlights the immediate value of such tools. It may influence the development of future kernel compilers and verification tools to be more amenable to automated reduction and proof. The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed graph. We realize this task-general view through a common interface based on random walks, allowing the same model to operate across homogeneous and multi-relational graphs with varying features, labels, and relational schemas. Wander can increase its structural context at inference time without changing its learned parameters and, under suitable assumptions, universally approximates the corresponding Bayes-optimal predictor on bounded connected graphs. Empirically, a single pretrained checkpoint achieves state-of-the-art or highly competitive results across node classification, homogeneous link prediction, and knowledge-graph link prediction. Moreover, joint pretraining across graph modalities and tasks preserves performance in specialized settings while enabling positive transfer and the composition of separately learned capabilities at inference time.
Primary: AITHYRA (Austrian Academy of Sciences)
All Institutions: AITHYRA, Austrian Academy of Sciences, Boehringer Ingelheim Stiftung
Wander introduces a unified graph foundation model using random walks as a shared interface, achieving state-of-the-art performance across node classification, homogeneous link prediction, and knowledge graph link prediction while demonstrating positive transfer and compositional generalization. The paper provides a rigorous theoretical framework proving universality and permutation equivariance, and empirically validates the model's ability to adapt its structural context at inference time, offering a promising path toward generalist graph learning.
The paper proposes "Wander," a graph foundation model that unifies node classification, homogeneous link prediction, and knowledge graph link prediction under a single probabilistic framework: the completion of a partially observed graph. The core architectural innovation is the use of random walks as a shared computational interface, replacing standard message passing. This allows the model to dynamically adjust its structural context at inference time without retraining. The methodology is rigorous, providing a theoretical foundation that proves Wander is a universal approximator of the Bayes-optimal predictor for bounded connected graphs and is permutation equivariant in distribution. The design separates feature, label, and structural channels, processing them via intra-node attention, in-context learning updates, and stochastic structural updates via sampled walks. This is a sophisticated and well-motivated approach to the fragmentation problem in graph foundation models.
The experimental evaluation is extensive and addresses four key questions: generalist performance, transfer through joint training, compositional generalization, and inference-time adaptation. The model achieves state-of-the-art or highly competitive results across 54 knowledge graphs, 10 homogeneous link prediction datasets, and 26 node classification datasets. Notably, the paper demonstrates positive transfer from joint pretraining (improving homogeneous link prediction by ~3% over single-task training) and compositional generalization (combining node features and edge types, which were never seen together during training). The inference-time adaptation experiments on the synthetic Grids task are particularly compelling, showing that increasing walk length significantly improves long-range reasoning accuracy from 56.6% to 96.6%.
The paper provides detailed implementation details in the appendix, including initialization strategies, specific attention mechanisms (RoPE, RMSNorm), random walk sampling protocols, and loss functions. The pretraining data generation process is described, combining synthetic graph models (SBM, ER, Watts-Strogatz) with real-world KGs. While the code is not explicitly linked in the provided text, the level of detail suggests high reproducibility. The use of standard benchmarks and clear evaluation protocols (MRR, Hits@10, Recall@20, Accuracy) further supports reproducibility.
The pretraining prior is limited, relying heavily on synthetic data for attributed graphs and only three real-world knowledge graphs for multi-relational pretraining. Inference cost can be high due to the stochastic nature of random walk sampling, particularly when large budgets are required for long-range reasoning. The model does not yet support node regression or graph classification tasks. The theoretical guarantees rely on capacity assumptions that may not hold for the finite-sized models used in practice.
This paper makes a significant contribution to the field of graph machine learning by demonstrating that a single model can effectively handle diverse graph modalities and tasks. The random walk interface offers a flexible alternative to message passing, with potential applications in dynamic graphs and scenarios where computational resources can be allocated adaptively at inference time. The positive transfer and compositional generalization results suggest that joint pretraining across tasks is a viable strategy for building more robust graph foundation models. Wander introduces a unified graph foundation model using random walks as a shared interface, achieving state-of-the-art performance across node classification, homogeneous link prediction, and knowledge graph link prediction while demonstrating positive transfer and compositional generalization. The paper provides a rigorous theoretical framework proving universality and permutation equivariance, and empirically validates the model's ability to adapt its structural context at inference time, offering a promising path toward generalist graph learning.
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-$k$ misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.
Primary: New York University
All Institutions: New York University, NYU Shanghai
The paper introduces a distributionally robust training objective for MoE models that improves routing reliability by strengthening non-top expert paths. It provides a rigorous theoretical framework and empirical evidence that optimizing for expert competence, rather than just load balance, leads to more robust and specialized MoE models at scale.
The paper proposes Distributionally Robust Mixture-of-Experts Training (DRMoET), a drop-in training objective that treats layer-wise experts as endogenous robustness groups. Unlike standard load-balancing losses that optimize for traffic distribution, DRMoET optimizes for expert competence under imperfect routing. It utilizes an entropy-regularized softmax update on EMA-smoothed, activation-weighted expert losses to strengthen plausible non-top routing paths. The method is theoretically grounded, providing convergence guarantees for the entropy-regularized robust objective, and is computationally efficient, adding negligible overhead to the standard MoE forward/backward pass. The distinction between "allocation" (load balancing) and "competence" (routing robustness) is a well-motivated and clear conceptual contribution.
The authors conduct rigorous experiments at two scales (746M and 10.3B total parameters) using the FLAME-MoE recipe. Results show consistent improvements in downstream task averages over both standard FLAME-MoE and auxiliary-loss-free baselines. Mechanistic analyses are strong, including expert-loss variance reduction, forced misrouting probes (showing 4.3% lower excess loss), and improved domain-expert specialization metrics. The ablation studies effectively isolate the contribution of activation-weighted credit and EMA decay. The inclusion of throughput analysis confirms the method's practical viability.
The paper provides detailed hyperparameters, data sources (DCLM), and training recipes. Code and project page are available. The method is described as a "drop-in" objective, suggesting ease of integration into existing MoE training pipelines. The specific implementation details of the EMA and dual variable updates are clearly specified in the algorithm box and appendix.
The evaluation is limited to two model scales and specific expert configurations. The paper acknowledges that it does not test highly imbalanced pretraining mixtures or long-tail evaluation suites, which would be a more direct stress test for robustness. The gains, while consistent, are moderate (e.g., 1.42 percentage points at 10.3B scale), which may be considered small in the context of large-scale LLM training where other factors often dominate.
As MoE architectures become standard for scaling LLMs, addressing the reliability of sparse routing is increasingly important. This work provides a practical tool to improve the robustness of MoE models without architectural changes, potentially leading to more reliable inference under distribution shift or imperfect gating. It shifts the focus from mere utilization balance to expert quality, a perspective likely to influence future MoE training strategies. The paper introduces a distributionally robust training objective for MoE models that improves routing reliability by strengthening non-top expert paths. It provides a rigorous theoretical framework and empirical evidence that optimizing for expert competence, rather than just load balance, leads to more robust and specialized MoE models at scale.
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
Primary: University of California, Berkeley
All Institutions: University of California, Berkeley, Massachusetts Institute of Technology (MIT)
The paper demonstrates that base models' reasoning capabilities are heavily influenced by specific starting token cues learned from training data, and that these cues can be manipulated to significantly enhance performance. This is a significant contribution to mechanistic interpretability and practical LLM optimization, offering a cost-effective alternative to RL for improving reasoning tasks.
The paper employs a rigorous causal intervention framework to dissect the role of "token cues" in Large Language Model (LLM) reasoning. The methodology is three-pronged: (1) Empirical demonstration that fixing specific starting tokens (e.g., ".\n\nOkay") significantly boosts base model performance on math and coding tasks, rivaling RL-trained models; (2) Causal data interventions (likely using activation patching or data ablation techniques) to prove that these cues are learned from specific document types in the training data, rather than being inherent to the token embeddings; (3) A safety case study showing that cues also modulate refusal behaviors. The approach is sophisticated, moving beyond correlation to causation by editing the model's internal representations or training data associations to isolate the effect of the cue.
The experiments are extensive and compelling. The authors test across multiple model families (OLMo-3, Qwen3) and tasks (MATH-500, coding). The jump in accuracy (e.g., 42% to 78% for OLMo-3-7B) is substantial and practically significant. The causal interventions are well-designed to rule out confounding factors. The safety analysis adds depth, showing that the same mechanism that aids reasoning also influences safety alignment, which is a critical insight for the field.
The paper appears to be from a high-quality group (Berkeley/MIT) with clear methodological descriptions. While specific code links are not provided in the snippet, the detailed description of the causal interventions and the use of open-source models (OLMo, Qwen) suggests high reproducibility. The abstract-only score of 60 suggests the full text provides sufficient detail for replication.
The study focuses on base models and the transition to RL. It may not fully account for how these cues interact with more complex agentic workflows or multi-turn conversations. The "chicken" example, while illustrative, is a synthetic intervention; the real-world generalizability of arbitrary word cues is limited, though the core finding about training data associations is robust.
This paper has high impact because it challenges the prevailing narrative that RL is the sole driver of reasoning capabilities in LLMs. It suggests that data curation and prompt engineering (specifically starting tokens) are underutilized levers for improving base model performance. This has immediate implications for model developers, who can optimize pre-training data and inference prompts without expensive RL fine-tuning. It also has safety implications, as it reveals how minor prompt changes can shift model behavior between compliance and refusal. The paper demonstrates that base models' reasoning capabilities are heavily influenced by specific starting token cues learned from training data, and that these cues can be manipulated to significantly enhance performance. This is a significant contribution to mechanistic interpretability and practical LLM optimization, offering a cost-effective alternative to RL for improving reasoning tasks.
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
Primary: TU Darmstadt
All Institutions: TU Darmstadt, Hessian.AI Service Center, Konrad Zuse School of Excellence in Learning and Intelligent Systems
The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
The paper introduces Privacy-Leaking Watermarks (PLWs), a novel adversarial attack on unified multimodal models. The core innovation lies in exploiting the shared latent space of unified architectures to embed trigger-dependent watermarks in generated images based on prior conversational context. The methodology involves a two-stage training process: first, training a watermark encoder/extractor pair, and second, fine-tuning the multimodal model (using LoRA) to condition the watermark embedding on specific semantic triggers found in the chat history. This approach is technically sound and cleverly leverages the specific architectural properties of unified models (where text and image generation are not strictly decoupled) to create a covert side-channel for privacy leakage.
The experiments are rigorous, testing the attack across two model families (including OmniGen2) and 13 sensitive attribute triggers. The results are striking, achieving up to 100% True Positive Rate at a 1% False Positive Rate, demonstrating that the watermarks are both highly detectable and effectively hidden from casual observation. The evaluation correctly isolates the threat model to locally deployed models where users assume privacy, providing a clear and compelling demonstration of the vulnerability.
The paper provides a strong reproducibility statement, detailing the training configurations, LoRA target modules, and specific seeds used. While the authors explicitly state they do not release poisoned checkpoints (a responsible security practice), the detailed appendices and methodological transparency allow for independent verification of the attack's feasibility.
The primary limitation is the requirement for white-box access to modify and redistribute model weights, which restricts the threat to malicious model providers or compromised supply chains rather than arbitrary users. Additionally, the attack is specific to unified multimodal architectures; it may not apply to traditional pipeline-based multimodal systems where text and image generation are separate modules.
This paper has significant implications for the deployment of unified multimodal models, particularly in consumer-facing applications where local deployment is marketed as a privacy feature. It highlights a critical gap in the security model of these architectures and will likely drive research into secure design patterns, watermarking defenses, and the separation of modalities in unified models. The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.
Primary: Georgia Institute of Technology
All Institutions: Georgia Institute of Technology, Simon Fraser University, University of Pennsylvania, University of Texas at Austin, Rutgers University, Foothill College
The paper presents a novel co-evolutionary framework for proof auto-formalization that jointly optimizes a large open-weight LLM and its agentic harness, achieving state-of-the-art results on a new research-level benchmark while significantly improving cost-efficiency compared to frontier API-based agents.
The paper proposes a joint optimization framework for both the language model and its agentic harness. The Semantic Alignment Model (SAM) introduces a contrastive loss to align natural language and formal language representations, which is a logical extension of existing alignment techniques but applied here to a large MoE model. The core novelty lies in HarnessEvolve, an evolutionary search algorithm that treats the agent's control flow (harness) as a mutable program, using diagnostic certificates from verifiers to guide mutations. This is a significant departure from static scaffolding or simple prompt optimization. The Agentic RLSF component adapts Group-Relative Policy Optimization for multi-turn agent interactions, using a fine-grained reward ladder that distinguishes between type errors, incompleteness, and semantic drift. This multi-objective reward design is crucial for preventing the "silent correction" of flawed proofs, a known failure mode in prior work.
The introduction of LoCoBench is a major contribution, providing 58.9k instances and a 771-instance validation set that is out-of-distribution (Mizar-to-Lean translation). The evaluation against 39 baselines is comprehensive. The results show a substantial improvement over the base model (15.7% to 36.7% pass@4 semantic correctness) and demonstrate that the framework can enhance frontier coding agents (Claude Code, Codex) when used as a specialized skill. The cost-efficiency analysis is particularly strong, showing that the open-weight system achieves competitive accuracy at a fraction of the cost of frontier API calls.
The paper provides extensive details on the training setup, including hardware, hyperparameters for LoRA and GRPO, and the specific configuration of the evolutionary search. The release of the benchmark and the open-weight model base (Leanstral-1.5) enhances reproducibility. However, the reliance on a frontier coding agent (Claude Opus 5) for the mutation step in HarnessEvolve introduces a dependency on a proprietary service, which may limit full reproducibility for all researchers.
The primary limitation is the heavy reliance on a frontier LLM for the harness evolution process, which contradicts the goal of full autonomy and open-source accessibility. The semantic correctness check relies on an extended BEq+ prover, which may have its own coverage limitations. Additionally, the "proof faithfulness" metric is not fully automated, relying on LLM judges which are noted to be less reliable.
This work has significant implications for the accessibility of formal verification. By demonstrating that open-weight models can be post-trained to perform research-level auto-formalization, it lowers the barrier to entry for formal methods. The co-evolution of model and harness offers a generalizable paradigm for improving agentic systems in other domains where tool-use and control flow are critical. The paper presents a novel co-evolutionary framework for proof auto-formalization that jointly optimizes a large open-weight LLM and its agentic harness, achieving state-of-the-art results on a new research-level benchmark while significantly improving cost-efficiency compared to frontier API-based agents.
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
The paper demonstrates that high accuracy on scientific benchmarks does not imply evidence grounding, introducing a verified counterfactual protocol to measure and address this gap. By showing that agents often rely on prior knowledge rather than supplied data, and that standard evaluators fail to penalize this, the work provides a crucial methodological correction for the evaluation of AI scientists.
The paper introduces a rigorous framework for evaluating "evidence grounding" in AI scientific agents. The core methodological contribution is the construction of verified counterfactuals (null, withdrawal, flip) for benchmark items, where the effect of the intervention on the ground truth is pre-registered and verified via reference statistics. This allows for the calculation of Evidence-Grounded Accuracy (EGA), which distinguishes between answers derived from data versus those recalled from prior knowledge or reached via elimination. The formalization of "noticing without acting" and the separation of agent grounding from evaluator validity are strong theoretical contributions. The protocol for cleaning metadata to prevent leakage (canonical writer) is a crucial practical detail that enhances the validity of the counterfactuals.
The experiments are well-designed but limited in scale. The authors evaluate 18 single-cell questions from BAISBench and 4 synthetic problems from GeneBench-Pro. While small, this is sufficient to demonstrate the phenomenon of high accuracy coexisting with low evidence grounding. The finding that Claude agents achieve 95% accuracy but only 41% EGA is a significant empirical result. The gene-name anonymization experiment effectively isolates prior knowledge as a confounding factor. The evaluation of evaluators (benchmark scores and LLM judges) showing they reward ungrounded answers is a critical finding for the field.
The paper provides high reproducibility. It details the canonical writer process, the specific interventions applied to the data, the agent scaffolds used (Claude Code, Codex CLI, Biomni), and the evaluation metrics. The pre-registration of reference statistics and the verification of every edit make the counterfactuals robust. The code and data protocols are described in sufficient detail for replication, although the specific datasets (BAISBench, GeneBench-Pro) are external dependencies.
The primary limitation is the small number of tasks (22 total). The findings may not generalize to all types of scientific tasks or domains beyond single-cell biology and synthetic genetics. The manual construction of counterfactuals is labor-intensive, limiting scalability. The study focuses on a specific set of models (Claude, GPT-5.6) and scaffolds, so results may vary for other architectures. The "flip" intervention assumes a binary or clear alternative, which may not apply to all scientific questions.
This paper has high impact on the evaluation of AI agents in scientific discovery. It challenges the validity of current benchmarks that rely solely on final answer accuracy. The proposed metrics (EGA, counterfactual validity) provide a new standard for assessing whether AI agents truly use the data provided. This work will likely influence the design of future benchmarks for AI scientists and the interpretation of existing leaderboard scores. It highlights a critical gap in current AI evaluation practices that affects the trustworthiness of AI-generated scientific insights. The paper demonstrates that high accuracy on scientific benchmarks does not imply evidence grounding, introducing a verified counterfactual protocol to measure and address this gap. By showing that agents often rely on prior knowledge rather than supplied data, and that standard evaluators fail to penalize this, the work provides a crucial methodological correction for the evaluation of AI scientists.
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Primary: University of New South Wales (UNSW Sydney)
All Institutions: University of New South Wales (UNSW Sydney)
ASCENT enables stable online test-time training of LLM agents by self-distilling verified experience through a frozen teacher model, significantly improving task success and efficiency in long-horizon environments without requiring external references or multiple attempts.
The paper proposes ASCENT, a method for Online Agentic Test-Time Training (OaTTT). The core innovation is using the frozen initial model as a teacher that conditions on the "privileged information" of a verified successful trajectory (hindsight) to generate soft targets for the current student model (via LoRA). This avoids the instability of direct imitation or reinforcement learning on single-attempt trajectories, which the authors demonstrate collapses the policy. The method effectively combines rejection sampling (only updating on verified successes) with on-policy distillation. The theoretical framing of the population target and the analysis of why direct imitation fails (sharpening around generated tokens vs. full-vocabulary matching) are strong contributions.
The experiments are extensive, covering ALFWorld, WebShop, and AppWorld with two model scales (Qwen3.5-4B and 9B). The paper provides strong evidence for the method's efficacy, showing significant improvements over the base model and various in-context adaptation baselines (MemP, ACE, etc.). The ablation studies on privileged information content and distillation divergence are thorough. The demonstration that direct imitation baselines fail catastrophically is a valuable negative result for the community.
The paper provides detailed descriptions of the protocol, loss functions, and hyperparameters. The use of standard benchmarks and open-weight models (Qwen) aids reproducibility. However, the specific implementation of the "validity filter" and the exact serialization of privileged information $z_i$ could benefit from more code-level detail, though the project page is provided.
The method requires access to the model's weights (open-source models only) and the ability to run a frozen copy of the model as a teacher, which doubles inference cost during the update phase. The reliance on sparse, episode-level verification limits the granularity of learning signals. The paper acknowledges that the teacher's privileged context may allow it to reconstruct student tokens, potentially limiting the independence of the hindsight signal.
This work is significant for the deployment of LLM agents in dynamic environments where continuous improvement is desired without offline retraining. It bridges the gap between in-context learning (memory-based) and parametric adaptation, offering a stable mechanism for weight updates. The findings on the instability of single-attempt RL/imitation are likely to influence future agent training designs. ASCENT enables stable online test-time training of LLM agents by self-distilling verified experience through a frozen teacher model, significantly improving task success and efficiency in long-horizon environments without requiring external references or multiple attempts.
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology, Qualcomm AI Research
Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
The paper proposes Score-Calibrated Flow (SCF), a method for training flow matching models to sample from unnormalized densities without access to target samples. The core innovation is a "self-consistency" approach that derives optimality conditions for the velocity field based on the target score (provided by the critic in RL) and the structure of the flow itself. The authors prove that the unique solution to these conditions is the ideal flow model that would be learned via standard Conditional Flow Matching (CFM) if target samples were available. The training objective is a stop-gradient regression on self-generated endpoints, avoiding the high variance of importance sampling and the computational cost of backpropagating through sampling trajectories. The theoretical grounding is strong, with clear proofs of uniqueness for both the terminal density and the velocity field.
The experiments include a toy example comparing SCF against recent samplers (AS, ASBS, BMS, FS) and extensive RL benchmarks (DeepMind Control Suite, HumanoidBench). SCF is compared against 8 strong baselines, including SAC and various generative policy methods (DIME, QSM, QFlex, etc.). The results show SCF matching or improving upon state-of-the-art methods while significantly reducing training time. The inclusion of wall-clock time comparisons is a strong practical contribution.
The paper provides detailed algorithmic descriptions and theoretical derivations. However, as a preprint, specific hyperparameters and code availability are not confirmed in the text. The method relies on standard flow matching components, making it relatively easy to implement if the specific coefficient schedules (mentioned in the appendix) are clear.
The method assumes access to the gradient of the potential function (critic), which is standard in RL but may not be available in all sampling tasks. The theoretical guarantees rely on certain regularity conditions (e.g., exponential moments) that may be hard to verify in practice for complex, high-dimensional distributions. The performance gains, while present, are incremental over strong baselines in some tasks.
This work addresses a fundamental bottleneck in applying generative models to online RL and other settings where only unnormalized densities are available. By providing a stable, efficient, and theoretically grounded training procedure, it could become a standard tool for training expressive policies in RL and for sampling in physics and molecular dynamics simulations. Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
Primary: Northwestern University
All Institutions: Northwestern University, Carnegie Mellon University, UNC Chapel Hill
The paper establishes visual transition reasoning as a shared, trainable primitive for multimodal LLMs and provides a systematic training recipe validated by a new benchmark, WOVEN. By demonstrating that small, targeted subsets of transition supervision can broadly improve performance across diverse spatial, physical, and temporal benchmarks, the work offers a data-efficient and theoretically grounded approach to enhancing visual world modeling capabilities in current MLLMs.
The paper proposes a novel framework for "visual transition reasoning," formalizing it as inference over $(s, a, s')$ triplets. The core methodological contribution is the WOVEN dataset, which uses video-pretrained generative models (Wan2.2) to create controlled rollouts with known actions and endpoints. This allows for the construction of multiple-choice questions with error-typed distractors, enabling fine-grained analysis of model failures. The training methodology involves supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO) on controlled subsets of the data. The systematic ablation study comparing supervision by reasoning operation versus action type is a strong methodological choice that isolates the source of transfer.
The experimental evaluation is extensive and rigorous. The authors evaluate 38 frontier MLLMs to establish a baseline deficit. They then demonstrate that training on small subsets (~2,000 items) of WOVEN significantly improves performance on 22 of 26 external benchmarks. A key strength is the causal intervention experiment, where replacing 30-50% of task-specific training data with WOVEN data yields comparable accuracy, proving the utility of the dataset as a data-efficient substitute. The correlation analysis ($r=0.86$) between WOVEN accuracy gains and external benchmark gains provides strong evidence for the "shared primitive" hypothesis.
The paper provides high reproducibility. Code and datasets are released on GitHub and HuggingFace. The generation process using Wan2.2 is specified, and the training recipes (SFT + GRPO) are detailed. The use of standard benchmarks for evaluation ensures that results can be verified against existing baselines.
The reliance on synthetic video generation (Wan2.2) may introduce domain gaps compared to real-world physics, although the authors argue for realism. The evaluation is primarily on multiple-choice tasks, which may not fully capture open-ended reasoning capabilities. The "frontier" models evaluated include some very recent releases (GPT-5.4, Qwen3-VL), which may limit immediate reproducibility for labs without access to these specific API versions or weights.
This paper has significant potential to influence how multimodal models are trained for embodied AI and physical reasoning. By identifying visual transition reasoning as a reusable primitive, it offers a data-efficient path to improving spatial and temporal understanding. The systematic recipe for selecting supervision (by reasoning operation rather than action) provides actionable guidelines for practitioners building world models. The paper establishes visual transition reasoning as a shared, trainable primitive for multimodal LLMs and provides a systematic training recipe validated by a new benchmark, WOVEN. By demonstrating that small, targeted subsets of transition supervision can broadly improve performance across diverse spatial, physical, and temporal benchmarks, the work offers a data-efficient and theoretically grounded approach to enhancing visual world modeling capabilities in current MLLMs.
Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/
Primary: NVIDIA
All Institutions: NVIDIA
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. GATOR introduces a novel generative-agentic framework for 3D object reconstruction that couples a modality mixer for initial asset generation with an iterative agentic refinement loop, achieving strong geometric and appearance fidelity from casual images while recovering scene-relative pose, though the lack of detailed quantitative results in the provided text limits the full assessment of its empirical superiority over existing feed-forward and iterative baselines.
The paper proposes GATOR, a framework combining a "local modality mixer" for feature fusion (RGB, mask, pointmap) with a "multimodal agent" for iterative refinement. The core novelty lies in the agentic loop (edit-render-review) that uses the initial generative output to guide targeted structural and texture corrections. This hybrid approach of one-shot generation followed by agentic refinement is a significant architectural shift from purely feed-forward reconstruction models. The use of stage-specific adapters for text-guided conditioning is a standard but effective technique. The methodology is sound, leveraging recent trends in agentic AI for 3D tasks, though the specific implementation details of the "agent's" decision-making process (e.g., how it selects which parts to refine) are critical to its success and likely complex.
The paper claims strong geometric and appearance fidelity on synthetic objects, cluttered tabletops, and indoor scenes. It includes time-budget comparisons and scene-level simulation readiness tests. However, the provided text is a placeholder with section headers but no actual quantitative results, tables, or figures. Without specific metrics (e.g., Chamfer Distance, FID, PSNR) or comparisons to state-of-the-art baselines (like Zero-1-to-3, Wonder3D, or recent agentic 3D papers), the experimental evaluation cannot be fully verified. The claim of "strong fidelity" is unsubstantiated in the provided text.
The project page is provided, which is a positive sign for reproducibility. However, the lack of detailed hyperparameters, agent prompt engineering details, and training data specifics in the provided text makes independent reproduction difficult. The "agentic" component often relies on proprietary LLMs or specific prompt strategies that may not be fully open-sourced.
The primary limitation is the reliance on an agentic loop, which can be computationally expensive and slow compared to feed-forward methods, despite the "time-budget" claims. The performance on real-world, uncontrolled casual images (as opposed to synthetic or curated datasets) is not clearly demonstrated in the provided text. The "scene-aligned" pose recovery from sparse observations is a difficult problem, and the robustness of this component to severe occlusions or ambiguous viewpoints is a potential weakness.
This work has high potential impact in robotics, AR/VR, and digital twin creation, where accurate 3D asset recovery from casual photos is crucial. The agentic approach could be extended to other 3D tasks like scene editing or animation. It represents a step towards more interactive and intelligent 3D reconstruction pipelines. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. GATOR introduces a novel generative-agentic framework for 3D object reconstruction that couples a modality mixer for initial asset generation with an iterative agentic refinement loop, achieving strong geometric and appearance fidelity from casual images while recovering scene-relative pose, though the lack of detailed quantitative results in the provided text limits the full assessment of its empirical superiority over existing feed-forward and iterative baselines.
Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at https://facebookresearch.github.io/GenIA.
Primary: University of Tübingen
All Institutions: Tübingen AI Center, University of Tübingen, Meta Reality Labs
GenIA introduces a test-time input alignment framework that grounds generative 3D priors in geometric and photometric observations without retraining. The paper presents a rigorous methodology combining geometric pose derivation, visibility-biased attention, and differentiable rendering to improve reconstruction fidelity across static and dynamic scenarios, offering a significant advancement in 3D asset recovery from sparse inputs.
The paper proposes GenIA, a framework for test-time input-aligned generation that grounds the SAM3D generative prior in geometric and photometric observations without retraining the foundation model. The core innovation lies in decoupling the learned rotation prior from translation and scale, which are derived directly from geometry. It employs visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during the denoising process to align appearance. An optional post-denoising refinement step adapts the appearance latent, lightweight decoder adapters, and object placement. The method is designed to handle monocular, sparse multi-view, and dynamic inputs, supporting externally supplied geometry for dynamic objects to recover shared canonical appearance and stable world-space placement.
The authors evaluate GenIA across synthetic and real benchmarks, comparing it against recent optimization-based, per-frame image-to-3D, and video-to-4D methods. The results indicate improvements in pose prediction and object reconstruction. The inclusion of dynamic object recovery from temporal shapes adds a significant dimension to the evaluation, demonstrating the framework's versatility beyond static reconstruction.
The paper provides a project page URL, which likely contains code or further details. The methodology relies on SAM3D, a specific foundation model, which may limit immediate reproducibility if that model is not publicly available or if the specific adapters are not released. However, the test-time nature of the alignment suggests that the core logic is implementable given the base model.
The approach is tightly coupled to the SAM3D foundation model, which may limit its generalizability to other generative priors. The reliance on differentiable rendering and attention mechanisms during denoising could introduce computational overhead, potentially limiting real-time applications. The performance on dynamic objects depends on the quality of the externally supplied temporal shapes.
This work has significant implications for 3D content creation, AR/VR, and robotics, where accurate and complete 3D asset reconstruction from sparse observations is critical. By enabling test-time alignment without retraining, it offers a flexible and efficient path to integrating generative priors with real-world sensor data. GenIA introduces a test-time input alignment framework that grounds generative 3D priors in geometric and photometric observations without retraining. The paper presents a rigorous methodology combining geometric pose derivation, visibility-biased attention, and differentiable rendering to improve reconstruction fidelity across static and dynamic scenarios, offering a significant advancement in 3D asset recovery from sparse inputs.
Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.
Primary: Meta
All Institutions: Meta
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a unified two-stage distillation framework for discrete multimodal dLLMs that effectively compresses inference steps for both image generation and text understanding, achieving state-of-the-art efficiency-performance trade-offs through targeted remedies for entropy collapse and parallel decoding repetition.
The paper proposes Omni-Diffusion-Distill, a two-stage distillation framework specifically designed for unified multimodal diffusion large language models (dLLMs) that operate in a fully discrete token space. The methodology is well-structured, addressing the specific challenges of compressing iterative decoding for both image generation and text understanding within a single architecture. Stage 1 employs off-policy trajectory replay to teach the student to skip decoding steps, while Stage 2 utilizes on-policy refinement on the student's own rollouts to mitigate exposure bias. The core technical contributions are two targeted remedies for modality-specific degradation: an entropy-matched guidance mechanism to prevent entropy collapse in image generation (caused by fitting sharpened CFG targets), and a pairwise collision penalty to reduce repetition in parallel text decoding. The adaptation of distribution matching distillation (DMD) to discrete tokens via an auxiliary network and direct gradient specification is a sound technical adaptation of continuous diffusion techniques to the discrete domain.
The experimental evaluation is comprehensive and rigorous. The authors distill the Lumina-DiMOO model and compare against strong baselines (Di[M]O, CDLM, T3D) across 17 benchmarks covering Text-to-Image (T2I), Image-to-Image (I2I), and Multimodal Understanding (MMU). The results demonstrate significant efficiency gains (18.2x and 21.2x wall-clock speedups) with minimal performance degradation, often outperforming the few-step teacher. The inclusion of ablation studies for the collision penalty and entropy-matched guidance, along with detailed analysis of degeneration metrics (repetition, entropy), provides strong evidence for the effectiveness of the proposed components. The use of GPT-5.5 for open-ended evaluation is a modern and appropriate choice, though it introduces some dependency on external models.
The paper provides high reproducibility. It details the training data mixture, optimization settings, and algorithms for both stages. The specific hyperparameters for the loss terms and the bisection method for entropy matching are described. While the code is not explicitly linked in the provided text, the level of detail in the appendices (algorithms, loss derivations, data sources) suggests that reproduction is feasible for a skilled practitioner. The use of standard benchmarks and clear evaluation protocols further supports reproducibility.
The primary limitation is the scope of validation; the method is only tested on one specific unified dLLM architecture (Lumina-DiMOO). It is unclear how well the framework generalizes to other unified models with different attention mechanisms or tokenization schemes. Additionally, the student model still trails the full-step teacher on certain fine-grained tasks like text rendering and detailed captioning. The reliance on a large teacher model and the computational cost of collecting teacher trajectories, while one-time, remains a barrier for smaller labs.
This work has significant potential impact on the deployment of unified multimodal models. By drastically reducing inference costs while maintaining quality, it makes real-time multimodal interaction more feasible. The techniques developed for handling entropy collapse and parallel decoding artifacts are likely to be applicable to other discrete generative models beyond dLLMs. The framework bridges the gap between high-quality iterative samplers and efficient few-step inference, a critical area for practical AI applications. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a unified two-stage distillation framework for discrete multimodal dLLMs that effectively compresses inference steps for both image generation and text understanding, achieving state-of-the-art efficiency-performance trade-offs through targeted remedies for entropy collapse and parallel decoding repetition.
Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.
Primary: Rice University
All Institutions: Rice University, Baylor College of Medicine, Texas Children's Hospital
EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
The paper proposes EchoDino, a self-supervised foundation model for echocardiography by adapting the DINOv3 framework. The core methodological contribution is the domain adaptation of a natural image foundation model to pediatric cardiac ultrasound using 3.7 million unlabeled frames. The authors introduce two specific architectural/algorithmic extensions: EchoDino-PATCH, which re-aggregates spatial patch tokens to preserve local anatomical detail for segmentation and measurement, and MEMS (Motion-biased Entropy Maximization Sampling), a frame selection strategy for video-level tasks that prioritizes high-motion, feature-diverse frames over uniform sampling. The approach relies on a frozen encoder with lightweight task-specific readouts, which is a standard but effective paradigm for foundation models. The adaptation of DINOv3's teacher-student architecture to the specific augmentation and cropping needs of echocardiography is well-motivated.
The experimental evaluation is extensive and rigorous. The model is tested across nine datasets, including five internal pediatric cohorts, one external pediatric cohort, and three external adult datasets. This cross-domain evaluation (pediatric to adult) is a strong point, demonstrating the generalizability of the learned representations. The tasks cover a wide spectrum: global (view classification), localized (measurement, SHD detection), dense (segmentation), and temporal (EF prediction, sweep recognition). The results show consistent and significant improvements over strong baselines like DINOv3, PanEcho, and EchoPrime. The use of bootstrap confidence intervals and patient-disjoint splits adds statistical rigor. The ablation of MEMS vs. uniform sampling clearly isolates the benefit of the proposed sampling strategy.
The paper provides a link to the code repository. However, the primary pretraining data (TCH-Complex) is not publicly available due to clinical data governance, which limits full reproducibility of the pretraining phase. The downstream evaluation on public datasets (EchoNet, CAMUS) is reproducible. The hyperparameters and training details are described in the supplementary methods, aiding partial reproducibility.
The main limitation is the lack of public access to the large-scale pediatric pretraining corpus, which prevents independent verification of the pretraining process. The SHD detection is evaluated as a binary composite classifier, which may mask performance on specific rare lesions. The paper acknowledges that aggregate metrics do not guarantee individual point-of-care actionability and that uncertainty estimation is needed for clinical deployment.
This work has significant potential impact in the field of medical imaging AI, particularly in pediatric cardiology where labeled data is scarce. By demonstrating that a single frozen encoder can support diverse tasks across age groups, it offers a scalable framework for developing modular clinical tools. The approach of adapting general-purpose vision foundation models to specific medical imaging modalities is a trend that this paper contributes to meaningfully. EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
Primary: University of Michigan
All Institutions: University of Michigan, Southeast University, Alibaba Group
The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
The paper proposes RAESR, a super-resolution model that operates in the latent space of a frozen DINOv3-L vision transformer rather than a VAE or pixel space. The core methodological contribution is the argument that self-supervised representation spaces (specifically DINOv3) embed degraded images closer to the "manifold" of clean images than reconstruction-oriented spaces (VAEs), thereby simplifying the mapping back to the clean manifold. The authors validate this geometric intuition with rigorous ablations, including latent line interpolation experiments and layer-wise information recovery analysis. The decoder is a 415M parameter transformer trained with reconstruction and adversarial losses, using a multi-layer DINOv3-B critic for fine-grained realism supervision. The approach is technically sound, well-motivated, and the ablations are exceptionally thorough, providing strong evidence for the central hypothesis.
The experiments are comprehensive, comparing RAESR against 12 state-of-the-art methods across four real-world benchmarks (RealSR, DRealSR, LSDIR, DIV2K-Val). RAESR achieves the best fidelity-perception trade-off, outperforming heavy diffusion-based models in both quality and efficiency (37ms per image). The inclusion of a human preference study and detailed per-benchmark breakdowns adds robustness. The ablation studies, particularly the comparison against a VAE substrate with identical training recipes, are critical and convincingly demonstrate that the performance gain stems from the choice of latent space rather than just the decoder architecture.
The paper provides extensive details on training schedules, hyperparameters, data processing, and evaluation protocols in the appendices. The use of standard datasets and public baselines facilitates reproduction. However, the reliance on specific frozen checkpoints (DINOv3-L, RAEv2) and the complex multi-stage training process may pose some barriers for independent replication without access to the authors' code or precise checkpoint versions.
The model is a single-pass restorer, lacking the ability to trade compute for quality on difficult images, unlike iterative diffusion methods. The claim about substrate suitability is currently supported primarily by DINOv3-L and tested against one alternative (SD3 VAE), leaving the behavior of other representation encoders unexplored. The model size (719M inference parameters) is still significant compared to lightweight GANs, though it is competitive with diffusion models.
This work shifts the perspective in super-resolution research from merely improving generative priors to carefully selecting the representation space for restoration. It highlights the utility of self-supervised vision foundation models for low-level vision tasks, potentially inspiring similar approaches in other restoration tasks like denoising or inpainting. The efficiency gains over diffusion models make it attractive for real-time applications. The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Primary: ETH Zurich (Inferred from "eth-lre" in GitHub URL and Swiss AI Initiative funding)
All Institutions: ETH Zurich, Swiss National Supercomputing Centre (CSCS)
The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
The paper addresses a critical flaw in existing RL-based tutoring methods: the "telling" problem, where models learn to simply provide answers to maximize immediate student success metrics. The authors propose a novel reward structure based on three conditions: (1) Near-transfer testing (student solves a variant problem, not the exact one), (2) Masked post-test (tutor's utterances are hidden during the test, forcing reliance on student-generated notes/reasoning), and (3) Binary reward gates (strict penalties for factual errors or solution handover). This design effectively decouples "helpfulness" from "teaching," forcing the model to elicit student reasoning rather than provide it. The use of a frozen LLM student with genuine errors (rather than prompted confusion) is a strong methodological choice that ensures the training signal reflects real diagnostic and adaptive challenges.
The experimental setup is rigorous and comprehensive. The authors train models across multiple sizes (4B to 27B) and architectures (Qwen3 variants), demonstrating the robustness of the recipe. The ablation study is particularly strong, isolating the contribution of each condition (transfer, masking, gates) and showing that the gates specifically reduce handover while the masked transfer test improves out-of-domain generalization. The evaluation uses independent benchmarks (MathTutorBench, TutorMoments) and judges (Gemini) distinct from the training setup, mitigating overfitting concerns. The results show that Eduardo-27B matches or exceeds frontier models (Gemini 3.1 Pro, Claude Opus 4.8) in pedagogical quality while using significantly fewer thinking tokens, a crucial efficiency gain for interactive applications.
High. The authors open-source the training environment, the 8,671-problem near-transfer dataset, and the trained models. Detailed hyperparameters, prompts, and compute costs are provided in the appendices. The use of standard RL algorithms (DPPO/GRPO) and open-source base models further enhances reproducibility.
The primary limitation is the reliance on a single frozen LLM student (Llama-3.1-8B-Instruct) for training, which may lead to overfitting to that specific model's failure modes. The paper acknowledges this and suggests future work with diverse student ensembles. Additionally, the evaluation is limited to mathematics, and the "affective gap" (lack of socio-emotional support) is noted as a consequence of optimizing purely for cognitive gain. The ablation study is limited to one seed and one model size (4B), which slightly weakens the statistical confidence in the ablation results.
This work has significant implications for the development of AI tutors and, more broadly, for any domain where an agent must build user capabilities rather than just provide answers. The "masked near-transfer" reward design is a generalizable principle that could be applied to other educational or collaborative tasks. The efficiency gains (fewer thinking tokens) make high-quality tutoring more feasible in real-time interactive settings. The open-sourcing of the dataset and models will likely accelerate research in AI for Education. The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/
Primary: Ben-Gurion University of the Negev
All Institutions: Ben-Gurion University of the Negev
GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
The paper proposes a novel architectural shift in neural audio codecs by replacing the standard Residual Vector Quantization (RVQ) or Finite Scalar Quantization (FSQ) bottleneck with a parametric 1D Gaussian Splatting decomposition. The core idea is to fit the encoder's latent representation as a weighted sum of Gaussian primitives via an inner optimization loop during training. To address the computational cost of this iterative fitting at inference, the authors introduce a "GS Predictor Net" that amortizes the optimization into a single forward pass. The method is technically sound, leveraging differentiable rendering concepts from 3D graphics (Gaussian Splatting) and applying them to 1D signal processing. The use of post-training scalar quantization on the fitted parameters rather than learned codebooks is a distinct design choice that enables fine-grained bitrate control without retraining.
The experimental evaluation is rigorous and comprehensive. The authors compare GS-Codec against strong, well-established baselines such as EnCodec, DAC, and WavTokenizer on standard datasets (LibriTTS, LJSpeech, LibriSpeech). The results show that GS-Codec matches or exceeds these baselines on key metrics like UTMOS (perceptual quality), STOI (intelligibility), and SIM (speaker similarity) at comparable bitrates. The inclusion of human listening tests (MUSHRA/MOS) adds significant credibility to the perceptual quality claims. The ablation studies on primitive count and bit depth provide clear insights into the rate-quality tradeoff. The encoding time analysis honestly reports the latency trade-off, showing that while the iterative version is slow, the Predictor Net brings it close to competitive levels, though still slightly slower than feed-forward baselines.
The paper provides high reproducibility. It details the SEANet backbone hyperparameters, the specific Gaussian Splatting configuration (number of primitives, inner loop steps, learning rates), and the training schedule. The code and audio samples are available via the provided URL. The use of standard metrics and open-source baselines facilitates easy comparison. The detailed appendix on quantization ranges and loss weights further supports reproducibility.
The primary limitation is the encoding latency. Even with the Predictor Net, the encoding time is higher than standard feed-forward codecs like EnCodec and DAC, which may be a bottleneck for real-time applications. The paper also notes that performance at very low bitrates (<3 kbps) is not as strong as specialized low-rate codecs. The method is currently validated primarily on English speech, and generalization to other languages or non-speech audio (music, sound effects) is not extensively explored.
This work opens a new direction in neural audio compression by demonstrating that parametric signal decompositions can serve as effective bottlenecks, challenging the dominance of codebook-based quantization. The ability to control bitrate post-training by varying the number of primitives and bit depth is a practical advantage for deployment. The cross-domain insight from 3D Gaussian Splatting to 1D audio signals is interesting and may inspire further research into using geometric or parametric priors in other signal processing tasks. GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.
Primary: Institute for AI Industry Research (AIR)
All Institutions: Institute for AI Industry Research (AIR), Tsinghua University, University of Electronic Science and Technology of China
The paper introduces a plug-and-play higher-order action supervision scheme that significantly enhances the robustness and performance of offline RL and IL policies by enforcing temporal consistency through first-order action derivatives. By leveraging Jacobian-vector products and integrating with continuous-time HJB frameworks, the method provides a theoretically grounded and empirically validated solution to control instability, demonstrating substantial gains in low-data regimes and complex high-dimensional tasks.
The paper proposes a "higher-order action supervision" framework that augments standard zeroth-order action prediction (mimicking $a_t$) with first-order action supervision (mimicking $\dot{a}_t$ or $a_{t+1}-a_t$). The core insight is leveraging the time derivative of the policy function $\pi(s)$, specifically $\nabla_s \pi(s) \dot{s}$, to provide a supervision signal for action dynamics without modifying the network architecture. This is implemented via Jacobian-vector products (JVPs), which are computationally efficient. The method is generalized to deterministic, stochastic (via Itô calculus/OU processes), and flow-matching policies. It also integrates with the Hamilton-Jacobi-Bellman (HJB) equation for value function regularization in continuous-time RL. The theoretical contribution provides error bounds showing that supervising first-order dynamics reduces policy error accumulation in closed-loop rollouts.
The experiments are extensive, covering OGBench (high-dimensional, sparse rewards) and D4RL (standard locomotion). The paper demonstrates significant improvements in both performance scores and robustness (reduced variance, smoother action trajectories). Notably, it shows substantial gains in low-data regimes (10k transitions), where standard methods often fail due to OOD extrapolation errors, while the proposed method recovers expert behavior. Ablations confirm the contribution of both the policy-side and value-side first-order terms. Comparisons against smoothness penalties (CAPS, L2C2) and architectural constraints (LipsNet) show the proposed method is more effective and less prone to underfitting or OOD errors.
The paper provides detailed derivations in the appendix and specifies the use of JVPs, which are standard in modern frameworks (PyTorch 2.0, JAX). Hyperparameters for the new loss terms are discussed (often set to 1.0 or swept). The code is not explicitly linked in the text provided, but the method is described with sufficient mathematical precision for implementation. The reliance on standard offline RL baselines (TD3+BC, IQL, FQL) makes replication feasible.
The method relies on finite-difference approximations for state and action derivatives, which can be noisy with coarse sampling. It is primarily evaluated on continuous control tasks; applicability to discrete actions or high-dimensional visual inputs (pixels) is limited without a good state representation. The computational cost increases slightly due to JVP calculations, though it remains lightweight. The theoretical guarantees assume certain smoothness conditions on the policy and environment dynamics.
This work addresses a critical bottleneck in deploying learned policies to real-world robotics and autonomous systems: control instability and jitter. By providing a plug-and-play module that enhances temporal consistency, it has high practical value. It bridges the gap between discrete-time MDP formulations and continuous-time control theory, offering a principled way to incorporate physical priors into neural policy learning. The approach could be widely adopted as a standard regularization term in offline RL pipelines. The paper introduces a plug-and-play higher-order action supervision scheme that significantly enhances the robustness and performance of offline RL and IL policies by enforcing temporal consistency through first-order action derivatives. By leveraging Jacobian-vector products and integrating with continuous-time HJB frameworks, the method provides a theoretically grounded and empirically validated solution to control instability, demonstrating substantial gains in low-data regimes and complex high-dimensional tasks.
Real-time whole-body controllers for legged robots typically plan through a fixed nominal model and degrade when the deployed dynamics change. Adaptive methods typically require a model structure that contact dynamics do not provide, or they need offline training for each anticipated condition. We present Look-back and Look-ahead Adaptive Model Predictive Path Integral control (LLA-MPPI). The method converts whole-body adaptation into selection over a bank of GPU-batched contact simulators with different physical or structural parameters. Windowed prediction errors select the simulator that best explains recent motion. A whole-body MPPI planner optimizes controls through the selected model. The framework requires no offline training, and its selected hypotheses are physically interpretable. Across four simulated tasks, it achieves 97.5% success while the strongest baseline reaches 74% and an oracle with the true model reaches 98.5%. Hardware validation on a Unitree Go2 shows the robot walking under a payload added mid-run, walking after one leg is disabled, and pushing a box to its goal while increasing its mass on the fly. Code, videos, and project details are available at: https://lla-control.github.io
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology
LLA-MPPI introduces a novel adaptive control framework that leverages GPU-batched physics simulation to select the best-fitting dynamics model from a bank of hypotheses, enabling real-time whole-body control of legged robots under significant model mismatch without offline training. The paper demonstrates near-oracle performance in simulation and successful hardware validation, establishing a new standard for adaptive model-based control in contact-rich environments.
The paper proposes LLA-MPPI, a framework that decouples adaptive model identification from trajectory optimization. The core innovation is the use of a "bank" of GPU-batched physics simulators (via MuJoCo Warp) to perform model selection. Instead of using classical estimators (like Kalman filters or particle filters) that require specific linear or affine structures, the method treats adaptation as a classification problem over a discrete set of physical hypotheses. The "Look-back" stage runs on the GPU, evaluating thousands of candidate dynamics models in parallel against recent state transitions to select the best-fitting model. The "Look-ahead" stage runs on the CPU, using MPPI (Model Predictive Path Integral) control to plan trajectories using the selected model. This asynchronous, heterogeneous architecture is technically sound and cleverly leverages modern hardware capabilities (GPU parallelism for wide/shallow simulation, CPU threads for narrow/deep planning). The removal of the need for differentiable or closed-form dynamics is a significant methodological advantage for contact-rich robotics.
The evaluation is rigorous and comprehensive. The authors test the method on four distinct simulated tasks (asymmetric payload, leg lock, slope/friction variation, variable-mass box pushing) and validate on hardware (Unitree Go2). The comparison against an oracle (true model) shows the method achieves 97.5% success vs. 98.5% for the oracle, indicating near-perfect adaptation. Comparisons against baselines (Nominal MPPI, Disturbance Observer, Continuous Parameter Estimator) show substantial improvements in success rate and fall rate. The inclusion of a comparison with Domain Randomized PPO is valuable, highlighting the trade-off between offline training costs and online adaptability. The hardware results are particularly convincing, demonstrating real-world applicability in scenarios like mid-run payload addition and leg failure.
The paper provides a project URL with code, videos, and details. The specific hardware (RTX 2080 Ti, i9-9980XE) and software versions (MuJoCo 3.4.0, MuJoCo-Warp 1.12.0) are clearly stated. The hyperparameters for the model bank size, window size, and switching margins are detailed for each task. This level of detail supports high reproducibility.
The method relies on a finite, design-time hypothesis bank. If the true dynamics fall outside the discretized range of the bank, adaptation will fail. The paper acknowledges this and suggests future work on adaptive bank refinement. Additionally, the method requires significant GPU resources to maintain the bank of simulators, which may not be available on all robotic platforms. The distinguishability of models depends on excitation; if the robot's motion does not excite a particular uncertainty (e.g., friction when not slipping), the model selection may be ambiguous, though the hysteresis mechanism helps mitigate this.
This work has high potential impact in the field of legged robotics and adaptive control. By demonstrating that high-fidelity, contact-rich dynamics can be adapted to in real-time without offline training or complex analytical derivations, it opens the door to more robust and versatile robotic systems. The technique of using GPU-batched simulation for model selection could be extended to other domains involving complex, non-differentiable dynamics, such as soft robotics or multi-agent systems. LLA-MPPI introduces a novel adaptive control framework that leverages GPU-batched physics simulation to select the best-fitting dynamics model from a bank of hypotheses, enabling real-time whole-body control of legged robots under significant model mismatch without offline training. The paper demonstrates near-oracle performance in simulation and successful hardware validation, establishing a new standard for adaptive model-based control in contact-rich environments.
Animal evidence shows that precise voluntary movements arise from rotational neural population dynamics in motor cortex, but their physical effects remain unknown. We developed a robotic analog of biological motor systems with artificial muscles, multimodal sensors, and a neural network controller trained via reinforcement learning. The robotic analog exhibited accurate movements, robustness to damage, and neural population dynamics akin to animals. This task-driven, embodied model illuminates the causal link between neural population dynamics and motor outcomes. We discovered that neural rotations generate oscillatory maneuvers orthogonal to the reaching direction, optimizing trajectory adjustments, which is confirmed by primate neural data. The model also revealed counterintuitive neural energy principles under sensor and motor redundancies, and striking Eureka moments during motor learning, bridging biological and artificial systems. These findings provide new perspectives on how neural dynamics contribute to accurate and flexible movement, inspiring future intelligent robots with animal-like mobility.
Primary: Peking University
All Institutions: Peking University, State Key Laboratory of Transvascular Implantation Devices, The Second Affiliated Hospital of Zhejiang University School of Medicine, IDG/McGovern Institute for Brain Research, Peking-Tsinghua Center for Life Sciences, School of Advanced Manufacturing and Robotics, School of Life Sciences
The paper introduces a robotic analog of biological motor systems to causally link neural rotational dynamics to physical movement outcomes. By combining a biomimetic robot arm with an RNN controller trained via reinforcement learning, the authors demonstrate that rotational neural dynamics drive orthogonal trajectory adjustments, a finding validated in primate data, thereby providing a new embodied framework for decoding the functional role of neural population dynamics in motor control.
The paper proposes a "robotic analog" of biological motor systems, combining a continuum robot arm with liquid crystal elastomer (LCE) artificial muscles, multimodal sensors (vision, IMUs, self-sensing muscles), and a Recurrent Neural Network (RNN) controller trained via Soft Actor-Critic (SAC) reinforcement learning. The core methodological innovation lies in using this embodied, transparent system to causally link neural population dynamics (specifically rotational dynamics) to physical motor outcomes. The authors employ dimensionality reduction techniques like jPCA and CCA to analyze the latent dynamics of the RNN controller, comparing them to primate motor cortex data. The approach is rigorous, moving beyond simple simulation to physical embodiment to test hypotheses about neural dynamics that are difficult to test in vivo.
The experiments are extensive, covering simulation and real-world physical tests. The robot achieves sub-5mm reaching errors and demonstrates robustness to muscle failure without retraining. Key findings include the emergence of rotational dynamics in the RNN layers that drive orthogonal maneuvers, a phenomenon confirmed in primate data. The paper also analyzes neural energy efficiency under different sensory feedback conditions and identifies "Eureka moments" during learning where performance and dynamics shift abruptly. The validation against primate neural data strengthens the biological relevance of the findings.
The paper provides detailed descriptions of the hardware fabrication (LCE fibers, IMU setup) and the simulation environment (MuJoCo, constitutive models). However, specific code repositories or hyperparameters for the RL training are not explicitly provided in the text, which may limit immediate reproducibility. The physical setup is complex and likely expensive to replicate.
The system is limited to a single arm and relatively simple reaching tasks. The "neural" dynamics are those of an artificial RNN, not biological neurons, so while the analogy is strong, it is not a direct proof of biological mechanisms. The "Eureka moment" finding, while interesting, is based on a specific training trajectory and may not generalize to all RL algorithms or tasks.
This work bridges the gap between computational neuroscience and robotics, offering a new paradigm for studying motor control. It provides insights into how neural dynamics contribute to movement precision and flexibility, which could inspire more biologically plausible robot controllers. The findings on rotational dynamics and energy efficiency have implications for both understanding brain function and designing efficient robotic systems. The paper introduces a robotic analog of biological motor systems to causally link neural rotational dynamics to physical movement outcomes. By combining a biomimetic robot arm with an RNN controller trained via reinforcement learning, the authors demonstrate that rotational neural dynamics drive orthogonal trajectory adjustments, a finding validated in primate data, thereby providing a new embodied framework for decoding the functional role of neural population dynamics in motor control.
World action models jointly learn visual predictionand robot actions, providing a way to use observations ofscene evolution for policy learning. Their video and actionlosses, however, provide no explicit target for the geometricconsequences of a demonstrated action sequence. Moreover,visual features taken after temporal attention can contain futureobservations, making them unsuitable as the sole current visualinput to an auxiliary predictor. We introduce ACG-WAMand its auxiliary objective, the Action-Conditioned GeometricJoint-Embedding Predictive Architecture (ACG-JEPA), whichpredicts geometric features at several horizons from the currentobservation and intervening actions, using the future slot of afrozen VGGT encoding of each current and future image pairas the target. We apply this supervision from the head and wristcameras to a shared visual embedding before temporal mixing,and remove the teacher and auxiliary modules at inference.On 50 RoboTwin 2.0 tasks, ACG-WAM achieves 93.46%success in clean scenes, with the best randomized success(92.68%) and mean across both settings (93.07%) among thecompared methods; across three tasks on a real robot, itachieves 85.00% success and 91.67% partial completion score,exceeding Motus by 10.00 and 9.17 percentage points, respec-tively. Code:https://github.com/RoboOpus/ACG-WAM.Website:https://RoboOpus.github.io/ACG-WAM.
Primary: Beijing Institute of Technology
All Institutions: Beijing Institute of Technology, LimX Dynamics
ACG-WAM introduces an action-conditioned geometric prediction objective to World Action Models, significantly improving bimanual manipulation performance in simulation and on real robots. The paper presents a rigorous technical contribution that addresses specific challenges in joint video-action learning, such as information leakage and the lack of explicit geometric targets, resulting in state-of-the-art results on the RoboTwin 2.0 benchmark and real-world tasks.
The paper proposes ACG-WAM, a World Action Model (WAM) that augments the Motus backbone with an auxiliary geometric prediction objective. The core innovation is the Action-Conditioned Geometric Joint-Embedding Predictive Architecture (ACG-JEPA). This module uses a frozen VGGT teacher to encode current and future image pairs, extracting the "future slot" features as targets. The student predictor takes current visual features (crucially, extracted *before* temporal attention to avoid information leakage from future frames) and the intervening action sequence to predict these geometric targets. This design explicitly supervises the geometric consequences of actions, addressing a gap in standard video/action loss functions. The method is well-motivated by the need for spatial reasoning in bimanual manipulation and the technical challenge of preventing future information leakage in auxiliary predictors.
The evaluation is robust, covering 50 tasks on the RoboTwin 2.0 benchmark (clean and randomized settings) and 3 real-world tasks on a TRON2 platform. ACG-WAM achieves state-of-the-art results, outperforming strong baselines like Motus, LingBot-VA, and MECo-WAM. Specifically, it achieves 93.46% success in clean scenes and 92.68% in randomized scenes on RoboTwin 2.0, and 85.00% success on the real robot. Ablation studies convincingly demonstrate the contribution of action conditioning, multi-horizon prediction, and the joint-encoding target design versus simple endpoint subtraction. The real-world results are particularly strong, showing significant gains over the base Motus model.
The paper provides high reproducibility. Code is released on GitHub. Implementation details are thorough, specifying the backbone (Wan2.2-TI2V-5B), teacher model (VGGT), training hyperparameters (learning rates, batch sizes, loss weights), and inference settings. The use of standard benchmarks (RoboTwin 2.0) and clear evaluation metrics (Success Rate, Partial Completion Score) further supports reproducibility.
The method relies on a large, frozen teacher model (VGGT) and a complex backbone (Motus/Wan2.2), which may limit accessibility for groups with lower computational resources. The real-world evaluation is limited to only three tasks, which is a small sample size for generalizing claims about physical robustness. Additionally, the paper acknowledges the use of ChatGPT for drafting, which, while disclosed, is a minor point regarding the originality of the text rather than the science.
This work contributes to the growing field of World Action Models for robotics. By explicitly incorporating geometric supervision, it offers a pathway to improve spatial reasoning in manipulation policies. The technique of supervising features before temporal mixing to avoid leakage is a useful insight for other architectures involving joint video-action prediction. The strong performance on both simulation and real hardware suggests practical utility for industrial and research robotics applications. ACG-WAM introduces an action-conditioned geometric prediction objective to World Action Models, significantly improving bimanual manipulation performance in simulation and on real robots. The paper presents a rigorous technical contribution that addresses specific challenges in joint video-action learning, such as information leakage and the lack of explicit geometric targets, resulting in state-of-the-art results on the RoboTwin 2.0 benchmark and real-world tasks.
New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and optimization. What agents produce depends on the interfaces they are given. These interfaces may expose a machine's raw capabilities or encode the known-good methods for exploiting its hardware features efficiently. We test the performance impact of these two interfaces on Intel GPUs through controlled experiments with three coding models on three kernels. With raw SYCL and execution feedback, Opus 4.8, the strongest of the three models tested, writes a single-GPU GEMM kernel reaching only 57.5% of Intel's tuned oneDNN library. We present SyclKittens, a hardware-aware tile programming model for Intel GPUs that encodes these known-good methods, so its operations place operands in matrix-engine layouts, prefetch through the L1 cache, and communicate over the fabric that links GPU stacks into a node. With SyclKittens under the same agent, task, and feedback budget, the agent reaches 82.1%, showing that encoded methods turn hardware capabilities into performance. Workload-specific policies stay programmable, and engineers and agents jointly refine these schedules in SyclKittens to build a kernel suite that reaches ~96% of oneDNN in geometric mean across GEMM shapes. The suite runs Llama-3.1-8B inference 1.59x faster than torch.compile on one GPU and up to 2.91x faster than a matched multi-GPU decode path on Intel's oneCCL. SyclKittens is open source and available at https://github.com/intel/SyclKittens.
Primary: Intel Corporation
All Institutions: Massachusetts Institute of Technology, Intel Corporation, Stanford University, California Institute of Technology
SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.
The paper introduces SyclKittens, a tile-based programming model for Intel GPUs that abstracts low-level hardware details (operand pipelines, data movement, matrix engine layouts) into high-level, hardware-aware operations. The core methodological contribution is the demonstration that providing coding agents with a structured, domain-specific interface (SyclKittens) rather than raw SYCL primitives significantly improves the performance of generated kernels. The authors employ a controlled experimental design comparing raw SYCL with execution feedback against SyclKittens under identical agent, task, and feedback budgets. This approach effectively isolates the impact of the programming interface on agent performance, providing a rigorous evaluation of how "known-good" hardware methods encoded in a DSL can guide LLMs to produce efficient code.
The experiments are conducted on Intel Max GPUs, focusing on critical AI workloads such as GEMM, attention, and normalization. The results show a substantial improvement: an agent using raw SYCL reaches only 57.5% of the performance of Intel's tuned oneDNN library, whereas the same agent using SyclKittens reaches 82.1%. Furthermore, a co-designed kernel suite using SyclKittens achieves ~96% of oneDNN performance in geometric mean across GEMM shapes. End-to-end inference benchmarks for Llama-3.1-8B demonstrate 1.59x speedup over torch.compile on a single GPU and up to 2.91x speedup over a matched multi-GPU decode path using Intel's oneCCL. The evaluation is strong in its direct comparison of interfaces and its end-to-end application metrics.
The paper is highly reproducible given that SyclKittens is open-sourced on GitHub. The authors provide detailed appendices separating the controlled-agent studies from the co-designed suite, including specific measurement protocols, warmup procedures, and aggregation methods. The use of specific coding models (Opus 4.8, etc.) and defined feedback budgets allows for precise replication of the agent experiments. However, the specific prompts and agent configurations may require careful extraction from the appendices to fully replicate the agent's behavior.
The evaluation is limited to Intel GPUs, which may not generalize directly to other hardware architectures (e.g., NVIDIA, AMD) without significant adaptation. The performance gains are relative to Intel's own oneDNN library, so the absolute competitive standing against other state-of-the-art libraries on different hardware is not assessed. The reliance on specific coding models means the results may vary with future model improvements or different agent architectures. Additionally, the paper focuses on a limited set of kernels (GEMM, attention, norms), and the applicability to more complex or irregular workloads is not fully explored.
This work has significant implications for the intersection of AI and systems programming. By demonstrating that high-level, hardware-aware abstractions can enable coding agents to write near-optimal GPU kernels, it suggests a new paradigm for kernel development that reduces the need for deep, architecture-specific expertise. This could accelerate the adoption of new AI accelerators by lowering the barrier to entry for kernel optimization. The open-source release of SyclKittens provides a valuable tool for the community to explore similar approaches on Intel hardware and potentially adapt the concepts to other platforms. SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.