Last 7 Days (October 02 – October 08, 2026)
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Primary: ETH Zurich (Inferred from "eth-lre" in GitHub URL and Swiss AI Initiative funding)
All Institutions: ETH Zurich, Swiss National Supercomputing Centre (CSCS)
The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
The paper addresses a critical flaw in existing RL-based tutoring methods: the "telling" problem, where models learn to simply provide answers to maximize immediate student success metrics. The authors propose a novel reward structure based on three conditions: (1) Near-transfer testing (student solves a variant problem, not the exact one), (2) Masked post-test (tutor's utterances are hidden during the test, forcing reliance on student-generated notes/reasoning), and (3) Binary reward gates (strict penalties for factual errors or solution handover). This design effectively decouples "helpfulness" from "teaching," forcing the model to elicit student reasoning rather than provide it. The use of a frozen LLM student with genuine errors (rather than prompted confusion) is a strong methodological choice that ensures the training signal reflects real diagnostic and adaptive challenges.
The experimental setup is rigorous and comprehensive. The authors train models across multiple sizes (4B to 27B) and architectures (Qwen3 variants), demonstrating the robustness of the recipe. The ablation study is particularly strong, isolating the contribution of each condition (transfer, masking, gates) and showing that the gates specifically reduce handover while the masked transfer test improves out-of-domain generalization. The evaluation uses independent benchmarks (MathTutorBench, TutorMoments) and judges (Gemini) distinct from the training setup, mitigating overfitting concerns. The results show that Eduardo-27B matches or exceeds frontier models (Gemini 3.1 Pro, Claude Opus 4.8) in pedagogical quality while using significantly fewer thinking tokens, a crucial efficiency gain for interactive applications.
High. The authors open-source the training environment, the 8,671-problem near-transfer dataset, and the trained models. Detailed hyperparameters, prompts, and compute costs are provided in the appendices. The use of standard RL algorithms (DPPO/GRPO) and open-source base models further enhances reproducibility.
The primary limitation is the reliance on a single frozen LLM student (Llama-3.1-8B-Instruct) for training, which may lead to overfitting to that specific model's failure modes. The paper acknowledges this and suggests future work with diverse student ensembles. Additionally, the evaluation is limited to mathematics, and the "affective gap" (lack of socio-emotional support) is noted as a consequence of optimizing purely for cognitive gain. The ablation study is limited to one seed and one model size (4B), which slightly weakens the statistical confidence in the ablation results.
This work has significant implications for the development of AI tutors and, more broadly, for any domain where an agent must build user capabilities rather than just provide answers. The "masked near-transfer" reward design is a generalizable principle that could be applied to other educational or collaborative tasks. The efficiency gains (fewer thinking tokens) make high-quality tutoring more feasible in real-time interactive settings. The open-sourcing of the dataset and models will likely accelerate research in AI for Education. The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Primary: University at Buffalo (inferred from Empire AI Consortium and NSF grants)
All Institutions: University at Buffalo, Empire AI Consortium, Inc, Simons Foundation, Secunda Family Foundation, Modal Labs
LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
The paper proposes LayerRoPE, a method that reinterprets the growth of hidden state norms in deep Transformers not as a pathology to be suppressed, but as an emergent depth-positional encoding. The core insight is that the RMSNorm weights ($\gamma$) across layers exhibit systematic magnitude growth and directional rotation, effectively encoding the layer index. LayerRoPE makes this explicit by replacing the $L$ independent per-layer $\gamma$ vectors with a single shared vector modulated by depth-conditioned scalars (magnitude and rotation) derived from a RoPE-like mechanism along the depth axis. This reduces parameter count and introduces a structured, learnable depth schedule. The methodology is sound, leveraging the analogy between positional encoding in sequence space and depth encoding in network space. The ablation studies confirm that both magnitude and rotation components are necessary and complementary.
The experimental evaluation is extensive and rigorous. The authors train a model ladder from 58M to 1.3B parameters on C4, comparing LayerRoPE against Pre-Norm, Post-Norm, Peri-Norm, and Layer-Norm Scaling. LayerRoPE consistently outperforms baselines, reaching Pre-Norm's 1.3B loss with 3.4x less compute. Crucially, the paper demonstrates superior depth scaling, with LayerRoPE being the only method to show monotonic loss improvement up to 512 layers, whereas baselines diverge or plateau. The paper also shows improved learning rate robustness (wider basin, delayed divergence) and successful transfer to looped latent models (Parcae) and Vision Transformers (ViT) without tuning. The analysis of the learned depth schedule reveals that LayerRoPE widens the residual stream rather than damping it, which is a counter-intuitive and significant finding.
The paper provides detailed hyperparameters, training recipes, and compute accounting. It specifies the initialization of depth slopes, rotation base frequency, and optimizer settings. The use of standard datasets (C4, ImageNet) and architectures (LLaMA-style, ViT) enhances reproducibility. However, the specific code implementation is not linked in the text provided, which is a minor gap for immediate reproduction, though the description is sufficiently detailed for re-implementation.
The primary limitation is the scale of the experiments. While 1.3B parameters and 512 layers are significant, the method's behavior at frontier scales (100B+) is extrapolated or tested at reduced budgets (6.7B). The paper relies on C4 for language modeling, which may not fully capture the benefits on more diverse or multilingual data. Additionally, the "widening" of the residual stream requires careful monitoring to ensure it doesn't lead to numerical instability in mixed-precision training, although the paper notes stability improvements.
This paper has high potential impact by challenging a fundamental assumption in Transformer architecture design: that activation norm growth is a defect. By providing a principled way to exploit this growth for depth encoding, it offers a path to training deeper, more stable, and more compute-efficient models. The transferability to looped models and ViTs suggests broad applicability across modalities. The reduction in learning rate sensitivity is particularly valuable for practitioners, as it simplifies hyperparameter tuning for deep networks. LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology
LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
The paper proposes LASER, a framework for offline RL that combines support-constrained latent space policies with explicit entropy regularization. The core methodological contribution is the use of "adjoint matching" to optimize the entropy-regularized policy objective in latent space without backpropagation through time (BPTT). This is a significant technical advancement over previous flow-based offline RL methods (like ReFORM or DSRL) which either relied on implicit regularization or computationally expensive BPTT. The derivation of the optimal Q-tilted distribution and its implementation via stochastic optimal control (SOC) is mathematically sound and provides a clear path for stable training. The separation of support constraints (handled by the flow decoder) from policy extraction stability (handled by entropy regularization) is a well-motivated architectural choice.
The experiments are conducted on the OGBench benchmark, covering 40 tasks across locomotion and manipulation. The results show that LASER achieves state-of-the-art performance, particularly on noisy datasets where support constraints are critical. A key strength is the demonstration of robustness: LASER uses a fixed set of hyperparameters across all tasks, outperforming baselines that require task-specific tuning. The ablation studies effectively isolate the contribution of entropy regularization and the efficiency of adjoint matching vs. BPTT (showing a 3.6x speedup). The comparison against recent flow-based baselines (IFQL, FQL, QAM) is appropriate and rigorous.
The paper provides a clear algorithmic summary and references a project website. The use of standard benchmarks (OGBench) and fixed hyperparameters enhances reproducibility. However, the specific implementation details of the adjoint matching solver and the flow matching training are deferred to the appendix, which is standard but requires careful reading. The code availability is implied by the project URL.
The method relies on the quality of the behavior cloning (BC) decoder to enforce support constraints; if the decoder is imperfect, OOD actions can still occur. The computational cost of solving ODEs during training remains high, though mitigated by adjoint matching. The paper acknowledges that the value function learning is relatively simple compared to more advanced pessimistic methods, suggesting room for improvement in the critic component.
This work advances the state of the art in safe offline RL, which is critical for deploying RL agents in real-world robotic systems where online interaction is costly or dangerous. The technique of using adjoint matching for entropy-regularized flow policies could be adapted to other generative model-based RL settings, potentially influencing how expressive policies are trained in broader RL research. LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
Primary: Columbia University
All Institutions: Columbia University
The paper introduces a decoupled exploration-distillation framework for RLVR that effectively boosts reasoning performance and diversity without degrading general capabilities. By separating the aggressive exploration policy from the stable optimization policy, it resolves the tension between novelty and correctness, offering a robust and scalable improvement over standard RLVR methods.
The paper proposes Exploration-Distillation (ExpDis), a framework that decouples exploration from optimization in Reinforcement Learning with Verifiable Rewards (RLVR). The core insight is that applying novelty bonuses directly to the policy being optimized leads to degradation of general capabilities because the verifiable reward only supervises a narrow slice of the model's knowledge. By training separate "explorer" policies with aggressive novelty incentives (using Random Network Distillation) and distilling only their correct, high-quality trajectories into a "student" policy via SFT followed by standard RLVR, the method isolates the benefits of exploration from the risks of instability. The methodology is sound, building on established concepts like intrinsic motivation and expert iteration, but the specific application to LLM reasoning with the decoupled architecture is a novel and effective contribution.
The experiments are rigorous and well-controlled. The authors compare against strong baselines (DAPO, GRPO, Dr. GRPO) and a direct novelty-bonus variant of DAPO. They demonstrate that ExpDis outperforms DAPO at the same wall-clock budget across multiple model families (Qwen3, Ministral) and benchmarks (AIME, MATH, etc.). Crucially, they show that while direct novelty bonuses degrade general knowledge benchmarks (MMLU, GPQA), ExpDis preserves or improves these capabilities while boosting reasoning performance. The analysis of pass@$k$ scaling and diversity metrics provides strong evidence that the method genuinely expands the model's solution space rather than just sharpening existing modes.
The paper provides high reproducibility. Code and checkpoints are released on GitHub and HuggingFace. Hyperparameters, including the novelty bonus weight schedule and filtering criteria, are detailed in the appendix. The use of standard datasets (DAPO-Math-17K) and public benchmarks further aids reproducibility.
The method requires training multiple models (explorers + student), which increases memory and compute overhead compared to single-model RLVR, although the paper argues this is offset by the ability to use aggressive exploration. The current evaluation is limited to mathematical reasoning; it is unclear if the benefits generalize to other domains like code or open-ended generation without similar verifiable rewards. The optimal number of explorers and rounds may be sensitive to the specific task and model scale.
This work has significant implications for the scaling of LLM reasoning. As models become proficient at known strategies, the need for discovering new strategies grows. ExpDis provides a safe mechanism to do so, potentially enabling the discovery of novel reasoning paths that standard RLVR would miss due to entropy collapse. It offers a practical path for practitioners to enhance model diversity and capability without the risk of catastrophic forgetting or quality degradation. The paper introduces a decoupled exploration-distillation framework for RLVR that effectively boosts reasoning performance and diversity without degrading general capabilities. By separating the aggressive exploration policy from the stable optimization policy, it resolves the tension between novelty and correctness, offering a robust and scalable improvement over standard RLVR methods.
Large neural networks often acquire capabilities that small models fail to learn. Does this stem from large models learning more representative features, or from being more robust to unaccounted-for adverse effects introduced during training? We define and quantify one such adverse effect, primacy bias, as the extent to which exposure to early data distributions impairs later learning. We show that small models can allocate learning capacity inefficiently toward early distributions, whereas sufficiently overparameterized models are robust to this effect. This inefficiency is particularly consequential in pretraining, where foundation models often encounter heterogeneous data distributions sequentially rather than jointly. As a result, small foundation models can struggle to learn distributions encountered late in training, which is particularly harmful when later data emphasizes desirable capabilities such as code, mathematics, and reasoning. Motivated by these findings, we introduce Exposure Therapy (ET), a simple regularization that promotes more efficient allocation of learning capacity during sequential pretraining. We demonstrate that ET improves foundation models' performance on late data distributions as well as overall capability in models up to the billion-parameter scale. Overall, our results suggest that some benefits of large foundation models may arise from greater robustness to adverse training effects, rather than from learning more representative features, and that improved training algorithms can recover some of these advantages in smaller models.
Primary: University of Illinois Urbana-Champaign
All Institutions: Purdue University, University of Illinois Urbana-Champaign
The paper identifies and quantifies "primacy bias" in neural network training, demonstrating that overparameterization provides robustness to adverse effects of sequential data exposure, and introduces "Exposure Therapy" as a simple, effective regularization technique to recover this robustness in smaller models, thereby offering a practical algorithmic alternative to scaling for improving late-stage learning in foundation models.
The paper introduces a well-defined metric, "primacy sensitivity," to quantify the impact of early data distribution on later learning performance. The core methodological contribution is "Exposure Therapy" (ET), a simple regularization technique that sparsely replaces training updates with samples from later-stage data distributions during the early, high-plasticity phase of training. The approach is theoretically grounded in the concept of capacity allocation and loss of plasticity, distinguishing itself from standard replay methods by focusing on preserving capacity for *new* distributions rather than retaining *old* ones. The methodology is clean, computationally efficient (no extra steps or data), and directly addresses a practical pain point in sequential pretraining.
The experimental evaluation is rigorous and multi-faceted. It begins with controlled MLP experiments on vision tasks (MNIST, Fashion-MNIST, KMNIST) to isolate the scale-dependent nature of primacy bias, providing mechanistic insights via neuron overlap analysis. It then scales up to causal Transformer pretraining (100M to 1B parameters) using diverse domains (code, math, German, Finnish, English). The results consistently show that ET mitigates the performance gap between small and large models in sequential settings. Ablation studies on reverse-order pretraining and varying mixture ratios further validate the robustness of the findings. The scale (up to 1B parameters) is appropriate for demonstrating the trend, though not frontier-scale.
The paper provides a reproducibility statement, indicating that code for toy experiments, foundation model experiments, and ET will be released. Detailed hyperparameters, architecture specifications, and training schedules are provided in the appendices. The use of public datasets (CodeSearchNet, OpenWebMath, Wikipedia, FineWeb, TinyStories) enhances reproducibility. The clear definition of the ET algorithm (Algorithm 1) allows for straightforward implementation.
The primary limitation is the scale of the foundation model experiments (max 1B parameters), which may not fully capture dynamics at frontier scales (100B+). The mechanistic analysis is limited to MLPs, and the authors acknowledge that the precise reason for parameter avoidance in large models remains an open question. The pretraining setup uses a simplified two-phase structure, whereas real-world pretraining often involves more complex, multi-stage curricula. Hyperparameter tuning for ET is not rigorously explored for large-scale models, relying instead on qualitative trends from toy models.
This work has significant implications for the efficiency of foundation model training. By demonstrating that some benefits of scale can be recovered through algorithmic improvements (ET), it offers a path to training smaller, more efficient models that retain capabilities typically associated with larger models. This is particularly relevant for resource-constrained settings and for understanding the interplay between model capacity, data order, and learning dynamics. The concept of "primacy bias" provides a new lens for analyzing curriculum learning and data mixing strategies. The paper identifies and quantifies "primacy bias" in neural network training, demonstrating that overparameterization provides robustness to adverse effects of sequential data exposure, and introduces "Exposure Therapy" as a simple, effective regularization technique to recover this robustness in smaller models, thereby offering a practical algorithmic alternative to scaling for improving late-stage learning in foundation models.
Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable. In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples during training. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data. To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation. This work fills a critical gap between traditional statistical multiview learning and the growing field of unpaired data learning.
Primary: Bar-Ilan University
All Institutions: Bar-Ilan University
The paper presents a novel and theoretically grounded method, Unpaired CCA, that enables the recovery of maximally correlated linear projections from strictly unpaired data by leveraging a connection to the Quadratic Assignment Problem. The rigorous theoretical analysis, combined with strong empirical results across diverse datasets demonstrating superior performance over existing unpaired alignment baselines, establishes this as a significant contribution to the field of multiview learning and representation learning.
The paper introduces Unpaired CCA (UCCA), a method to learn linear projections maximizing correlation between two views without paired samples. The core theoretical contribution is establishing an equivalence between the optimal unpaired pairing problem (relaxed to the Frobenius sphere) and the Quadratic Assignment Problem (QAP). Specifically, it proves that maximizing the squared Frobenius norm of the permuted cross-correlation matrix is equivalent to the Koopmans-Beckmann QAP formulation using Gram matrices. This allows the use of efficient approximate QAP solvers (like 2-Opt) to find a proxy permutation. The algorithm then uses K-Means centroids as anchors to reduce the QAP size, solves for the anchor matching, and applies standard CCA to the resulting pseudo-pairs. The theoretical link between the orthogonal group relaxation and the Frobenius sphere relaxation is well-established under mild majorization assumptions, which are empirically validated.
The experiments are rigorous and well-designed. The authors evaluate on four diverse real-world datasets (SNARE, Flickr, COCO, Handwritten) with strictly unpaired training splits. They compare against relevant baselines like SCA, UCA, and adapted versions of J-MDS/SCOT. The primary metric, Total Correlation (TC) of the true underlying pairing, clearly shows UCCA outperforming all baselines and approaching the paired CCA upper bound. Cross-view classification results further validate the utility of the learned representations. Ablations on anchor selection, QAP solvers, and data imbalance provide a comprehensive understanding of the method's robustness. The validation of the theoretical equivalence (Kendall Tau distances) is a strong empirical backing for the theoretical claims.
The paper provides a GitHub link with code. The methodology is clearly described, including the specific QAP solvers used (2-Opt vs FAQ) and the anchor extraction strategy (K-Means). Hyperparameters and implementation details are deferred to the appendix but are standard for this type of work. The use of standard datasets and clear evaluation protocols enhances reproducibility.
The method is currently limited to the bi-view setting; extension to multi-view is left for future work. There is a performance gap compared to the paired CCA upper bound, though it is small. The reliance on QAP solvers, while efficient, introduces a dependency on the quality of the approximation, though ablations show robustness to solver choice. The assumption that the data can be represented by a permutation of samples (even if unknown) is a strong structural assumption that may not hold for all generative processes.
This work bridges a significant gap between classical statistical multiview learning and modern unpaired data learning. It provides a principled, non-generative approach to aligning unpaired multimodal data, which is highly relevant given the scarcity of paired data in many domains (e.g., medical imaging, scientific data). The method could be widely adopted as a preprocessing step or a standalone tool for multimodal representation learning, potentially enabling new applications where paired data is prohibitively expensive or impossible to obtain. The paper presents a novel and theoretically grounded method, Unpaired CCA, that enables the recovery of maximally correlated linear projections from strictly unpaired data by leveraging a connection to the Quadratic Assignment Problem. The rigorous theoretical analysis, combined with strong empirical results across diverse datasets demonstrating superior performance over existing unpaired alignment baselines, establishes this as a significant contribution to the field of multiview learning and representation learning.
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
Primary: University at Buffalo (inferred from Empire AI Consortium and NSF grants)
All Institutions: University at Buffalo, Empire AI Consortium, Inc, Simons Foundation, Secunda Family Foundation, Modal Labs
LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
The paper proposes LayerRoPE, a method that reinterprets the growth of hidden state norms in deep Transformers not as a pathology to be suppressed, but as an emergent depth-positional encoding. The core insight is that the RMSNorm weights ($\gamma$) across layers exhibit systematic magnitude growth and directional rotation, effectively encoding the layer index. LayerRoPE makes this explicit by replacing the $L$ independent per-layer $\gamma$ vectors with a single shared vector modulated by depth-conditioned scalars (magnitude and rotation) derived from a RoPE-like mechanism along the depth axis. This reduces parameter count and introduces a structured, learnable depth schedule. The methodology is sound, leveraging the analogy between positional encoding in sequence space and depth encoding in network space. The ablation studies confirm that both magnitude and rotation components are necessary and complementary.
The experimental evaluation is extensive and rigorous. The authors train a model ladder from 58M to 1.3B parameters on C4, comparing LayerRoPE against Pre-Norm, Post-Norm, Peri-Norm, and Layer-Norm Scaling. LayerRoPE consistently outperforms baselines, reaching Pre-Norm's 1.3B loss with 3.4x less compute. Crucially, the paper demonstrates superior depth scaling, with LayerRoPE being the only method to show monotonic loss improvement up to 512 layers, whereas baselines diverge or plateau. The paper also shows improved learning rate robustness (wider basin, delayed divergence) and successful transfer to looped latent models (Parcae) and Vision Transformers (ViT) without tuning. The analysis of the learned depth schedule reveals that LayerRoPE widens the residual stream rather than damping it, which is a counter-intuitive and significant finding.
The paper provides detailed hyperparameters, training recipes, and compute accounting. It specifies the initialization of depth slopes, rotation base frequency, and optimizer settings. The use of standard datasets (C4, ImageNet) and architectures (LLaMA-style, ViT) enhances reproducibility. However, the specific code implementation is not linked in the text provided, which is a minor gap for immediate reproduction, though the description is sufficiently detailed for re-implementation.
The primary limitation is the scale of the experiments. While 1.3B parameters and 512 layers are significant, the method's behavior at frontier scales (100B+) is extrapolated or tested at reduced budgets (6.7B). The paper relies on C4 for language modeling, which may not fully capture the benefits on more diverse or multilingual data. Additionally, the "widening" of the residual stream requires careful monitoring to ensure it doesn't lead to numerical instability in mixed-precision training, although the paper notes stability improvements.
This paper has high potential impact by challenging a fundamental assumption in Transformer architecture design: that activation norm growth is a defect. By providing a principled way to exploit this growth for depth encoding, it offers a path to training deeper, more stable, and more compute-efficient models. The transferability to looped models and ViTs suggests broad applicability across modalities. The reduction in learning rate sensitivity is particularly valuable for practitioners, as it simplifies hyperparameter tuning for deep networks. LayerRoPE reinterprets residual stream norm growth as an emergent depth-positional encoding, replacing per-layer normalization weights with a shared, depth-conditioned RoPE-like mechanism that improves depth scaling, learning rate robustness, and compute efficiency across LLMs and ViTs. The paper provides a compelling theoretical and empirical case for widening rather than damping the residual stream, demonstrating state-of-the-art results in depth scaling up to 512 layers and significant compute savings in model training.
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology
LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
The paper proposes LASER, a framework for offline RL that combines support-constrained latent space policies with explicit entropy regularization. The core methodological contribution is the use of "adjoint matching" to optimize the entropy-regularized policy objective in latent space without backpropagation through time (BPTT). This is a significant technical advancement over previous flow-based offline RL methods (like ReFORM or DSRL) which either relied on implicit regularization or computationally expensive BPTT. The derivation of the optimal Q-tilted distribution and its implementation via stochastic optimal control (SOC) is mathematically sound and provides a clear path for stable training. The separation of support constraints (handled by the flow decoder) from policy extraction stability (handled by entropy regularization) is a well-motivated architectural choice.
The experiments are conducted on the OGBench benchmark, covering 40 tasks across locomotion and manipulation. The results show that LASER achieves state-of-the-art performance, particularly on noisy datasets where support constraints are critical. A key strength is the demonstration of robustness: LASER uses a fixed set of hyperparameters across all tasks, outperforming baselines that require task-specific tuning. The ablation studies effectively isolate the contribution of entropy regularization and the efficiency of adjoint matching vs. BPTT (showing a 3.6x speedup). The comparison against recent flow-based baselines (IFQL, FQL, QAM) is appropriate and rigorous.
The paper provides a clear algorithmic summary and references a project website. The use of standard benchmarks (OGBench) and fixed hyperparameters enhances reproducibility. However, the specific implementation details of the adjoint matching solver and the flow matching training are deferred to the appendix, which is standard but requires careful reading. The code availability is implied by the project URL.
The method relies on the quality of the behavior cloning (BC) decoder to enforce support constraints; if the decoder is imperfect, OOD actions can still occur. The computational cost of solving ODEs during training remains high, though mitigated by adjoint matching. The paper acknowledges that the value function learning is relatively simple compared to more advanced pessimistic methods, suggesting room for improvement in the critic component.
This work advances the state of the art in safe offline RL, which is critical for deploying RL agents in real-world robotic systems where online interaction is costly or dangerous. The technique of using adjoint matching for entropy-regularized flow policies could be adapted to other generative model-based RL settings, potentially influencing how expressive policies are trained in broader RL research. LASER introduces a robust offline RL algorithm that uses latent-space adjoint matching to achieve entropy-regularized policy optimization with flow-based policies, effectively decoupling support constraints from policy extraction stability. The paper makes a significant contribution to the field by identifying entropy regularization as a critical, previously under-exploited component in support-constrained latent-space RL, and by providing an efficient, BPTT-free optimization method that yields state-of-the-art results on the OGBench benchmark with superior hyperparameter robustness.
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.
Primary: Tsinghua University
All Institutions: Tsinghua University
The paper presents a rigorous and effective framework for evolving RL algorithms using LLMs, demonstrating significant performance gains and discovering interpretable, transferable mechanisms. Its combination of structured search (PCE) and efficient evaluation (PPE) addresses key bottlenecks in program evolution for RL, offering a practical and impactful contribution to the field of automated machine learning.
The paper introduces RLDiscover, a framework for the self-evolution of model-free deep RL algorithms using LLMs. The core methodological contribution is the combination of Progressive Co-Evolution (PCE) and Progressive Probabilistic Evaluation (PPE). PCE addresses the difficulty of joint search over coupled algorithmic components by first using a UCB-inspired bandit to identify high-value components for isolated mutation, then transitioning to joint co-evolution. PPE addresses the high cost and noise of RL evaluation by implementing a multi-fidelity funnel (L0: 50k steps, L1: 200k steps, L2: 1M steps) that allocates compute efficiently. The decomposition of RL algorithms into five specific components (state encoding, value target, critic loss, policy loss, exploration bonus) is a structured and logical choice that enables targeted LLM mutations. The methodology is sound, leveraging established concepts (bandits, multi-fidelity optimization) in a novel application context.
The experiments are extensive, covering three major RL families (SAC, PPO, DQN) across 26 task-algorithm pairs from four benchmark suites (DMControl, Gymnasium, Meta-World, MetaDrive). The results show substantial improvements, with median gains of 32-84% over baselines. A particularly strong aspect is the "SOTA-init" experiment, where evolution starts from strong open-source implementations (CleanRL, denisyarats/pytorch_sac) and still yields improvements, validating that the gains are not merely due to fixing a weak baseline. The ablation studies are rigorous, comparing the full method against variants removing PCE or PPE components, and analyzing the trade-off between search breadth and evaluation fidelity. The cross-task transfer analysis provides evidence that the evolved programs contain generalizable mechanisms rather than just task-specific hacks.
The paper provides detailed descriptions of the component interfaces, the PCE algorithm, and the PPE promotion rules. It explicitly states that code and data will be released upon acceptance. The "AI use statement" clarifies that the LLM is the mutation operator, not just a writing aid, and details the audit process for generated programs. The reproducibility statement is thorough, specifying training budgets, seeds, and evaluation protocols. The inclusion of per-seed results and the discussion of variance (e.g., the bimodal outcome of MountainCar) enhances transparency.
The authors acknowledge that PPO and DQN results are suboptimal due to budget constraints carried over from SAC. The five-component decomposition is only active at the L0 screening stage, limiting structured exploration at higher fidelities. L0 screening at 50k steps may discard strategies with slow warm-up. The LLM identity is withheld for anonymity, which is standard for submission but limits immediate reproducibility of the exact mutation behavior until publication.
This paper contributes to the growing field of LLM-driven algorithm discovery. By demonstrating that LLMs can evolve interpretable, reusable RL mechanisms (like robust critic losses and adaptive Q-blending) that transfer across tasks, it suggests a path toward automated RL algorithm design. The findings on "convergent discovery" of specific mechanisms provide insights into what makes RL algorithms robust, potentially guiding human designers. The efficiency gains from PPE (15x reduction in compute) make this approach more practical for broader adoption. The paper presents a rigorous and effective framework for evolving RL algorithms using LLMs, demonstrating significant performance gains and discovering interpretable, transferable mechanisms. Its combination of structured search (PCE) and efficient evaluation (PPE) addresses key bottlenecks in program evolution for RL, offering a practical and impactful contribution to the field of automated machine learning.
Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.
Primary: Technical University of Denmark (DTU)
All Institutions: Technical University of Denmark (DTU), Novo Nordisk Foundation
The paper introduces a novel framework for molecular discovery that leverages information-dense stochastic synthesis to achieve exponential speedups in finding rare active molecules. By rigorously combining Bayesian experimental design with combinatorial chemistry, it provides a theoretically grounded and empirically promising solution to the sparsity problem in lab-in-the-loop systems, though its practical impact awaits wet-lab validation.
The paper proposes "Lab-in-the-loop learning with Information-Dense Synthesis" (LIDS), a framework that shifts molecular discovery from testing individual molecules to testing complex mixtures (pools). The core innovation is the application of Bayesian experimental design, specifically maximizing Expected Information Gain (EIG), to the design of these mixtures. The authors provide a rigorous theoretical foundation, proving that under specific conditions (e.g., discrete activity, no noise), this approach can reduce the number of experiments from $O(d)$ to $O(1)$ or $O(\log d)$. They introduce a tractable sparse function class (a Bayesian neural network with exponential nonlinearity) to model the molecule-activity map, which allows for closed-form computation of the inner product between the activity map and the mixture distribution. The method uses variational inference to approximate the posterior and optimize the synthesis parameters. The theoretical contribution is strong, particularly the proof that standard acquisition functions like UCB and EI fail to exploit stochastic synthesis, whereas EIG does.
The experiments are conducted entirely in simulation. The authors evaluate LIDS on two types of oracles: (1) a synthetic sparse DNA sequence oracle and (2) a protein language model (ProGen2) based oracle for peptide design. The results show significant speedups, with LIDS finding optimal molecules in 2-5 experiments compared to 100+ for baselines like Thompson Sampling and Gaussian Process Bayesian Optimization. The ablation studies effectively demonstrate that the gains come from the stochastic synthesis design rather than the specific surrogate model. However, the lack of wet-lab validation is a significant weakness, as the assumptions of linearity in the assay and the ability to synthesize arbitrary mixtures with high fidelity are not tested in a real-world setting.
The paper provides detailed algorithmic descriptions and hyperparameters for the simulations. It specifies the use of NumPyro for probabilistic programming and provides details on the variational distribution architecture and training procedures. However, the code is not explicitly linked in the provided text, and the reliance on specific protein language models and synthetic oracles makes direct reproduction of the exact results dependent on access to these models and the specific simulation setup.
The primary limitation is the absence of experimental validation in a physical laboratory. The method assumes that mixtures can be synthesized and tested with high fidelity and that the assay response is linear with respect to the mixture composition, which may not hold for all biological assays. Additionally, the computational cost of LIDS is significantly higher than individual synthesis methods, requiring substantial GPU resources for EIG optimization and posterior updates. The method also relies on known noise levels, which may not always be available in practice.
If validated in the wet lab, this approach could revolutionize automated molecular discovery by drastically reducing the number of experiments needed to find rare active molecules. It offers a new paradigm for scaling laboratory feedback by increasing the information density of each experiment rather than just the number of experiments. This could have significant implications for drug discovery, materials science, and other fields where searching large combinatorial spaces is a bottleneck. The paper introduces a novel framework for molecular discovery that leverages information-dense stochastic synthesis to achieve exponential speedups in finding rare active molecules. By rigorously combining Bayesian experimental design with combinatorial chemistry, it provides a theoretically grounded and empirically promising solution to the sparsity problem in lab-in-the-loop systems, though its practical impact awaits wet-lab validation.
A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents' private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent's contribution scales with its squared centrality in a network reweighted by agents' preferences across activities. This decomposition motivates Poll, a polling algorithm in which the planner samples one agent per round, walks briefly through the agent's neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, Poll uses significantly fewer operations than exact computation and other distributed methods, requiring up to three orders of magnitude less communication on a real-world network with over 300,000 agents; statistically, its query complexity scales with topology and preference heterogeneity rather than explicitly with population size; and economically, it converges to welfare-maximizing prices while admitting behavior-specific implementations that induce truthful reports and detect adversarial deviations.
Primary: MIT
All Institutions: MIT
The paper introduces a scalable polling algorithm for network intervention that leverages centrality-based welfare decomposition to achieve computational, statistical, and economic efficiency. By combining stochastic approximation with mechanism design techniques like decoy signals, it offers a robust solution to the entangled challenges of learning, strategy, and computation in large-scale networked systems, representing a significant advance in algorithmic game theory and distributed optimization.
The paper proposes a novel framework for network intervention where a planner sets non-discriminatory prices in a multi-activity network game with heterogeneous private technologies. The core theoretical contribution is a centrality-based decomposition of the welfare kernel, showing that optimal prices depend on agents' squared centralities in a network reweighted by their preferences. This insight motivates "Poll," a stochastic approximation algorithm that samples one agent per round, performs local random walks, and updates prices based on local reports. The methodology is rigorous, combining game theory, network science, and stochastic optimization. The introduction of "decoy signals" to induce truthful reporting via Knightian uncertainty is a particularly creative and theoretically sound mechanism design contribution that avoids the need for monetary payments.
The experiments are extensive, covering synthetic networks (hierarchical, ER, regular) and a large real-world dataset (DBLP coauthorship network with ~317k agents). The paper demonstrates significant computational efficiency, claiming up to three orders of magnitude less communication than distributed baselines like Federated Learning and Consensus. It also validates the learning efficiency by showing query complexity scales with heterogeneity rather than population size, and the economic efficiency by demonstrating robustness to strategic manipulation under various ambiguity attitudes. The results are convincing and align well with the theoretical predictions.
The paper provides detailed descriptions of the algorithm, the network models, and the experimental setup. However, no code repository or specific implementation details (e.g., hyperparameter sweep ranges, exact random seeds) are provided in the text. The reliance on specific graph structures and preference distributions makes full reproduction dependent on the authors' released code, which is not linked.
The model assumes linear-quadratic utilities, which may not capture all real-world strategic behaviors. The incentive guarantee relies on agents being ambiguity-averse (max-min preferences), which is a strong behavioral assumption. The "decoy" mechanism requires the planner to generate and send multiple signals, which adds communication overhead, though still less than full aggregation. The analysis assumes the planner does not know the network topology, but the algorithm's performance depends on the network's spectral properties, which might be hard to estimate in dynamic networks.
This work has significant implications for platform design, public policy (subsidies, tariffs), and distributed AI systems. It provides a scalable, privacy-preserving, and incentive-compatible method for coordinating large-scale networks of strategic agents. The insights into centrality and preference heterogeneity could inform future designs of federated learning systems and market mechanisms. The paper introduces a scalable polling algorithm for network intervention that leverages centrality-based welfare decomposition to achieve computational, statistical, and economic efficiency. By combining stochastic approximation with mechanism design techniques like decoy signals, it offers a robust solution to the entangled challenges of learning, strategy, and computation in large-scale networked systems, representing a significant advance in algorithmic game theory and distributed optimization.
AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain "reduced-concurrency" versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.
Primary: Microsoft Research
All Institutions: University of California, Riverside, Math, Inc., Microsoft Research
The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
The paper proposes RESOLVE, a three-stage pipeline for validating GPU kernels: (1) binary instrumentation to test for nondeterminism/races, (2) agent-based rewriting to create "reduced-concurrency" versions that are bitwise equivalent to the originals, and (3) formal verification of the reduced versions using F*/Pulse. The core novelty lies in the decoupling of concurrency correctness (handled via testing) from functional correctness (handled via proof), and the use of LLM agents to bridge the gap between complex production kernels and the limited semantic coverage of formal verifiers. The approach is clever but relies heavily on the agent's ability to produce correct reductions, which is a non-trivial assumption. The use of bitwise equivalence as an oracle for the reduction step is a strong design choice that avoids numerical tolerance issues during the intermediate phase.
The evaluation covers KernelBench, fused GEMMs in CUTLASS/Triton/Gluon, and mega-kernels. The finding of four previously unreported issues (including two bugs) in state-of-the-art frameworks is a significant empirical result, demonstrating the tool's practical utility. The comparison against existing tolerance-based tests shows that RESOLVE catches errors that standard testing misses. However, the evaluation is somewhat limited in scale (only three mega-kernels) and lacks a detailed analysis of the cost (time/compute) of the agent-based reduction and proof steps. The claim that agents can repair the issues with minimal performance impact is supported but would benefit from more extensive benchmarking.
The paper describes the pipeline clearly, but the reliance on "agents" to perform rewrites and proofs introduces variability. Without a fixed prompt strategy or a deterministic agent framework, exact reproduction of the results may be difficult. The use of NVBit and F*/Pulse is standard, but the specific agent configurations are not detailed enough for full reproduction. The code availability is not explicitly stated in the provided text, which is a gap for a systems paper.
The primary limitation is the dependence on LLM agents for the reduction and proof steps. If the agent fails to produce a valid reduction or proof, the pipeline stalls. The paper does not deeply analyze the failure modes of the agent or the success rate of the reduction step across a larger corpus. Additionally, the formal verification step is limited to the subset of CUDA/Kuiper supported by F*/Pulse, meaning kernels with exotic hardware features may still require manual intervention or may not be verifiable. The performance overhead of the validation pipeline itself is not thoroughly quantified.
This work has high potential impact on the reliability of AI-generated code, particularly in high-stakes domains like autonomous driving or financial modeling where GPU kernel correctness is critical. It provides a framework for integrating formal methods into the agentic coding loop, which is a growing area of interest. The discovery of bugs in production frameworks like CUTLASS and Triton highlights the immediate value of such tools. It may influence the development of future kernel compilers and verification tools to be more amenable to automated reduction and proof. The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed graph. We realize this task-general view through a common interface based on random walks, allowing the same model to operate across homogeneous and multi-relational graphs with varying features, labels, and relational schemas. Wander can increase its structural context at inference time without changing its learned parameters and, under suitable assumptions, universally approximates the corresponding Bayes-optimal predictor on bounded connected graphs. Empirically, a single pretrained checkpoint achieves state-of-the-art or highly competitive results across node classification, homogeneous link prediction, and knowledge-graph link prediction. Moreover, joint pretraining across graph modalities and tasks preserves performance in specialized settings while enabling positive transfer and the composition of separately learned capabilities at inference time.
Primary: AITHYRA (Austrian Academy of Sciences)
All Institutions: AITHYRA, Austrian Academy of Sciences, Boehringer Ingelheim Stiftung
Wander introduces a unified graph foundation model using random walks as a shared interface, achieving state-of-the-art performance across node classification, homogeneous link prediction, and knowledge graph link prediction while demonstrating positive transfer and compositional generalization. The paper provides a rigorous theoretical framework proving universality and permutation equivariance, and empirically validates the model's ability to adapt its structural context at inference time, offering a promising path toward generalist graph learning.
The paper proposes "Wander," a graph foundation model that unifies node classification, homogeneous link prediction, and knowledge graph link prediction under a single probabilistic framework: the completion of a partially observed graph. The core architectural innovation is the use of random walks as a shared computational interface, replacing standard message passing. This allows the model to dynamically adjust its structural context at inference time without retraining. The methodology is rigorous, providing a theoretical foundation that proves Wander is a universal approximator of the Bayes-optimal predictor for bounded connected graphs and is permutation equivariant in distribution. The design separates feature, label, and structural channels, processing them via intra-node attention, in-context learning updates, and stochastic structural updates via sampled walks. This is a sophisticated and well-motivated approach to the fragmentation problem in graph foundation models.
The experimental evaluation is extensive and addresses four key questions: generalist performance, transfer through joint training, compositional generalization, and inference-time adaptation. The model achieves state-of-the-art or highly competitive results across 54 knowledge graphs, 10 homogeneous link prediction datasets, and 26 node classification datasets. Notably, the paper demonstrates positive transfer from joint pretraining (improving homogeneous link prediction by ~3% over single-task training) and compositional generalization (combining node features and edge types, which were never seen together during training). The inference-time adaptation experiments on the synthetic Grids task are particularly compelling, showing that increasing walk length significantly improves long-range reasoning accuracy from 56.6% to 96.6%.
The paper provides detailed implementation details in the appendix, including initialization strategies, specific attention mechanisms (RoPE, RMSNorm), random walk sampling protocols, and loss functions. The pretraining data generation process is described, combining synthetic graph models (SBM, ER, Watts-Strogatz) with real-world KGs. While the code is not explicitly linked in the provided text, the level of detail suggests high reproducibility. The use of standard benchmarks and clear evaluation protocols (MRR, Hits@10, Recall@20, Accuracy) further supports reproducibility.
The pretraining prior is limited, relying heavily on synthetic data for attributed graphs and only three real-world knowledge graphs for multi-relational pretraining. Inference cost can be high due to the stochastic nature of random walk sampling, particularly when large budgets are required for long-range reasoning. The model does not yet support node regression or graph classification tasks. The theoretical guarantees rely on capacity assumptions that may not hold for the finite-sized models used in practice.
This paper makes a significant contribution to the field of graph machine learning by demonstrating that a single model can effectively handle diverse graph modalities and tasks. The random walk interface offers a flexible alternative to message passing, with potential applications in dynamic graphs and scenarios where computational resources can be allocated adaptively at inference time. The positive transfer and compositional generalization results suggest that joint pretraining across tasks is a viable strategy for building more robust graph foundation models. Wander introduces a unified graph foundation model using random walks as a shared interface, achieving state-of-the-art performance across node classification, homogeneous link prediction, and knowledge graph link prediction while demonstrating positive transfer and compositional generalization. The paper provides a rigorous theoretical framework proving universality and permutation equivariance, and empirically validates the model's ability to adapt its structural context at inference time, offering a promising path toward generalist graph learning.
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-$k$ misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.
Primary: New York University
All Institutions: New York University, NYU Shanghai
The paper introduces a distributionally robust training objective for MoE models that improves routing reliability by strengthening non-top expert paths. It provides a rigorous theoretical framework and empirical evidence that optimizing for expert competence, rather than just load balance, leads to more robust and specialized MoE models at scale.
The paper proposes Distributionally Robust Mixture-of-Experts Training (DRMoET), a drop-in training objective that treats layer-wise experts as endogenous robustness groups. Unlike standard load-balancing losses that optimize for traffic distribution, DRMoET optimizes for expert competence under imperfect routing. It utilizes an entropy-regularized softmax update on EMA-smoothed, activation-weighted expert losses to strengthen plausible non-top routing paths. The method is theoretically grounded, providing convergence guarantees for the entropy-regularized robust objective, and is computationally efficient, adding negligible overhead to the standard MoE forward/backward pass. The distinction between "allocation" (load balancing) and "competence" (routing robustness) is a well-motivated and clear conceptual contribution.
The authors conduct rigorous experiments at two scales (746M and 10.3B total parameters) using the FLAME-MoE recipe. Results show consistent improvements in downstream task averages over both standard FLAME-MoE and auxiliary-loss-free baselines. Mechanistic analyses are strong, including expert-loss variance reduction, forced misrouting probes (showing 4.3% lower excess loss), and improved domain-expert specialization metrics. The ablation studies effectively isolate the contribution of activation-weighted credit and EMA decay. The inclusion of throughput analysis confirms the method's practical viability.
The paper provides detailed hyperparameters, data sources (DCLM), and training recipes. Code and project page are available. The method is described as a "drop-in" objective, suggesting ease of integration into existing MoE training pipelines. The specific implementation details of the EMA and dual variable updates are clearly specified in the algorithm box and appendix.
The evaluation is limited to two model scales and specific expert configurations. The paper acknowledges that it does not test highly imbalanced pretraining mixtures or long-tail evaluation suites, which would be a more direct stress test for robustness. The gains, while consistent, are moderate (e.g., 1.42 percentage points at 10.3B scale), which may be considered small in the context of large-scale LLM training where other factors often dominate.
As MoE architectures become standard for scaling LLMs, addressing the reliability of sparse routing is increasingly important. This work provides a practical tool to improve the robustness of MoE models without architectural changes, potentially leading to more reliable inference under distribution shift or imperfect gating. It shifts the focus from mere utilization balance to expert quality, a perspective likely to influence future MoE training strategies. The paper introduces a distributionally robust training objective for MoE models that improves routing reliability by strengthening non-top expert paths. It provides a rigorous theoretical framework and empirical evidence that optimizing for expert competence, rather than just load balance, leads to more robust and specialized MoE models at scale.
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
Primary: University of California, Berkeley
All Institutions: University of California, Berkeley, Massachusetts Institute of Technology (MIT)
The paper demonstrates that base models' reasoning capabilities are heavily influenced by specific starting token cues learned from training data, and that these cues can be manipulated to significantly enhance performance. This is a significant contribution to mechanistic interpretability and practical LLM optimization, offering a cost-effective alternative to RL for improving reasoning tasks.
The paper employs a rigorous causal intervention framework to dissect the role of "token cues" in Large Language Model (LLM) reasoning. The methodology is three-pronged: (1) Empirical demonstration that fixing specific starting tokens (e.g., ".\n\nOkay") significantly boosts base model performance on math and coding tasks, rivaling RL-trained models; (2) Causal data interventions (likely using activation patching or data ablation techniques) to prove that these cues are learned from specific document types in the training data, rather than being inherent to the token embeddings; (3) A safety case study showing that cues also modulate refusal behaviors. The approach is sophisticated, moving beyond correlation to causation by editing the model's internal representations or training data associations to isolate the effect of the cue.
The experiments are extensive and compelling. The authors test across multiple model families (OLMo-3, Qwen3) and tasks (MATH-500, coding). The jump in accuracy (e.g., 42% to 78% for OLMo-3-7B) is substantial and practically significant. The causal interventions are well-designed to rule out confounding factors. The safety analysis adds depth, showing that the same mechanism that aids reasoning also influences safety alignment, which is a critical insight for the field.
The paper appears to be from a high-quality group (Berkeley/MIT) with clear methodological descriptions. While specific code links are not provided in the snippet, the detailed description of the causal interventions and the use of open-source models (OLMo, Qwen) suggests high reproducibility. The abstract-only score of 60 suggests the full text provides sufficient detail for replication.
The study focuses on base models and the transition to RL. It may not fully account for how these cues interact with more complex agentic workflows or multi-turn conversations. The "chicken" example, while illustrative, is a synthetic intervention; the real-world generalizability of arbitrary word cues is limited, though the core finding about training data associations is robust.
This paper has high impact because it challenges the prevailing narrative that RL is the sole driver of reasoning capabilities in LLMs. It suggests that data curation and prompt engineering (specifically starting tokens) are underutilized levers for improving base model performance. This has immediate implications for model developers, who can optimize pre-training data and inference prompts without expensive RL fine-tuning. It also has safety implications, as it reveals how minor prompt changes can shift model behavior between compliance and refusal. The paper demonstrates that base models' reasoning capabilities are heavily influenced by specific starting token cues learned from training data, and that these cues can be manipulated to significantly enhance performance. This is a significant contribution to mechanistic interpretability and practical LLM optimization, offering a cost-effective alternative to RL for improving reasoning tasks.
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
Primary: TU Darmstadt
All Institutions: TU Darmstadt, Hessian.AI Service Center, Konrad Zuse School of Excellence in Learning and Intelligent Systems
The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
The paper introduces Privacy-Leaking Watermarks (PLWs), a novel adversarial attack on unified multimodal models. The core innovation lies in exploiting the shared latent space of unified architectures to embed trigger-dependent watermarks in generated images based on prior conversational context. The methodology involves a two-stage training process: first, training a watermark encoder/extractor pair, and second, fine-tuning the multimodal model (using LoRA) to condition the watermark embedding on specific semantic triggers found in the chat history. This approach is technically sound and cleverly leverages the specific architectural properties of unified models (where text and image generation are not strictly decoupled) to create a covert side-channel for privacy leakage.
The experiments are rigorous, testing the attack across two model families (including OmniGen2) and 13 sensitive attribute triggers. The results are striking, achieving up to 100% True Positive Rate at a 1% False Positive Rate, demonstrating that the watermarks are both highly detectable and effectively hidden from casual observation. The evaluation correctly isolates the threat model to locally deployed models where users assume privacy, providing a clear and compelling demonstration of the vulnerability.
The paper provides a strong reproducibility statement, detailing the training configurations, LoRA target modules, and specific seeds used. While the authors explicitly state they do not release poisoned checkpoints (a responsible security practice), the detailed appendices and methodological transparency allow for independent verification of the attack's feasibility.
The primary limitation is the requirement for white-box access to modify and redistribute model weights, which restricts the threat to malicious model providers or compromised supply chains rather than arbitrary users. Additionally, the attack is specific to unified multimodal architectures; it may not apply to traditional pipeline-based multimodal systems where text and image generation are separate modules.
This paper has significant implications for the deployment of unified multimodal models, particularly in consumer-facing applications where local deployment is marketed as a privacy feature. It highlights a critical gap in the security model of these architectures and will likely drive research into secure design patterns, watermarking defenses, and the separation of modalities in unified models. The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.
Primary: Georgia Institute of Technology
All Institutions: Georgia Institute of Technology, Simon Fraser University, University of Pennsylvania, University of Texas at Austin, Rutgers University, Foothill College
The paper presents a novel co-evolutionary framework for proof auto-formalization that jointly optimizes a large open-weight LLM and its agentic harness, achieving state-of-the-art results on a new research-level benchmark while significantly improving cost-efficiency compared to frontier API-based agents.
The paper proposes a joint optimization framework for both the language model and its agentic harness. The Semantic Alignment Model (SAM) introduces a contrastive loss to align natural language and formal language representations, which is a logical extension of existing alignment techniques but applied here to a large MoE model. The core novelty lies in HarnessEvolve, an evolutionary search algorithm that treats the agent's control flow (harness) as a mutable program, using diagnostic certificates from verifiers to guide mutations. This is a significant departure from static scaffolding or simple prompt optimization. The Agentic RLSF component adapts Group-Relative Policy Optimization for multi-turn agent interactions, using a fine-grained reward ladder that distinguishes between type errors, incompleteness, and semantic drift. This multi-objective reward design is crucial for preventing the "silent correction" of flawed proofs, a known failure mode in prior work.
The introduction of LoCoBench is a major contribution, providing 58.9k instances and a 771-instance validation set that is out-of-distribution (Mizar-to-Lean translation). The evaluation against 39 baselines is comprehensive. The results show a substantial improvement over the base model (15.7% to 36.7% pass@4 semantic correctness) and demonstrate that the framework can enhance frontier coding agents (Claude Code, Codex) when used as a specialized skill. The cost-efficiency analysis is particularly strong, showing that the open-weight system achieves competitive accuracy at a fraction of the cost of frontier API calls.
The paper provides extensive details on the training setup, including hardware, hyperparameters for LoRA and GRPO, and the specific configuration of the evolutionary search. The release of the benchmark and the open-weight model base (Leanstral-1.5) enhances reproducibility. However, the reliance on a frontier coding agent (Claude Opus 5) for the mutation step in HarnessEvolve introduces a dependency on a proprietary service, which may limit full reproducibility for all researchers.
The primary limitation is the heavy reliance on a frontier LLM for the harness evolution process, which contradicts the goal of full autonomy and open-source accessibility. The semantic correctness check relies on an extended BEq+ prover, which may have its own coverage limitations. Additionally, the "proof faithfulness" metric is not fully automated, relying on LLM judges which are noted to be less reliable.
This work has significant implications for the accessibility of formal verification. By demonstrating that open-weight models can be post-trained to perform research-level auto-formalization, it lowers the barrier to entry for formal methods. The co-evolution of model and harness offers a generalizable paradigm for improving agentic systems in other domains where tool-use and control flow are critical. The paper presents a novel co-evolutionary framework for proof auto-formalization that jointly optimizes a large open-weight LLM and its agentic harness, achieving state-of-the-art results on a new research-level benchmark while significantly improving cost-efficiency compared to frontier API-based agents.
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
The paper demonstrates that high accuracy on scientific benchmarks does not imply evidence grounding, introducing a verified counterfactual protocol to measure and address this gap. By showing that agents often rely on prior knowledge rather than supplied data, and that standard evaluators fail to penalize this, the work provides a crucial methodological correction for the evaluation of AI scientists.
The paper introduces a rigorous framework for evaluating "evidence grounding" in AI scientific agents. The core methodological contribution is the construction of verified counterfactuals (null, withdrawal, flip) for benchmark items, where the effect of the intervention on the ground truth is pre-registered and verified via reference statistics. This allows for the calculation of Evidence-Grounded Accuracy (EGA), which distinguishes between answers derived from data versus those recalled from prior knowledge or reached via elimination. The formalization of "noticing without acting" and the separation of agent grounding from evaluator validity are strong theoretical contributions. The protocol for cleaning metadata to prevent leakage (canonical writer) is a crucial practical detail that enhances the validity of the counterfactuals.
The experiments are well-designed but limited in scale. The authors evaluate 18 single-cell questions from BAISBench and 4 synthetic problems from GeneBench-Pro. While small, this is sufficient to demonstrate the phenomenon of high accuracy coexisting with low evidence grounding. The finding that Claude agents achieve 95% accuracy but only 41% EGA is a significant empirical result. The gene-name anonymization experiment effectively isolates prior knowledge as a confounding factor. The evaluation of evaluators (benchmark scores and LLM judges) showing they reward ungrounded answers is a critical finding for the field.
The paper provides high reproducibility. It details the canonical writer process, the specific interventions applied to the data, the agent scaffolds used (Claude Code, Codex CLI, Biomni), and the evaluation metrics. The pre-registration of reference statistics and the verification of every edit make the counterfactuals robust. The code and data protocols are described in sufficient detail for replication, although the specific datasets (BAISBench, GeneBench-Pro) are external dependencies.
The primary limitation is the small number of tasks (22 total). The findings may not generalize to all types of scientific tasks or domains beyond single-cell biology and synthetic genetics. The manual construction of counterfactuals is labor-intensive, limiting scalability. The study focuses on a specific set of models (Claude, GPT-5.6) and scaffolds, so results may vary for other architectures. The "flip" intervention assumes a binary or clear alternative, which may not apply to all scientific questions.
This paper has high impact on the evaluation of AI agents in scientific discovery. It challenges the validity of current benchmarks that rely solely on final answer accuracy. The proposed metrics (EGA, counterfactual validity) provide a new standard for assessing whether AI agents truly use the data provided. This work will likely influence the design of future benchmarks for AI scientists and the interpretation of existing leaderboard scores. It highlights a critical gap in current AI evaluation practices that affects the trustworthiness of AI-generated scientific insights. The paper demonstrates that high accuracy on scientific benchmarks does not imply evidence grounding, introducing a verified counterfactual protocol to measure and address this gap. By showing that agents often rely on prior knowledge rather than supplied data, and that standard evaluators fail to penalize this, the work provides a crucial methodological correction for the evaluation of AI scientists.
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Primary: University of New South Wales (UNSW Sydney)
All Institutions: University of New South Wales (UNSW Sydney)
ASCENT enables stable online test-time training of LLM agents by self-distilling verified experience through a frozen teacher model, significantly improving task success and efficiency in long-horizon environments without requiring external references or multiple attempts.
The paper proposes ASCENT, a method for Online Agentic Test-Time Training (OaTTT). The core innovation is using the frozen initial model as a teacher that conditions on the "privileged information" of a verified successful trajectory (hindsight) to generate soft targets for the current student model (via LoRA). This avoids the instability of direct imitation or reinforcement learning on single-attempt trajectories, which the authors demonstrate collapses the policy. The method effectively combines rejection sampling (only updating on verified successes) with on-policy distillation. The theoretical framing of the population target and the analysis of why direct imitation fails (sharpening around generated tokens vs. full-vocabulary matching) are strong contributions.
The experiments are extensive, covering ALFWorld, WebShop, and AppWorld with two model scales (Qwen3.5-4B and 9B). The paper provides strong evidence for the method's efficacy, showing significant improvements over the base model and various in-context adaptation baselines (MemP, ACE, etc.). The ablation studies on privileged information content and distillation divergence are thorough. The demonstration that direct imitation baselines fail catastrophically is a valuable negative result for the community.
The paper provides detailed descriptions of the protocol, loss functions, and hyperparameters. The use of standard benchmarks and open-weight models (Qwen) aids reproducibility. However, the specific implementation of the "validity filter" and the exact serialization of privileged information $z_i$ could benefit from more code-level detail, though the project page is provided.
The method requires access to the model's weights (open-source models only) and the ability to run a frozen copy of the model as a teacher, which doubles inference cost during the update phase. The reliance on sparse, episode-level verification limits the granularity of learning signals. The paper acknowledges that the teacher's privileged context may allow it to reconstruct student tokens, potentially limiting the independence of the hindsight signal.
This work is significant for the deployment of LLM agents in dynamic environments where continuous improvement is desired without offline retraining. It bridges the gap between in-context learning (memory-based) and parametric adaptation, offering a stable mechanism for weight updates. The findings on the instability of single-attempt RL/imitation are likely to influence future agent training designs. ASCENT enables stable online test-time training of LLM agents by self-distilling verified experience through a frozen teacher model, significantly improving task success and efficiency in long-horizon environments without requiring external references or multiple attempts.
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology, Qualcomm AI Research
Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
The paper proposes Score-Calibrated Flow (SCF), a method for training flow matching models to sample from unnormalized densities without access to target samples. The core innovation is a "self-consistency" approach that derives optimality conditions for the velocity field based on the target score (provided by the critic in RL) and the structure of the flow itself. The authors prove that the unique solution to these conditions is the ideal flow model that would be learned via standard Conditional Flow Matching (CFM) if target samples were available. The training objective is a stop-gradient regression on self-generated endpoints, avoiding the high variance of importance sampling and the computational cost of backpropagating through sampling trajectories. The theoretical grounding is strong, with clear proofs of uniqueness for both the terminal density and the velocity field.
The experiments include a toy example comparing SCF against recent samplers (AS, ASBS, BMS, FS) and extensive RL benchmarks (DeepMind Control Suite, HumanoidBench). SCF is compared against 8 strong baselines, including SAC and various generative policy methods (DIME, QSM, QFlex, etc.). The results show SCF matching or improving upon state-of-the-art methods while significantly reducing training time. The inclusion of wall-clock time comparisons is a strong practical contribution.
The paper provides detailed algorithmic descriptions and theoretical derivations. However, as a preprint, specific hyperparameters and code availability are not confirmed in the text. The method relies on standard flow matching components, making it relatively easy to implement if the specific coefficient schedules (mentioned in the appendix) are clear.
The method assumes access to the gradient of the potential function (critic), which is standard in RL but may not be available in all sampling tasks. The theoretical guarantees rely on certain regularity conditions (e.g., exponential moments) that may be hard to verify in practice for complex, high-dimensional distributions. The performance gains, while present, are incremental over strong baselines in some tasks.
This work addresses a fundamental bottleneck in applying generative models to online RL and other settings where only unnormalized densities are available. By providing a stable, efficient, and theoretically grounded training procedure, it could become a standard tool for training expressive policies in RL and for sampling in physics and molecular dynamics simulations. Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at https://github.com/Ayanami-0123/Open-Mesh-Learning.
Primary: Peking University
All Institutions: Peking University
The paper identifies and mitigates catastrophic strategy collapse in RLVR for LLMs by introducing Mesh Learning. It provides a rigorous theoretical framework linking optimization dynamics to information-theoretic capacity constraints, supported by strong empirical results across multiple benchmarks and model families, establishing strategy preservation as a key principle for stable post-training.
The paper addresses a critical failure mode in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs), specifically the "catastrophic strategy collapse" observed in GRPO-style algorithms. The authors propose a theoretical framework combining optimization dynamics and information theory to explain this phenomenon, defining "strategies" via trajectory-level policy-update interactions. They introduce the Mirrored Entanglement Index (MEI) as an online diagnostic metric and propose "Mesh Learning," a method designed to preserve strategy diversity by preventing any single reasoning strategy from dominating the optimization landscape. The theoretical contribution is significant as it moves beyond empirical observation to provide a mechanistic explanation for why standard RLVR objectives lead to entropy collapse and reduced effective capacity.
The experimental section evaluates the proposed method across a robust suite of benchmarks including AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench. The models tested include the Qwen and Phi families, which are widely used in the community. The reported gains are substantial, with improvements of up to 13.4 percentage points on Qwen models and 11.5 percentage points on Phi models compared to strong baselines. The consistency of performance across diverse tasks (math, code, general reasoning) suggests that the method effectively preserves the general reasoning capabilities of the model while enhancing specific task performance, rather than overfitting to a narrow distribution of strategies.
The paper provides a public code repository (https://github.com/Ayanami-0123/Open-Mesh-Learning), which is a strong positive for reproducibility. Given the complexity of RLVR training and the specific hyperparameters likely required for "Mesh Learning," the availability of code is crucial. The use of standard, open-source model families (Qwen, Phi) further enhances the reproducibility of the results for other researchers.
The primary limitation is the computational cost associated with RLVR training, which remains high despite the proposed method. Additionally, the definition of "strategies" via trajectory-level interactions may be complex to interpret in non-mathematical or open-ended generation tasks, potentially limiting the direct applicability of the MEI metric to broader domains. The paper focuses on verifiable rewards, so its applicability to RLHF with human preference data (where rewards are less verifiable) is not directly addressed.
This work has significant implications for the stability and reliability of LLM post-training pipelines. By identifying and mitigating strategy collapse, it contributes to the development of more robust and versatile reasoning models. The concept of "strategy preservation" could influence future RLVR algorithms, potentially leading to a new class of training objectives that explicitly balance exploration and exploitation in the policy space. This is particularly relevant as the industry moves towards more complex, multi-step reasoning tasks where diverse strategies are essential. The paper identifies and mitigates catastrophic strategy collapse in RLVR for LLMs by introducing Mesh Learning. It provides a rigorous theoretical framework linking optimization dynamics to information-theoretic capacity constraints, supported by strong empirical results across multiple benchmarks and model families, establishing strategy preservation as a key principle for stable post-training.
Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.
Primary: Stanford University
All Institutions: Stanford University
The paper provides a rigorous theoretical and empirical analysis of how divergence choices in distillation control student entropy, proving that forward KL inflates entropy while reverse KL deflates it, and demonstrating that this divergence acts as an implicit regularizer crucial for preventing entropy collapse in self-distillation. By decoupling sampling and divergence objectives and validating these insights on large-scale models like OLMo and Qwen, the work offers actionable insights for improving the stability and diversity of distilled language models.
The paper proposes a rigorous theoretical framework to analyze the effect of divergence choices in knowledge distillation on the entropy of the student model. The core methodological contribution is the decoupling of the sampling distribution (on-policy vs. off-policy) from the divergence objective (forward KL, reverse KL, etc.) by defining the objective at the token level rather than the sequence level. This allows for independent analysis of how each component affects entropy. The authors introduce a toy model of language modeling (linear softmax head with random embeddings) that is tractable for theoretical analysis yet captures key phenomena of large language models. They prove that forward KL inflates student entropy relative to the teacher by the residual divergence, and that reverse KL deflates entropy until the task becomes too hard, at which point it inflates. The analysis extends to generalized Jensen-Shannon divergences and top-k restrictions.
The experimental evaluation is strong and multi-faceted. The authors validate their theoretical predictions using the OLMo 2 suite of models, demonstrating that the entropy of a model matches its cross-entropy loss at convergence, a finding they claim is novel. They conduct ablation studies on Qwen3 models to show that the divergence choice has a larger impact on entropy than the sampling strategy (on-policy vs. off-policy). Finally, they apply their insights to self-distillation on SciKnowEval and GSM8K, showing that specific divergence hyperparameters are necessary to prevent entropy collapse when conditioning on privileged information. The experiments are well-designed to test specific theoretical claims.
The paper provides a link to a GitHub repository containing code to reproduce all experiments and figures, along with the underlying data. The use of open-source models (OLMo, Qwen3) and public datasets (DeepScaleR, SciKnowEval, GSM8K) further enhances reproducibility. The theoretical proofs are detailed in the appendix, allowing for verification of the mathematical claims.
The theoretical analysis relies on a toy model with random embeddings, which may not fully capture the complexities of real-world language models with deep sequence modeling components. While the authors argue that the token-level objective is what matters in practice due to stop-gradients, the gap between the toy model and large-scale transformers remains a potential limitation. The experiments, while thorough, are limited to specific model families (OLMo, Qwen) and datasets, and the generalizability to other architectures or domains is not explicitly tested.
This paper has significant implications for the design of distillation pipelines in LLM training. By identifying the divergence as an implicit entropy regularizer, it provides practitioners with a new knob to control model uncertainty and diversity. The finding that forward KL inflates entropy and reverse KL deflates it (until failure) offers practical guidance for choosing objectives in post-training and self-distillation. It also challenges the common practice of deriving on-policy distillation from sequence-level reverse KL, suggesting that decoupling these choices can lead to better performance and stability. The paper provides a rigorous theoretical and empirical analysis of how divergence choices in distillation control student entropy, proving that forward KL inflates entropy while reverse KL deflates it, and demonstrating that this divergence acts as an implicit regularizer crucial for preventing entropy collapse in self-distillation. By decoupling sampling and divergence objectives and validating these insights on large-scale models like OLMo and Qwen, the work offers actionable insights for improving the stability and diversity of distilled language models.
Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.
Primary: Rice University
All Institutions: Rice University, Baylor College of Medicine, Texas Children's Hospital
EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
The paper proposes EchoDino, a self-supervised foundation model for echocardiography by adapting the DINOv3 framework. The core methodological contribution is the domain adaptation of a natural image foundation model to pediatric cardiac ultrasound using 3.7 million unlabeled frames. The authors introduce two specific architectural/algorithmic extensions: EchoDino-PATCH, which re-aggregates spatial patch tokens to preserve local anatomical detail for segmentation and measurement, and MEMS (Motion-biased Entropy Maximization Sampling), a frame selection strategy for video-level tasks that prioritizes high-motion, feature-diverse frames over uniform sampling. The approach relies on a frozen encoder with lightweight task-specific readouts, which is a standard but effective paradigm for foundation models. The adaptation of DINOv3's teacher-student architecture to the specific augmentation and cropping needs of echocardiography is well-motivated.
The experimental evaluation is extensive and rigorous. The model is tested across nine datasets, including five internal pediatric cohorts, one external pediatric cohort, and three external adult datasets. This cross-domain evaluation (pediatric to adult) is a strong point, demonstrating the generalizability of the learned representations. The tasks cover a wide spectrum: global (view classification), localized (measurement, SHD detection), dense (segmentation), and temporal (EF prediction, sweep recognition). The results show consistent and significant improvements over strong baselines like DINOv3, PanEcho, and EchoPrime. The use of bootstrap confidence intervals and patient-disjoint splits adds statistical rigor. The ablation of MEMS vs. uniform sampling clearly isolates the benefit of the proposed sampling strategy.
The paper provides a link to the code repository. However, the primary pretraining data (TCH-Complex) is not publicly available due to clinical data governance, which limits full reproducibility of the pretraining phase. The downstream evaluation on public datasets (EchoNet, CAMUS) is reproducible. The hyperparameters and training details are described in the supplementary methods, aiding partial reproducibility.
The main limitation is the lack of public access to the large-scale pediatric pretraining corpus, which prevents independent verification of the pretraining process. The SHD detection is evaluated as a binary composite classifier, which may mask performance on specific rare lesions. The paper acknowledges that aggregate metrics do not guarantee individual point-of-care actionability and that uncertainty estimation is needed for clinical deployment.
This work has significant potential impact in the field of medical imaging AI, particularly in pediatric cardiology where labeled data is scarce. By demonstrating that a single frozen encoder can support diverse tasks across age groups, it offers a scalable framework for developing modular clinical tools. The approach of adapting general-purpose vision foundation models to specific medical imaging modalities is a trend that this paper contributes to meaningfully. EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
Primary: University of Michigan
All Institutions: University of Michigan, Southeast University, Alibaba Group
The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
The paper proposes RAESR, a super-resolution model that operates in the latent space of a frozen DINOv3-L vision transformer rather than a VAE or pixel space. The core methodological contribution is the argument that self-supervised representation spaces (specifically DINOv3) embed degraded images closer to the "manifold" of clean images than reconstruction-oriented spaces (VAEs), thereby simplifying the mapping back to the clean manifold. The authors validate this geometric intuition with rigorous ablations, including latent line interpolation experiments and layer-wise information recovery analysis. The decoder is a 415M parameter transformer trained with reconstruction and adversarial losses, using a multi-layer DINOv3-B critic for fine-grained realism supervision. The approach is technically sound, well-motivated, and the ablations are exceptionally thorough, providing strong evidence for the central hypothesis.
The experiments are comprehensive, comparing RAESR against 12 state-of-the-art methods across four real-world benchmarks (RealSR, DRealSR, LSDIR, DIV2K-Val). RAESR achieves the best fidelity-perception trade-off, outperforming heavy diffusion-based models in both quality and efficiency (37ms per image). The inclusion of a human preference study and detailed per-benchmark breakdowns adds robustness. The ablation studies, particularly the comparison against a VAE substrate with identical training recipes, are critical and convincingly demonstrate that the performance gain stems from the choice of latent space rather than just the decoder architecture.
The paper provides extensive details on training schedules, hyperparameters, data processing, and evaluation protocols in the appendices. The use of standard datasets and public baselines facilitates reproduction. However, the reliance on specific frozen checkpoints (DINOv3-L, RAEv2) and the complex multi-stage training process may pose some barriers for independent replication without access to the authors' code or precise checkpoint versions.
The model is a single-pass restorer, lacking the ability to trade compute for quality on difficult images, unlike iterative diffusion methods. The claim about substrate suitability is currently supported primarily by DINOv3-L and tested against one alternative (SD3 VAE), leaving the behavior of other representation encoders unexplored. The model size (719M inference parameters) is still significant compared to lightweight GANs, though it is competitive with diffusion models.
This work shifts the perspective in super-resolution research from merely improving generative priors to carefully selecting the representation space for restoration. It highlights the utility of self-supervised vision foundation models for low-level vision tasks, potentially inspiring similar approaches in other restoration tasks like denoising or inpainting. The efficiency gains over diffusion models make it attractive for real-time applications. The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.
Primary: Harvard University
All Institutions: Harvard University, Johns Hopkins University, Kempner Institute
[One sentence main contribution]. MoSE3 introduces the first feed-forward model for dense SE(3) motion estimation from monocular video, utilizing a novel decomposition into point tracks and rigidity embeddings to handle the non-Euclidean nature of rotations, and contributes a large-scale synthetic dataset, Art-Kubric, to enable training and evaluation. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper addresses a fundamental gap in motion modeling by moving beyond 3-DoF translation to full 6-DoF rigid transforms. The technical approach is sound, leveraging differentiable optimization to fit SE(3) transforms within learned rigid clusters, which effectively circumvents the difficulties of direct manifold regression. The introduction of Art-Kubric is a substantial contribution to the field, providing the necessary ground truth data for this complex task. The experimental results are strong, showing SOTA performance on both the new SE(3) task and existing point tracking benchmarks, with impressive generalization to real-world data. This work is likely to influence future research in dense motion estimation and provide a valuable tool for applications requiring detailed understanding of scene dynamics.
The paper proposes MoSE3, a feed-forward architecture that predicts dense SE(3) motion (6-DoF: rotation and translation) from monocular RGB video. The core innovation lies in avoiding direct regression on the non-Euclidean SO(3) manifold by decomposing the problem into two jointly learned intermediates: 3D point tracks (translations) and rigidity embeddings. The SE(3) transforms are then recovered via differentiable fitting within soft rigid clusters defined by the rigidity embeddings. This approach elegantly handles the manifold constraint and the grouping of pixels into rigid bodies. The introduction of the Art-Kubric dataset, featuring dense SE(3) and rigidity labels for articulated objects with physical interactions, is a significant methodological contribution to address the lack of ground truth data for this specific task.
The experiments demonstrate state-of-the-art performance in SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks. The model also achieves state-of-the-art average 3D point tracking accuracy across three datasets. A particularly strong result is the generalization to real-world videos despite training solely on synthetic data, which validates the robustness of the learned representations. The evaluation covers both the novel SE(3) task and the established point tracking task, providing a comprehensive assessment of the model's capabilities.
The paper provides a project page URL. As an arXiv preprint, the code availability is not explicitly confirmed in the provided text, but the detailed description of the method and the release of a large-scale synthetic dataset (Art-Kubric) suggest a high level of reproducibility. The use of standard synthetic data generators (Kubric) for the dataset creation further aids reproducibility.
The primary limitation is the reliance on synthetic data for training, which may introduce a domain gap for certain real-world scenarios not covered by the synthetic generator, although the paper claims strong generalization. The method assumes rigid or articulated motion within clusters; highly deformable objects might not be captured accurately by the SE(3) fitting approach. The computational cost of differentiable fitting within clusters could be a bottleneck for very high-resolution videos or long sequences.
This work has significant implications for robotics, augmented reality, and video understanding. Dense SE(3) motion estimation provides a richer representation of scene dynamics than point tracking alone, enabling better understanding of object interactions, part-level motion, and scene structure. The Art-Kubric dataset will likely become a standard benchmark for motion estimation tasks involving articulated objects. The ability to predict 6-DoF motion from monocular video could enhance applications in robot manipulation, human motion analysis, and 3D scene reconstruction. [One sentence main contribution]. MoSE3 introduces the first feed-forward model for dense SE(3) motion estimation from monocular video, utilizing a novel decomposition into point tracks and rigidity embeddings to handle the non-Euclidean nature of rotations, and contributes a large-scale synthetic dataset, Art-Kubric, to enable training and evaluation. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper addresses a fundamental gap in motion modeling by moving beyond 3-DoF translation to full 6-DoF rigid transforms. The technical approach is sound, leveraging differentiable optimization to fit SE(3) transforms within learned rigid clusters, which effectively circumvents the difficulties of direct manifold regression. The introduction of Art-Kubric is a substantial contribution to the field, providing the necessary ground truth data for this complex task. The experimental results are strong, showing SOTA performance on both the new SE(3) task and existing point tracking benchmarks, with impressive generalization to real-world data. This work is likely to influence future research in dense motion estimation and provide a valuable tool for applications requiring detailed understanding of scene dynamics.
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Primary: ETH Zurich (Inferred from "eth-lre" in GitHub URL and Swiss AI Initiative funding)
All Institutions: ETH Zurich, Swiss National Supercomputing Centre (CSCS)
The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
The paper addresses a critical flaw in existing RL-based tutoring methods: the "telling" problem, where models learn to simply provide answers to maximize immediate student success metrics. The authors propose a novel reward structure based on three conditions: (1) Near-transfer testing (student solves a variant problem, not the exact one), (2) Masked post-test (tutor's utterances are hidden during the test, forcing reliance on student-generated notes/reasoning), and (3) Binary reward gates (strict penalties for factual errors or solution handover). This design effectively decouples "helpfulness" from "teaching," forcing the model to elicit student reasoning rather than provide it. The use of a frozen LLM student with genuine errors (rather than prompted confusion) is a strong methodological choice that ensures the training signal reflects real diagnostic and adaptive challenges.
The experimental setup is rigorous and comprehensive. The authors train models across multiple sizes (4B to 27B) and architectures (Qwen3 variants), demonstrating the robustness of the recipe. The ablation study is particularly strong, isolating the contribution of each condition (transfer, masking, gates) and showing that the gates specifically reduce handover while the masked transfer test improves out-of-domain generalization. The evaluation uses independent benchmarks (MathTutorBench, TutorMoments) and judges (Gemini) distinct from the training setup, mitigating overfitting concerns. The results show that Eduardo-27B matches or exceeds frontier models (Gemini 3.1 Pro, Claude Opus 4.8) in pedagogical quality while using significantly fewer thinking tokens, a crucial efficiency gain for interactive applications.
High. The authors open-source the training environment, the 8,671-problem near-transfer dataset, and the trained models. Detailed hyperparameters, prompts, and compute costs are provided in the appendices. The use of standard RL algorithms (DPPO/GRPO) and open-source base models further enhances reproducibility.
The primary limitation is the reliance on a single frozen LLM student (Llama-3.1-8B-Instruct) for training, which may lead to overfitting to that specific model's failure modes. The paper acknowledges this and suggests future work with diverse student ensembles. Additionally, the evaluation is limited to mathematics, and the "affective gap" (lack of socio-emotional support) is noted as a consequence of optimizing purely for cognitive gain. The ablation study is limited to one seed and one model size (4B), which slightly weakens the statistical confidence in the ablation results.
This work has significant implications for the development of AI tutors and, more broadly, for any domain where an agent must build user capabilities rather than just provide answers. The "masked near-transfer" reward design is a generalizable principle that could be applied to other educational or collaborative tasks. The efficiency gains (fewer thinking tokens) make high-quality tutoring more feasible in real-time interactive settings. The open-sourcing of the dataset and models will likely accelerate research in AI for Education. The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.
Primary: Princeton University
All Institutions: Princeton University
Queen introduces a 4B-parameter chess-language model that achieves Grandmaster-level play (2697 Elo) and high-quality explanations by integrating a silent expert encoder with an LM via cross-attention and iterative search distillation. This paper makes a significant contribution to the field by demonstrating that hybrid architectures combining specialized encoders with general LMs can outperform frontier LMs in domain-specific reasoning tasks, offering a scalable and efficient alternative to pure LLM approaches for tasks requiring deep domain expertise.
The paper proposes a novel hybrid architecture, "Queen," that integrates a silent expert chess encoder (Leela/BT5) with a general-purpose language model decoder (SmolLM3-3B) via a Flamingo-inspired gated cross-attention bridge. The methodology is rigorous, featuring a two-stage training process: (1) Domain Adaptation using a curated QA curriculum to teach the LM to interpret the encoder's latent representations, and (2) Iterative Search Distillation, a self-improvement loop inspired by Bellman updates where the model analyzes child positions and consolidates explanations, which are then distilled back into the model. This approach effectively bridges the gap between strong silent experts and fluent but weak LMs.
The experimental evaluation is comprehensive and compelling. Queen achieves a 2697 Elo rating, surpassing frontier LMs like GPT-5.6-Sol (2071) and Gemini-3.1-Pro (2201) by a significant margin while using 3 orders of magnitude fewer parameters. The paper introduces a robust evaluation framework covering accuracy (Elo), substantiation (no-mistake rate on puzzles), and coherence (LM-judged). Ablations clearly demonstrate the necessity of both the encoder and the curriculum. The comparison against a "HCE" variant shows that the method is not solely reliant on frontier model distillation for its core strength.
High. The authors provide code, model weights, and detailed appendices including hyperparameters, data construction pipelines, and prompt templates. The use of open-source components (SmolLM3, Leela) and public datasets (Lichess) further enhances reproducibility.
The primary limitation is the low "conceptual coherence" score (2.76/5), indicating that while the model plays well and structures its analysis correctly, it still hallucinates chess motifs or patterns. The method is currently specific to chess; while the authors argue for generality, the reliance on a specific "silent expert" encoder and the specific recursive distillation logic may not transfer trivially to other domains without significant adaptation.
This work offers a general recipe for coupling LMs with domain-specific expert encoders, with potential applications in robotics, computer use, and other games. It demonstrates that small, specialized models can outperform massive general-purpose LMs in specific reasoning tasks when properly grounded in expert knowledge. Queen introduces a 4B-parameter chess-language model that achieves Grandmaster-level play (2697 Elo) and high-quality explanations by integrating a silent expert encoder with an LM via cross-attention and iterative search distillation. This paper makes a significant contribution to the field by demonstrating that hybrid architectures combining specialized encoders with general LMs can outperform frontier LMs in domain-specific reasoning tasks, offering a scalable and efficient alternative to pure LLM approaches for tasks requiring deep domain expertise.
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/
Primary: Ben-Gurion University of the Negev
All Institutions: Ben-Gurion University of the Negev
GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
The paper proposes a novel architectural shift in neural audio codecs by replacing the standard Residual Vector Quantization (RVQ) or Finite Scalar Quantization (FSQ) bottleneck with a parametric 1D Gaussian Splatting decomposition. The core idea is to fit the encoder's latent representation as a weighted sum of Gaussian primitives via an inner optimization loop during training. To address the computational cost of this iterative fitting at inference, the authors introduce a "GS Predictor Net" that amortizes the optimization into a single forward pass. The method is technically sound, leveraging differentiable rendering concepts from 3D graphics (Gaussian Splatting) and applying them to 1D signal processing. The use of post-training scalar quantization on the fitted parameters rather than learned codebooks is a distinct design choice that enables fine-grained bitrate control without retraining.
The experimental evaluation is rigorous and comprehensive. The authors compare GS-Codec against strong, well-established baselines such as EnCodec, DAC, and WavTokenizer on standard datasets (LibriTTS, LJSpeech, LibriSpeech). The results show that GS-Codec matches or exceeds these baselines on key metrics like UTMOS (perceptual quality), STOI (intelligibility), and SIM (speaker similarity) at comparable bitrates. The inclusion of human listening tests (MUSHRA/MOS) adds significant credibility to the perceptual quality claims. The ablation studies on primitive count and bit depth provide clear insights into the rate-quality tradeoff. The encoding time analysis honestly reports the latency trade-off, showing that while the iterative version is slow, the Predictor Net brings it close to competitive levels, though still slightly slower than feed-forward baselines.
The paper provides high reproducibility. It details the SEANet backbone hyperparameters, the specific Gaussian Splatting configuration (number of primitives, inner loop steps, learning rates), and the training schedule. The code and audio samples are available via the provided URL. The use of standard metrics and open-source baselines facilitates easy comparison. The detailed appendix on quantization ranges and loss weights further supports reproducibility.
The primary limitation is the encoding latency. Even with the Predictor Net, the encoding time is higher than standard feed-forward codecs like EnCodec and DAC, which may be a bottleneck for real-time applications. The paper also notes that performance at very low bitrates (<3 kbps) is not as strong as specialized low-rate codecs. The method is currently validated primarily on English speech, and generalization to other languages or non-speech audio (music, sound effects) is not extensively explored.
This work opens a new direction in neural audio compression by demonstrating that parametric signal decompositions can serve as effective bottlenecks, challenging the dominance of codebook-based quantization. The ability to control bitrate post-training by varying the number of primitives and bit depth is a practical advantage for deployment. The cross-domain insight from 3D Gaussian Splatting to 1D audio signals is interesting and may inspire further research into using geometric or parametric priors in other signal processing tasks. GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
Real-time whole-body controllers for legged robots typically plan through a fixed nominal model and degrade when the deployed dynamics change. Adaptive methods typically require a model structure that contact dynamics do not provide, or they need offline training for each anticipated condition. We present Look-back and Look-ahead Adaptive Model Predictive Path Integral control (LLA-MPPI). The method converts whole-body adaptation into selection over a bank of GPU-batched contact simulators with different physical or structural parameters. Windowed prediction errors select the simulator that best explains recent motion. A whole-body MPPI planner optimizes controls through the selected model. The framework requires no offline training, and its selected hypotheses are physically interpretable. Across four simulated tasks, it achieves 97.5% success while the strongest baseline reaches 74% and an oracle with the true model reaches 98.5%. Hardware validation on a Unitree Go2 shows the robot walking under a payload added mid-run, walking after one leg is disabled, and pushing a box to its goal while increasing its mass on the fly. Code, videos, and project details are available at: https://lla-control.github.io
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology
LLA-MPPI introduces a novel adaptive control framework that leverages GPU-batched physics simulation to select the best-fitting dynamics model from a bank of hypotheses, enabling real-time whole-body control of legged robots under significant model mismatch without offline training. The paper demonstrates near-oracle performance in simulation and successful hardware validation, establishing a new standard for adaptive model-based control in contact-rich environments.
The paper proposes LLA-MPPI, a framework that decouples adaptive model identification from trajectory optimization. The core innovation is the use of a "bank" of GPU-batched physics simulators (via MuJoCo Warp) to perform model selection. Instead of using classical estimators (like Kalman filters or particle filters) that require specific linear or affine structures, the method treats adaptation as a classification problem over a discrete set of physical hypotheses. The "Look-back" stage runs on the GPU, evaluating thousands of candidate dynamics models in parallel against recent state transitions to select the best-fitting model. The "Look-ahead" stage runs on the CPU, using MPPI (Model Predictive Path Integral) control to plan trajectories using the selected model. This asynchronous, heterogeneous architecture is technically sound and cleverly leverages modern hardware capabilities (GPU parallelism for wide/shallow simulation, CPU threads for narrow/deep planning). The removal of the need for differentiable or closed-form dynamics is a significant methodological advantage for contact-rich robotics.
The evaluation is rigorous and comprehensive. The authors test the method on four distinct simulated tasks (asymmetric payload, leg lock, slope/friction variation, variable-mass box pushing) and validate on hardware (Unitree Go2). The comparison against an oracle (true model) shows the method achieves 97.5% success vs. 98.5% for the oracle, indicating near-perfect adaptation. Comparisons against baselines (Nominal MPPI, Disturbance Observer, Continuous Parameter Estimator) show substantial improvements in success rate and fall rate. The inclusion of a comparison with Domain Randomized PPO is valuable, highlighting the trade-off between offline training costs and online adaptability. The hardware results are particularly convincing, demonstrating real-world applicability in scenarios like mid-run payload addition and leg failure.
The paper provides a project URL with code, videos, and details. The specific hardware (RTX 2080 Ti, i9-9980XE) and software versions (MuJoCo 3.4.0, MuJoCo-Warp 1.12.0) are clearly stated. The hyperparameters for the model bank size, window size, and switching margins are detailed for each task. This level of detail supports high reproducibility.
The method relies on a finite, design-time hypothesis bank. If the true dynamics fall outside the discretized range of the bank, adaptation will fail. The paper acknowledges this and suggests future work on adaptive bank refinement. Additionally, the method requires significant GPU resources to maintain the bank of simulators, which may not be available on all robotic platforms. The distinguishability of models depends on excitation; if the robot's motion does not excite a particular uncertainty (e.g., friction when not slipping), the model selection may be ambiguous, though the hysteresis mechanism helps mitigate this.
This work has high potential impact in the field of legged robotics and adaptive control. By demonstrating that high-fidelity, contact-rich dynamics can be adapted to in real-time without offline training or complex analytical derivations, it opens the door to more robust and versatile robotic systems. The technique of using GPU-batched simulation for model selection could be extended to other domains involving complex, non-differentiable dynamics, such as soft robotics or multi-agent systems. LLA-MPPI introduces a novel adaptive control framework that leverages GPU-batched physics simulation to select the best-fitting dynamics model from a bank of hypotheses, enabling real-time whole-body control of legged robots under significant model mismatch without offline training. The paper demonstrates near-oracle performance in simulation and successful hardware validation, establishing a new standard for adaptive model-based control in contact-rich environments.
Animal evidence shows that precise voluntary movements arise from rotational neural population dynamics in motor cortex, but their physical effects remain unknown. We developed a robotic analog of biological motor systems with artificial muscles, multimodal sensors, and a neural network controller trained via reinforcement learning. The robotic analog exhibited accurate movements, robustness to damage, and neural population dynamics akin to animals. This task-driven, embodied model illuminates the causal link between neural population dynamics and motor outcomes. We discovered that neural rotations generate oscillatory maneuvers orthogonal to the reaching direction, optimizing trajectory adjustments, which is confirmed by primate neural data. The model also revealed counterintuitive neural energy principles under sensor and motor redundancies, and striking Eureka moments during motor learning, bridging biological and artificial systems. These findings provide new perspectives on how neural dynamics contribute to accurate and flexible movement, inspiring future intelligent robots with animal-like mobility.
Primary: Peking University
All Institutions: Peking University, State Key Laboratory of Transvascular Implantation Devices, The Second Affiliated Hospital of Zhejiang University School of Medicine, IDG/McGovern Institute for Brain Research, Peking-Tsinghua Center for Life Sciences, School of Advanced Manufacturing and Robotics, School of Life Sciences
The paper introduces a robotic analog of biological motor systems to causally link neural rotational dynamics to physical movement outcomes. By combining a biomimetic robot arm with an RNN controller trained via reinforcement learning, the authors demonstrate that rotational neural dynamics drive orthogonal trajectory adjustments, a finding validated in primate data, thereby providing a new embodied framework for decoding the functional role of neural population dynamics in motor control.
The paper proposes a "robotic analog" of biological motor systems, combining a continuum robot arm with liquid crystal elastomer (LCE) artificial muscles, multimodal sensors (vision, IMUs, self-sensing muscles), and a Recurrent Neural Network (RNN) controller trained via Soft Actor-Critic (SAC) reinforcement learning. The core methodological innovation lies in using this embodied, transparent system to causally link neural population dynamics (specifically rotational dynamics) to physical motor outcomes. The authors employ dimensionality reduction techniques like jPCA and CCA to analyze the latent dynamics of the RNN controller, comparing them to primate motor cortex data. The approach is rigorous, moving beyond simple simulation to physical embodiment to test hypotheses about neural dynamics that are difficult to test in vivo.
The experiments are extensive, covering simulation and real-world physical tests. The robot achieves sub-5mm reaching errors and demonstrates robustness to muscle failure without retraining. Key findings include the emergence of rotational dynamics in the RNN layers that drive orthogonal maneuvers, a phenomenon confirmed in primate data. The paper also analyzes neural energy efficiency under different sensory feedback conditions and identifies "Eureka moments" during learning where performance and dynamics shift abruptly. The validation against primate neural data strengthens the biological relevance of the findings.
The paper provides detailed descriptions of the hardware fabrication (LCE fibers, IMU setup) and the simulation environment (MuJoCo, constitutive models). However, specific code repositories or hyperparameters for the RL training are not explicitly provided in the text, which may limit immediate reproducibility. The physical setup is complex and likely expensive to replicate.
The system is limited to a single arm and relatively simple reaching tasks. The "neural" dynamics are those of an artificial RNN, not biological neurons, so while the analogy is strong, it is not a direct proof of biological mechanisms. The "Eureka moment" finding, while interesting, is based on a specific training trajectory and may not generalize to all RL algorithms or tasks.
This work bridges the gap between computational neuroscience and robotics, offering a new paradigm for studying motor control. It provides insights into how neural dynamics contribute to movement precision and flexibility, which could inspire more biologically plausible robot controllers. The findings on rotational dynamics and energy efficiency have implications for both understanding brain function and designing efficient robotic systems. The paper introduces a robotic analog of biological motor systems to causally link neural rotational dynamics to physical movement outcomes. By combining a biomimetic robot arm with an RNN controller trained via reinforcement learning, the authors demonstrate that rotational neural dynamics drive orthogonal trajectory adjustments, a finding validated in primate data, thereby providing a new embodied framework for decoding the functional role of neural population dynamics in motor control.
World action models jointly learn visual predictionand robot actions, providing a way to use observations ofscene evolution for policy learning. Their video and actionlosses, however, provide no explicit target for the geometricconsequences of a demonstrated action sequence. Moreover,visual features taken after temporal attention can contain futureobservations, making them unsuitable as the sole current visualinput to an auxiliary predictor. We introduce ACG-WAMand its auxiliary objective, the Action-Conditioned GeometricJoint-Embedding Predictive Architecture (ACG-JEPA), whichpredicts geometric features at several horizons from the currentobservation and intervening actions, using the future slot of afrozen VGGT encoding of each current and future image pairas the target. We apply this supervision from the head and wristcameras to a shared visual embedding before temporal mixing,and remove the teacher and auxiliary modules at inference.On 50 RoboTwin 2.0 tasks, ACG-WAM achieves 93.46%success in clean scenes, with the best randomized success(92.68%) and mean across both settings (93.07%) among thecompared methods; across three tasks on a real robot, itachieves 85.00% success and 91.67% partial completion score,exceeding Motus by 10.00 and 9.17 percentage points, respec-tively. Code:https://github.com/RoboOpus/ACG-WAM.Website:https://RoboOpus.github.io/ACG-WAM.
Primary: Beijing Institute of Technology
All Institutions: Beijing Institute of Technology, LimX Dynamics
ACG-WAM introduces an action-conditioned geometric prediction objective to World Action Models, significantly improving bimanual manipulation performance in simulation and on real robots. The paper presents a rigorous technical contribution that addresses specific challenges in joint video-action learning, such as information leakage and the lack of explicit geometric targets, resulting in state-of-the-art results on the RoboTwin 2.0 benchmark and real-world tasks.
The paper proposes ACG-WAM, a World Action Model (WAM) that augments the Motus backbone with an auxiliary geometric prediction objective. The core innovation is the Action-Conditioned Geometric Joint-Embedding Predictive Architecture (ACG-JEPA). This module uses a frozen VGGT teacher to encode current and future image pairs, extracting the "future slot" features as targets. The student predictor takes current visual features (crucially, extracted *before* temporal attention to avoid information leakage from future frames) and the intervening action sequence to predict these geometric targets. This design explicitly supervises the geometric consequences of actions, addressing a gap in standard video/action loss functions. The method is well-motivated by the need for spatial reasoning in bimanual manipulation and the technical challenge of preventing future information leakage in auxiliary predictors.
The evaluation is robust, covering 50 tasks on the RoboTwin 2.0 benchmark (clean and randomized settings) and 3 real-world tasks on a TRON2 platform. ACG-WAM achieves state-of-the-art results, outperforming strong baselines like Motus, LingBot-VA, and MECo-WAM. Specifically, it achieves 93.46% success in clean scenes and 92.68% in randomized scenes on RoboTwin 2.0, and 85.00% success on the real robot. Ablation studies convincingly demonstrate the contribution of action conditioning, multi-horizon prediction, and the joint-encoding target design versus simple endpoint subtraction. The real-world results are particularly strong, showing significant gains over the base Motus model.
The paper provides high reproducibility. Code is released on GitHub. Implementation details are thorough, specifying the backbone (Wan2.2-TI2V-5B), teacher model (VGGT), training hyperparameters (learning rates, batch sizes, loss weights), and inference settings. The use of standard benchmarks (RoboTwin 2.0) and clear evaluation metrics (Success Rate, Partial Completion Score) further supports reproducibility.
The method relies on a large, frozen teacher model (VGGT) and a complex backbone (Motus/Wan2.2), which may limit accessibility for groups with lower computational resources. The real-world evaluation is limited to only three tasks, which is a small sample size for generalizing claims about physical robustness. Additionally, the paper acknowledges the use of ChatGPT for drafting, which, while disclosed, is a minor point regarding the originality of the text rather than the science.
This work contributes to the growing field of World Action Models for robotics. By explicitly incorporating geometric supervision, it offers a pathway to improve spatial reasoning in manipulation policies. The technique of supervising features before temporal mixing to avoid leakage is a useful insight for other architectures involving joint video-action prediction. The strong performance on both simulation and real hardware suggests practical utility for industrial and research robotics applications. ACG-WAM introduces an action-conditioned geometric prediction objective to World Action Models, significantly improving bimanual manipulation performance in simulation and on real robots. The paper presents a rigorous technical contribution that addresses specific challenges in joint video-action learning, such as information leakage and the lack of explicit geometric targets, resulting in state-of-the-art results on the RoboTwin 2.0 benchmark and real-world tasks.
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)
Primary: Stanford University
All Institutions: Stanford University, UC Berkeley
EyeRobot 2.0 introduces a biologically inspired active gaze framework that enables precise bimanual manipulation using only a fixed stereo camera, outperforming standard wrist-camera baselines in occlusion-heavy tasks. The paper demonstrates that physically attending to 3D fixation points and canonicalizing actions in a fixation-relative frame significantly enhances policy learning, offering a robust alternative to the ubiquitous wrist-camera setup in robotic manipulation.
The paper proposes a hierarchical framework for active visual fixation in bimanual manipulation. The core innovation is the decoupling of low-level gaze servoing (trained with RL using geometric rewards) from high-level target selection (trained with RL using a BC-RL loop that optimizes for downstream gripper policy accuracy). The use of a fixation-relative SE(3) frame for action canonicalization is a strong technical contribution that reduces the complexity of the action space. The foveated processing of stereo images is well-motivated, though the specific implementation details of the "foveated transformer-decoder" are somewhat sparse in the main text, relying on the appendix. The method effectively addresses the occlusion issues inherent in wrist-mounted cameras by physically moving the viewpoint.
The experimental setup is rigorous, involving 7 real-world and 6 simulated tasks with over 1000 physical trials. The comparison against passive stereo and ego+wrist baselines is fair, as all policies are trained on the same data. The results are compelling: EyeRobot 2.0 significantly outperforms passive stereo and matches or exceeds ego+wrist performance, particularly in occlusion scenarios where wrist cameras fail. The ablation studies convincingly demonstrate the contribution of foveation, fixation-centric actions, and stereo depth. The inclusion of simulation experiments provides a controlled environment for analysis, although the primary claim is validated on real hardware.
The paper states that simulation code will be made public, which is a positive step. However, the reliance on specific hardware (I2RT YAM manipulator, custom leader arm) and specific software stacks (MuJoCo, GELLO, TRLC) may limit immediate reproducibility for labs without similar setups. The detailed description of the RL training loops and reward functions provides a clear path for implementation, but the lack of a public code release at the time of review (arXiv preprint) is a minor drawback.
The system is currently task-specific, requiring separate training for each task. The paper acknowledges that extending this to multi-task learning would require more complex prompt generation (e.g., VLMs). The framework assumes a fixed head position and only controls eye movements, which limits its applicability to mobile manipulation where head/neck movement is crucial. The computational cost of running two RL policies and a BC policy in real-time is not deeply analyzed, though the 30Hz gaze rate suggests it is feasible.
This work has significant implications for the design of robotic manipulators, potentially eliminating the need for wrist cameras and allowing for sleeker, more robust gripper designs. It also opens a pathway for leveraging egocentric human data for robot learning by mimicking human fixation patterns. The approach could be extended to other domains requiring precise visual attention, such as surgical robotics or inspection tasks. EyeRobot 2.0 introduces a biologically inspired active gaze framework that enables precise bimanual manipulation using only a fixed stereo camera, outperforming standard wrist-camera baselines in occlusion-heavy tasks. The paper demonstrates that physically attending to 3D fixation points and canonicalizing actions in a fixation-relative frame significantly enhances policy learning, offering a robust alternative to the ubiquitous wrist-camera setup in robotic manipulation.
New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and optimization. What agents produce depends on the interfaces they are given. These interfaces may expose a machine's raw capabilities or encode the known-good methods for exploiting its hardware features efficiently. We test the performance impact of these two interfaces on Intel GPUs through controlled experiments with three coding models on three kernels. With raw SYCL and execution feedback, Opus 4.8, the strongest of the three models tested, writes a single-GPU GEMM kernel reaching only 57.5% of Intel's tuned oneDNN library. We present SyclKittens, a hardware-aware tile programming model for Intel GPUs that encodes these known-good methods, so its operations place operands in matrix-engine layouts, prefetch through the L1 cache, and communicate over the fabric that links GPU stacks into a node. With SyclKittens under the same agent, task, and feedback budget, the agent reaches 82.1%, showing that encoded methods turn hardware capabilities into performance. Workload-specific policies stay programmable, and engineers and agents jointly refine these schedules in SyclKittens to build a kernel suite that reaches ~96% of oneDNN in geometric mean across GEMM shapes. The suite runs Llama-3.1-8B inference 1.59x faster than torch.compile on one GPU and up to 2.91x faster than a matched multi-GPU decode path on Intel's oneCCL. SyclKittens is open source and available at https://github.com/intel/SyclKittens.
Primary: Intel Corporation
All Institutions: Massachusetts Institute of Technology, Intel Corporation, Stanford University, California Institute of Technology
SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.
The paper introduces SyclKittens, a tile-based programming model for Intel GPUs that abstracts low-level hardware details (operand pipelines, data movement, matrix engine layouts) into high-level, hardware-aware operations. The core methodological contribution is the demonstration that providing coding agents with a structured, domain-specific interface (SyclKittens) rather than raw SYCL primitives significantly improves the performance of generated kernels. The authors employ a controlled experimental design comparing raw SYCL with execution feedback against SyclKittens under identical agent, task, and feedback budgets. This approach effectively isolates the impact of the programming interface on agent performance, providing a rigorous evaluation of how "known-good" hardware methods encoded in a DSL can guide LLMs to produce efficient code.
The experiments are conducted on Intel Max GPUs, focusing on critical AI workloads such as GEMM, attention, and normalization. The results show a substantial improvement: an agent using raw SYCL reaches only 57.5% of the performance of Intel's tuned oneDNN library, whereas the same agent using SyclKittens reaches 82.1%. Furthermore, a co-designed kernel suite using SyclKittens achieves ~96% of oneDNN performance in geometric mean across GEMM shapes. End-to-end inference benchmarks for Llama-3.1-8B demonstrate 1.59x speedup over torch.compile on a single GPU and up to 2.91x speedup over a matched multi-GPU decode path using Intel's oneCCL. The evaluation is strong in its direct comparison of interfaces and its end-to-end application metrics.
The paper is highly reproducible given that SyclKittens is open-sourced on GitHub. The authors provide detailed appendices separating the controlled-agent studies from the co-designed suite, including specific measurement protocols, warmup procedures, and aggregation methods. The use of specific coding models (Opus 4.8, etc.) and defined feedback budgets allows for precise replication of the agent experiments. However, the specific prompts and agent configurations may require careful extraction from the appendices to fully replicate the agent's behavior.
The evaluation is limited to Intel GPUs, which may not generalize directly to other hardware architectures (e.g., NVIDIA, AMD) without significant adaptation. The performance gains are relative to Intel's own oneDNN library, so the absolute competitive standing against other state-of-the-art libraries on different hardware is not assessed. The reliance on specific coding models means the results may vary with future model improvements or different agent architectures. Additionally, the paper focuses on a limited set of kernels (GEMM, attention, norms), and the applicability to more complex or irregular workloads is not fully explored.
This work has significant implications for the intersection of AI and systems programming. By demonstrating that high-level, hardware-aware abstractions can enable coding agents to write near-optimal GPU kernels, it suggests a new paradigm for kernel development that reduces the need for deep, architecture-specific expertise. This could accelerate the adoption of new AI accelerators by lowering the barrier to entry for kernel optimization. The open-source release of SyclKittens provides a valuable tool for the community to explore similar approaches on Intel hardware and potentially adapt the concepts to other platforms. SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.