Last 7 Days (September 24 – September 30, 2026)
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
Primary: Anthropic
All Institutions: Anthropic
Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
The paper proposes "Constitutional Adapters" (CAs), a method for distilling alignment principles into lightweight, inference-time interventions (LoRAs and ReLU steering vectors). The core methodological innovation is the use of a control-subtracted task arithmetic approach: training a lightweight object on a synthetic corpus derived from a model constitution ($D_C$) and subtracting an object trained on a register-matched control corpus ($D_K$). This subtraction aims to isolate the "constitutional" behavior from general fine-tuning artifacts like format or register shifts. The method leverages synthetic data generation via a strong model (Claude Opus 5) to create diverse scenarios that stress-test ethical principles, rather than relying on adversarial jailbreak examples. This allows the adapters to generalize to unseen attacks. The use of ReLU steering vectors, which allow for input-dependent steering strength, is a significant technical detail that improves upon fixed-direction steering.
The experimental evaluation is rigorous and comprehensive. The authors test across four distinct model families (Qwen3.5, Llama-3.1, Gemma-4, Nemotron 3) and various attack types (static, multi-query, multi-turn). Key findings include the superior robustness of CAs in long-context and multi-turn settings compared to system prompting, which degrades as context length increases. The paper also evaluates the trade-off between defense and benign compliance (over-refusal) by sweeping the steering strength, demonstrating a predictable and tunable frontier. The inclusion of misalignment evaluations (Petri Bloom) and capability benchmarks (MMLU, GSM8K, etc.) provides a holistic view of the method's impact, showing minimal capability degradation.
The paper provides detailed descriptions of the synthetic data pipeline, training hyperparameters, and evaluation protocols. It mentions that code and data will be released upon acceptance. The use of specific model versions (e.g., Claude Opus 5, Qwen3.5-4B) and clear definitions of metrics (StrongREJECT, over-refusal scoring) enhances reproducibility. However, the reliance on proprietary models for data generation (Claude) may limit full reproducibility for external researchers without access to those specific model versions.
The method is evaluated primarily on models up to 8 parameters, and it is unclear how well it scales to frontier-scale models (e.g., 70B+). The synthetic data generation relies on a strong, aligned model, which may introduce biases or limit the diversity of the training data. The paper does not extensively discuss the computational cost of generating the synthetic corpora or the potential for adversarial attacks specifically targeting the adapter weights. Additionally, the zero-shot transfer to post-trained checkpoints is a strong claim that might not hold for all post-training procedures.
This work has significant implications for the deployment of LLMs in API settings, where lightweight, tunable safety interventions are highly desirable. By providing a method to "turn up" safety without retraining or terminating sessions, CAs offer a practical solution for dynamic risk management. The approach also contributes to the broader field of AI alignment by demonstrating that complex, principle-based behaviors can be compressed into low-dimensional, portable objects. This could facilitate the development of modular safety systems that can be updated or adjusted independently of the base model. Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: Unknown (Likely ByteDance based on "YuE" naming convention and technical style, but not explicitly stated in provided text)
All Institutions: Unknown
YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
The paper proposes YuE2, a unified framework that integrates symbolic music generation (ABC notation) with audio generation (flow matching) within a single Mixture-of-Transformers (MoT) architecture. The core methodological contribution is the "symbolic planning" step, where the model first generates a readable score (melody, harmony, form) before expanding it into semantic tokens and acoustic latents. This is supported by two auxiliary models: MERT2, a music representation learner that sets new SOTA on MARBLE benchmarks, and SheetSage2, a full-song transcription model that generates the symbolic supervision signals required for training. The use of an AR-NAR MoT to handle both discrete symbolic/semantic tokens and continuous acoustic latents is a sophisticated architectural choice that allows for bidirectional attention in the acoustic stream while maintaining causal generation for the score.
The evaluation is extensive, covering automatic metrics (SongBench, SongEval, AudioBox) and expert listening tests. YuE2 outperforms public baselines and is competitive with proprietary systems like Suno v4.5/v5. The ablation study on symbolic planning is particularly strong, demonstrating that generating the score first significantly improves perceived musicality and overall quality compared to direct audio generation. The score-editing experiments show that the model can preserve unedited content while modifying specific sections, validating the utility of the symbolic interface.
The paper provides detailed architectural descriptions and training procedures. However, as a technical report from a likely industry lab, the full code and weights may not be immediately available to the public, which limits immediate reproducibility. The reliance on proprietary or large-scale datasets (346,000 hours of music) also poses a barrier for independent replication.
The model is large (3.58B parameters) and computationally expensive. The evaluation against proprietary systems is limited to expert listening and some automatic metrics, as direct API access for rigorous benchmarking is often restricted. The "best-of-8" selection strategy, while effective for benchmarking, may not reflect real-time interactive use cases where latency is critical.
This work bridges the gap between symbolic composition and audio production, offering a new paradigm for music generation that allows for explicit control over musical structure. The introduction of MERT2 and SheetSage2 as high-quality supervision tools will likely benefit the broader music AI community. The agentic editing capability suggests potential applications in professional music production workflows. YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
Primary: Anthropic
All Institutions: Anthropic
Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
The paper proposes "Constitutional Adapters" (CAs), a method for distilling alignment principles into lightweight, inference-time interventions (LoRAs and ReLU steering vectors). The core methodological innovation is the use of a control-subtracted task arithmetic approach: training a lightweight object on a synthetic corpus derived from a model constitution ($D_C$) and subtracting an object trained on a register-matched control corpus ($D_K$). This subtraction aims to isolate the "constitutional" behavior from general fine-tuning artifacts like format or register shifts. The method leverages synthetic data generation via a strong model (Claude Opus 5) to create diverse scenarios that stress-test ethical principles, rather than relying on adversarial jailbreak examples. This allows the adapters to generalize to unseen attacks. The use of ReLU steering vectors, which allow for input-dependent steering strength, is a significant technical detail that improves upon fixed-direction steering.
The experimental evaluation is rigorous and comprehensive. The authors test across four distinct model families (Qwen3.5, Llama-3.1, Gemma-4, Nemotron 3) and various attack types (static, multi-query, multi-turn). Key findings include the superior robustness of CAs in long-context and multi-turn settings compared to system prompting, which degrades as context length increases. The paper also evaluates the trade-off between defense and benign compliance (over-refusal) by sweeping the steering strength, demonstrating a predictable and tunable frontier. The inclusion of misalignment evaluations (Petri Bloom) and capability benchmarks (MMLU, GSM8K, etc.) provides a holistic view of the method's impact, showing minimal capability degradation.
The paper provides detailed descriptions of the synthetic data pipeline, training hyperparameters, and evaluation protocols. It mentions that code and data will be released upon acceptance. The use of specific model versions (e.g., Claude Opus 5, Qwen3.5-4B) and clear definitions of metrics (StrongREJECT, over-refusal scoring) enhances reproducibility. However, the reliance on proprietary models for data generation (Claude) may limit full reproducibility for external researchers without access to those specific model versions.
The method is evaluated primarily on models up to 8 parameters, and it is unclear how well it scales to frontier-scale models (e.g., 70B+). The synthetic data generation relies on a strong, aligned model, which may introduce biases or limit the diversity of the training data. The paper does not extensively discuss the computational cost of generating the synthetic corpora or the potential for adversarial attacks specifically targeting the adapter weights. Additionally, the zero-shot transfer to post-trained checkpoints is a strong claim that might not hold for all post-training procedures.
This work has significant implications for the deployment of LLMs in API settings, where lightweight, tunable safety interventions are highly desirable. By providing a method to "turn up" safety without retraining or terminating sessions, CAs offer a practical solution for dynamic risk management. The approach also contributes to the broader field of AI alignment by demonstrating that complex, principle-based behaviors can be compressed into low-dimensional, portable objects. This could facilitate the development of modular safety systems that can be updated or adjusted independently of the base model. Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges between components. We establish the first information-theoretic lower bounds on sample complexity for this target. For $p$ variables, maximum SCC size $s_{\max}$, and maximum external-parent count $d_B$, any estimator requires order $s_{\max}\log(ep/s_{\max})+d_B\log(ep/d_B)$ samples in the worst case over a regular model class. These bounds distinguish the costs of SCC membership and external-parent selection. Under principal invertibility and without correlation faithfulness, we establish a population block-exogeneity principle that identifies unknown root SCCs through residual independence and inclusion minimality. A sparse-adjustment characterization shows that small adjustment sets suffice to identify SCCs and their direct external parents, without regressing on all previously recovered variables. These characterizations yield BlockExo, which attains a structurally matching sample bound without knowing $s_{\max}$ or $d_B$ under suitable conditions. Simulations support the structural dependence of our sample bound and demonstrate BlockExo's sample-efficient recovery in comparisons with other methods for cyclic causal discovery.
Primary: Princeton University
All Institutions: Princeton University, Seoul National University
The paper establishes the first information-theoretic lower bounds for sample complexity in cyclic causal discovery and introduces BlockExo, an algorithm that achieves structurally optimal recovery. By decoupling the costs of SCC membership and external parent selection, the work provides a precise characterization of the statistical difficulty of learning feedback loops, offering a rigorous theoretical foundation that advances the field of causal inference beyond acyclic assumptions.
The paper proposes a rigorous theoretical framework for causal discovery in cyclic linear non-Gaussian (LiNG) models. The core contribution is the establishment of the first information-theoretic lower bounds on the sample complexity for recovering the condensation (SCC partition and inter-component edges). The authors introduce a "block-exogeneity" principle that allows for the identification of root SCCs via residual independence without requiring correlation faithfulness, a significant relaxation of standard assumptions. The proposed algorithm, BlockExo, utilizes sparse adjustment sets to identify SCCs and their external parents, achieving a sample complexity that matches the derived lower bounds structurally. The methodology is mathematically sound, leveraging the Darmois-Skitovitch theorem and Fano's inequality to bridge the gap between population-level identifiability and finite-sample guarantees.
The experimental section is limited but appropriate for a theory-heavy paper. It validates the structural dependence of the sample bound by showing that recovery curves align when normalized by the theoretical factors. It compares BlockExo against existing methods like Coarsening, DisjointCycles, and StableSpIn. BlockExo demonstrates superior sample efficiency in exact recovery tasks, particularly in overlapping cycle structures where baselines fail or require significantly more samples. However, the experiments are conducted on synthetic data only, with small dimensionality ($p=50$) and limited noise distributions, which restricts the generalizability of the empirical claims.
The paper provides a reproducibility statement indicating that code, configurations, and scripts will be released. The algorithmic details are clear, and the theoretical assumptions are explicitly stated. However, the lack of publicly available code at the time of review and the reliance on specific, potentially hard-to-tune thresholds (e.g., in the $PASS$ condition) may pose challenges for independent replication without the authors' implementation.
The primary limitation is the restriction to linear non-Gaussian models with causal sufficiency (no latent confounders). The method assumes principal invertibility, which, while standard, excludes certain unstable cyclic systems. The computational complexity of BlockExo is exponential in the sum of SCC size and external parent count ($p^{O(s_{max}+d_B)}$), which limits its applicability to dense or large-scale cyclic structures. Furthermore, the empirical evaluation lacks real-world data validation, leaving the practical utility of the method in complex biological or economic systems unproven.
This work provides a foundational statistical understanding of the costs associated with learning cyclic causal structures, a critical area in systems biology and economics. By establishing minimax optimality for condensation recovery, it sets a benchmark for future algorithms. The block-exogeneity principle offers a new tool for causal discovery that does not rely on faithfulness, potentially influencing the design of more robust causal inference methods in the broader machine learning community. The paper establishes the first information-theoretic lower bounds for sample complexity in cyclic causal discovery and introduces BlockExo, an algorithm that achieves structurally optimal recovery. By decoupling the costs of SCC membership and external parent selection, the work provides a precise characterization of the statistical difficulty of learning feedback loops, offering a rigorous theoretical foundation that advances the field of causal inference beyond acyclic assumptions.
Neural operators are typically trained in a supervised fashion, which requires a dataset to be generated with a classical solver. Training them physics-informed, i.e., purely from the governing equations, removes this large offline cost and allows fresh samples to be drawn at every optimization step, but has so far been limited to simplified problems and trails supervised training in accuracy. The obstacle is the ill-conditioning of physics-informed losses, which differential operators induce and which worsens as the discretization is refined. We therefore propose a preconditioned residual loss function and show mesh-independent conditioning for elliptic problems and greatly improved conditioning for saddle point problems. Realized through geometric and algebraic multigrid, the construction applies to linear and nonlinear equations, steady or time-dependent, on structured and unstructured meshes, is agnostic to the neural operator architecture, and adds no cost at inference. On the Poisson, Allen-Cahn and stationary Stokes equations, the resulting label-free training matches supervised training and is four to twenty-five times more accurate than previous physics-informed operator learning methods.
Primary: ETH Zurich
All Institutions: ETH Zurich
The paper proposes a preconditioned physics-informed loss function using multigrid methods to solve the ill-conditioning problem in neural operator training, achieving label-free accuracy that matches supervised learning. This is a significant technical contribution that bridges numerical linear algebra and deep learning, offering a practical and scalable solution to a long-standing optimization challenge in physics-informed machine learning.
The paper addresses a critical bottleneck in physics-informed neural operator (PINO) training: the severe ill-conditioning of residual losses induced by differential operators, which scales poorly with mesh refinement ($O(h^{-4})$ for elliptic problems). The authors propose a preconditioned least-squares loss function that incorporates multigrid preconditioners (geometric for structured grids, algebraic for unstructured) directly into the loss landscape. This approach is elegant because it decouples the preconditioning from the neural network architecture, requiring no changes to the model structure and adding zero cost at inference. The method generalizes to nonlinear problems (Allen-Cahn) and saddle-point problems (Stokes) by using approximate Jacobian inverses and block-diagonal preconditioners, respectively. The theoretical analysis correctly identifies that preconditioning the residual effectively transforms the Hessian of the loss to be mesh-independent or significantly better conditioned, thereby enabling efficient gradient-based optimization.
The experiments are rigorous and cover three distinct PDE classes: Poisson (linear elliptic), Allen-Cahn (nonlinear parabolic), and Stokes (saddle-point). The use of both FNO (Fourier Neural Operator) and GAOT (Geometry-Aware Operator Transformer) demonstrates architecture agnosticism. The results are compelling: the proposed method matches or surpasses supervised training accuracy while requiring no labeled data, and outperforms standard PINO and PI-DeepONet baselines by factors of 4 to 25. The "infinite data limit" experiment is particularly strong, showing that fresh sampling at each step removes the data-size bottleneck of supervised learning. The inclusion of unstructured meshes for the Stokes problem adds significant practical relevance, as many real-world PDEs do not admit structured grids.
The paper provides a public GitHub repository with code. The appendices contain detailed descriptions of the discretization, preconditioner construction (including specific multigrid parameters), and optimization hyperparameters. The use of standard libraries (neuraloperator, AMGX) and clear algorithmic descriptions (e.g., the V-cycle algorithm) ensures high reproducibility. The authors also provide a clear distinction between the "interpolated neural operator" and the raw network output, clarifying how the loss is computed.
The primary limitation is the dependence on the availability of an efficient preconditioner for the specific PDE class. For problems where multigrid or other fast solvers are not readily available (e.g., highly convection-dominated flows or complex wave equations), the method's advantage diminishes. Additionally, the method relies on an explicit finite element discretization of the residual, which may not be straightforward for all PDE formulations or black-box physics. The paper also notes that for the Stokes problem, the conditioning is only improved to $O(h^{-2})$, not $O(1)$, though this is still sufficient for convergence.
This work has high potential impact on the field of scientific machine learning. By enabling label-free training of neural operators that matches supervised performance, it removes the need for expensive offline data generation via classical solvers. This is particularly significant for physics foundation models and applications where generating high-fidelity training data is prohibitive. The technique is broadly applicable to any PDE where a fast linear solver or preconditioner exists, making it a valuable tool for the broader community of PDE surrogates. The paper proposes a preconditioned physics-informed loss function using multigrid methods to solve the ill-conditioning problem in neural operator training, achieving label-free accuracy that matches supervised learning. This is a significant technical contribution that bridges numerical linear algebra and deep learning, offering a practical and scalable solution to a long-standing optimization challenge in physics-informed machine learning.
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
Primary: Microsoft Research
All Institutions: Microsoft Research, New York University, Stanford University
The paper introduces "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs with minimal overhead. By shifting the burden of verification to the trainer via lightweight behavioral fingerprints and replay challenges, it addresses the reproducibility crisis in ML, enabling trusted baselines and scalable auditing of billion-parameter models.
The paper proposes "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs. The core innovation is shifting the burden of verification from the reader (who must reproduce expensive runs) to the trainer (who generates a lightweight, cryptographically signed "tape" of training steps). The method utilizes three key technical components: (1) a quasi-geometric parameter commitment using low-rank orthonormal projections to create fast behavioral fingerprints of model weights; (2) a non-geometric hash for optimizer states to ensure state continuity; and (3) a probabilistic replay challenge mechanism where the verifier randomly samples steps to verify against the declared training function. The system leverages stateless accelerator serializations (JAX/Torch) to seal the computation and uses white-box cryptography or zkVMs to anchor trust in the attestation engine. The theoretical contribution includes proofs of completeness and soundness, demonstrating that the probability of a malicious trainer successfully hiding invalid updates is bounded and amplifiable.
The authors validate the system on GPT-style models ranging from 100M to 2B parameters, testing both Data Parallel (DDP) and Fully Sharded Data Parallel (FSDP) configurations. Key results include a minimal overhead of 1.19% in speed and 1.11% in memory for a 2B parameter run, which is remarkably low for a security-critical system. The paper demonstrates high detection rates for data poisoning, parameter corruption, and adversarial optimization attacks against the sketch. The inclusion of a "self-regulating leaderboard" (dvbench.org) provides a practical ecosystem for adoption, allowing the community to submit and verify certified runs. The experiments convincingly show that the method scales to billion-parameter models without prohibitive cost.
The paper provides high reproducibility by releasing the code and a public leaderboard. The methodology is code-base agnostic, relying on standard serialization formats (StableHLO/XLA) which are widely supported. The detailed description of the cryptographic primitives (BLAKE3, ratcheted signatures) and the specific implementation details (Rust workers, Python facade) allow for independent verification. The open-source nature of the leaderboard further enhances reproducibility by providing real-world examples of certified runs.
The primary limitation is the reliance on a trusted attestation engine ($B_{att}$). While the paper argues for white-box cryptography or zkVMs, these are complex engineering solutions that may be vulnerable to side-channel attacks or implementation bugs not covered by the theoretical model. Additionally, the system assumes the training function $f$ is declared and bounded; it does not protect against logical errors in the training algorithm itself, only against deviations from the declared algorithm. The overhead, while low, is non-zero and may be significant for extremely latency-sensitive applications. The method is currently focused on pretraining and may require adaptation for fine-tuning or reinforcement learning workflows.
This work has the potential to fundamentally change how machine learning research is conducted and evaluated. By providing a technical solution to the reproducibility crisis, it enables standardized comparisons and trusted baselines, which are currently lacking in the field. The leaderboard initiative could foster a culture of verified results, reducing the impact of "slop" contributions and increasing the reliability of scientific claims. This could lead to more efficient use of computational resources, as researchers can trust published results without duplicating expensive training runs. The broader impact extends to AI safety and governance, as certified training runs provide an audit trail for model behavior and data usage. The paper introduces "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs with minimal overhead. By shifting the burden of verification to the trainer via lightweight behavioral fingerprints and replay challenges, it addresses the reproducibility crisis in ML, enabling trusted baselines and scalable auditing of billion-parameter models.
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
Primary: University of Chinese Academy of Sciences
All Institutions: University of Chinese Academy of Sciences, Meituan
The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
The paper proposes TGRL, a method that integrates temperature-based exploration into the RLVR training loop by treating temperature-induced diversity as a learnable signal. The core mechanism involves partitioning rollouts into low-temperature (reference) and high-temperature (exploration) groups. The method estimates "exploration gain" via the reward contrast between these groups and uses Jensen-Shannon (JS) divergence between temperature-scaled token distributions to allocate this gain as token-level credit. The theoretical analysis provides bounds on how JS divergence captures local sensitivity to temperature changes, justifying the credit allocation strategy. The approach is logically sound, bridging the gap between sampling diversity and policy gradient updates in a principled manner.
The experimental evaluation is extensive, covering 11 benchmarks across mathematical reasoning, code generation, and agentic tasks. The model sizes tested (Qwen3-4B, 14B, 32B) are relevant to current LLM research. The results show consistent improvements over strong baselines like GRPO and DAPO, with specific gains in CodeForces rating and LiveCodeBench Pass@16. The ablation studies effectively isolate the contributions of the temperature grouping and JS-based credit allocation. The wall-clock analysis demonstrating 36% faster convergence is a strong practical contribution.
The paper provides a public GitHub repository and detailed hyperparameters (temperatures, rollout budgets, warmup steps). The use of standard models (Qwen3) and public benchmarks enhances reproducibility. However, the specific implementation details of the JS divergence calculation and the exact handling of the mixed-group advantage normalization would require careful inspection of the code to fully replicate.
The method relies on the assumption that temperature scaling is a sufficient proxy for exploration diversity. It may not capture all forms of beneficial exploration, such as those requiring semantic shifts rather than just stochastic sampling. The computational overhead of computing JS divergence for every token in the high-temperature group could be significant for very long contexts, though the paper claims efficiency gains. The performance gains, while consistent, are moderate (e.g., 1.6% on math average), which may limit its adoption if simpler baselines are sufficient for many tasks.
This work contributes to the understanding of how to efficiently utilize rollout budgets in LLM post-training. By providing a mechanism to quantify and exploit exploration gain, it offers a path toward more sample-efficient RLVR training. The insights into token-level credit allocation based on distributional sensitivity could be applicable to other areas of LLM optimization, such as test-time compute scaling. The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis. Debate beats single-agent inference (3--7 points where tasks have headroom) but at matched budget conditions it ties or even loses to self-consistency sampling at 1.6$\times$ the wall-clock and 3.4$\times$ the token cost. Persona prompting reduces accuracy and a dose-response experiment over each model's full combinatorial persona space shows the cost is a persona tax, not a diversity tax: redundant personas hurt most, while maximally-diverse teams recover part of the loss. Furthermore, mixed-model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate's benefit comes from the first exchange of answers. We further identify a pervasive measurement hazard in which debate transcripts silently overflow serving context windows, whose correction alone moves our debate-versus-sampling comparison from $-1.8$ points to parity. Our results recast reported MAD gains as an ensemble-sampling effect and provide the budget-matched, contamination-checked baseline bar that future debate mechanisms should be required to clear.
Primary: Harvard University
All Institutions: Boston Children's Hospital, Harvard College, Harvard Medical School, Harvard School of Engineering and Applied Sciences, Cincinnati Children's Hospital
The paper rigorously refutes the hypothesis that cognitive diversity drives the benefits of multi-agent debate, demonstrating that gains are primarily due to ensemble sampling effects and identifying a critical measurement hazard related to context-window overflow. Through extensive experiments with 23 open-weight models and budget-matched controls, it provides a robust baseline for future research and highlights the inefficiency of standard debate protocols compared to simple self-consistency sampling.
The paper employs a rigorous experimental design to test the "cognitive diversity" hypothesis in Multi-Agent Debate (MAD). The methodology is strong, utilizing a large pool of 23 open-weight models across 11 vendor families to ensure generalizability. A key methodological strength is the use of generation-budget-matched controls (self-consistency sampling) rather than just comparing against single-agent inference, which isolates the effect of deliberation from the effect of increased compute. The introduction of a "persona tax" analysis via a dose-response experiment over the combinatorial persona space is a sophisticated approach to disentangling the cost of role-playing from the benefit of diversity. The identification of context-window overflow as a confounding variable is a critical methodological insight that adds significant value to the field's understanding of LLM evaluation pitfalls.
The experimental scale is impressive, with over 5,500 debate and control runs across five tasks. The results are consistent and statistically significant, showing that MAD gains are largely attributable to ensemble sampling effects rather than cognitive diversity. The finding that persona prompting reduces accuracy in small models is a surprising and important empirical result. The evaluation of heterogeneous model teams against majority voting controls further supports the conclusion that diversity does not inherently improve debate performance. The inclusion of a biography generation task with LLM judges adds depth, although the reliance on LLM judges introduces some subjectivity, which the authors acknowledge and mitigate with vendor-balanced panels.
The paper provides a high level of reproducibility. All models are open-weight, and the authors release the full budget-matched, contamination-checked grid as a baseline. The use of deterministic scripts for figure and table generation, along with released CSVs for scored outputs and persona ladders, ensures that the results can be verified. The explicit mention of fixed bootstrap seeds and the assertion that persona-team selection is exactly reproducible via programmatic checks further enhance reproducibility.
The primary limitation is the focus on small open-weight models. The authors acknowledge that larger frontier models might behave differently, particularly regarding the "persona tax," as they have more capacity to adopt unconventional cognitive frames. The study is also limited to specific tasks (math, knowledge, biography), and the results may not generalize to domains requiring novelty or synthesis. The reliance on LLM judges for the biography task, while mitigated, still introduces potential bias. Additionally, the study does not explore structured single-agent test-time methods or objectively verifiable domains like code generation.
This paper has significant broader impact by challenging a popular assumption in the LLM community regarding multi-agent debate. By demonstrating that MAD gains are largely an ensemble-sampling effect, it redirects research focus towards more efficient sampling strategies and rigorous budget-matched evaluations. The identification of context-window overflow as a measurement hazard is a valuable contribution that will likely influence future experimental designs. The paper provides a clear baseline for future debate mechanisms, encouraging more rigorous comparison against simple sampling controls. The paper rigorously refutes the hypothesis that cognitive diversity drives the benefits of multi-agent debate, demonstrating that gains are primarily due to ensemble sampling effects and identifying a critical measurement hazard related to context-window overflow. Through extensive experiments with 23 open-weight models and budget-matched controls, it provides a robust baseline for future research and highlights the inefficiency of standard debate protocols compared to simple self-consistency sampling.
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $α$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
Primary: Stanford University
All Institutions: Stanford University
The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
The paper introduces "rank confidence sequences," a rigorous statistical framework for evaluating model rankings in a sequential, anytime-valid manner. The core innovation lies in combining betting e-processes for pairwise comparisons with closed testing over weak orderings (permutations with ties). This approach allows the construction of confidence sets for the ranks of all models simultaneously, valid at any time step, without assuming independence between model scores on the same items (a common issue in benchmarking). The method handles both fixed finite benchmarks (with random item ordering) and superpopulation settings. The theoretical contribution is significant, providing finite-sample guarantees that hold under any stopping rule, which addresses the "peeking" problem inherent in real-time leaderboard monitoring. The use of integer programming for exact certification at scale is a clever computational addition, though the coNP-hardness of the general problem is acknowledged.
The experiments are well-designed to validate the theoretical claims. E1 demonstrates the failure of fixed-sample methods under repeated monitoring, showing a 30% error rate vs. the nominal 5%. E2 applies the method to real LLM leaderboard data (Open LLM Leaderboard), showing that while early certification is limited, it becomes robust as more items are revealed. E3 highlights the practical benefit of compute savings, showing that early stopping for specific models (e.g., top-3 certification) can save significant evaluation costs without compromising validity. E4 compares power against fixed-sample methods at pre-planned look times, showing comparable performance. The use of real-world LLM data adds substantial practical relevance.
The paper provides a GitHub link to the code. The algorithms are described in detail, including the betting strategies, the offset calculations, and the integer programming formulations. The specific parameters for the betting grid are provided. The experimental setups are described with references to public datasets. The level of detail suggests high reproducibility for researchers with statistical programming skills.
The method requires per-item scores in [0,1] and assumes a fixed set of models evaluated on common items. It does not handle adaptive item selection or models arriving over time. The computational cost, while manageable for typical leaderboard sizes (up to ~50 models), grows with the number of models and items, and the exact certification via integer programming can be complex to implement correctly. The method is primarily useful for ranking, not for estimating the absolute performance gap with high precision in early stages.
This work has high potential impact on the ML evaluation community. As leaderboards become more dynamic and expensive to run, the need for statistically valid, anytime-valid monitoring is critical. This paper provides a principled way to do so, potentially changing how benchmarks are reported and interpreted. It bridges the gap between statistical theory (e-processes, closed testing) and practical ML evaluation. The compute savings aspect is also a strong practical incentive for adoption. The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(π_θ\Vert π_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Primary: Unknown (Affiliations not explicitly listed in provided text, but authors Qinfeng Li, Wenqi Zhang, Guoqing Jiang, Liwei Chen, Xuanping Li, Zhiheng Qin, Yuntai Bao, Xuhong Zhang are associated with Alibaba Group / Qwen Team)
All Institutions: Alibaba Group (Inferred from author list and Qwen model usage)
The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The paper employs a rigorous empirical scaling analysis framework, adapting the methodology of reward model overoptimization studies (Gao et al., 2023) to the domain of On-Policy Distillation (OPD). The core methodological contribution is the characterization of OPD training dynamics as a function of the square root of token-level reverse KL divergence ($d$). The authors identify a "useful-transfer" regime where gold score increases linearly with $d$, followed by a noisy tail. They fit power laws to predict peak performance ($G_{peak}$) and transfer rates based on student/teacher scale and teacher quality. The inclusion of a theoretical derivation (Appendix) explaining why local KL geometry leads to linear transfer in $d$ adds significant depth, linking empirical observations to Fisher information geometry. The comparison between Vanilla-OPD and Delta-OPD (using policy shift contrasts) is well-motivated and clearly defined.
The experimental design is comprehensive, covering 25 teacher-student combinations across Qwen2.5 models (0.5B-14B) in weak-to-strong, same-base, and strong-to-weak configurations. The use of a single model family (Qwen2.5) is a limitation but allows for controlled isolation of scale effects. The evaluation on math reasoning (GSM8K/MATH) is standard but sufficient for this type of scaling study. Key findings include: (1) Peak student error is proportional to teacher remaining error, (2) Smaller teachers transfer better at matched scores (counter-intuitive and significant), (3) Bootstrapping does not improve over direct transfer from the smallest expert, and (4) Off-policy cold starts harm weak-to-strong transfer. The validation via leave-one-scale-out prediction is a strong methodological choice that demonstrates the predictive power of the fitted laws.
The paper provides detailed hyperparameters in the appendix and specifies the use of the `verl` framework. However, the reliance on a single random seed for all runs is a significant reproducibility weakness, although the authors justify this by citing standard practices in scaling law studies (Kaplan et al., Hoffmann et al.). The lack of seed variance estimation limits the confidence in the precise coefficients of the power laws, though the trends appear robust across the grid. The code and data are not explicitly linked in the provided text, which is a minor gap for immediate reproducibility.
The primary limitation is the restriction to a single model family (Qwen2.5) and a single task domain (math reasoning). The authors acknowledge that whether these scaling laws hold for other architectures, tasks, or post-training recipes is untested. The single-seed constraint means that stochastic variance in RL/OPD training is not captured, potentially masking instability in the "noisy tail" dynamics. The theoretical derivation assumes smoothness and differentiability that may not hold strictly in discrete token spaces or with clipping mechanisms, though the empirical fit is strong.
This paper has high practical impact for LLM practitioners. By providing predictive scaling laws for OPD, it enables engineers to estimate the outcome of distillation runs before committing significant compute resources. The finding that smaller teachers can be more effective than larger ones at matched scores challenges common assumptions about teacher quality and offers a cost-effective strategy for model family development. The negative result on bootstrapping is also valuable, preventing wasted effort on a seemingly intuitive but ineffective strategy. This work bridges the gap between theoretical scaling laws and practical post-training pipelines. The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.
Primary: NVIDIA
All Institutions: NVIDIA
The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
The paper employs a rigorous ablation study combined with theoretical analysis to deconstruct the Transolver architecture. The methodology is sound, isolating the "physics-attention" mechanism into slicing/deslicing, token attention, and pointwise MLPs. The theoretical contribution leverages the theory of Averaging Neural Operators (ANO) to prove that the slicing/deslicing component, even with a constant core (no attention), is sufficient for universal approximation of continuous operators. This provides a strong mathematical foundation for the empirical finding that the Transformer component is redundant. The implementation of a FlashAttention-style kernel for the slicing operation is a significant technical contribution, addressing the memory bottleneck of materializing slice weights.
The experiments are extensive, covering nine challenging 3D fluid dynamics benchmarks, including industrial-scale aerodynamics (DrivAerNet++, SHIFT-SUV, SHIFT-Wing, DrivAerML). The evaluation protocol is careful, using matched step budgets and identical data pipelines. The results clearly demonstrate that removing token attention does not degrade accuracy, while removing the global mixing (slicing) causes performance collapse. The efficiency gains from the new kernel are substantial, showing significant memory savings and speedups, particularly at large slice counts.
The paper provides detailed descriptions of the ablation variants, training protocols, and kernel implementation. The use of standard datasets and clear metric definitions (relative L1 error) supports reproducibility. The code for the FlashSlice kernel is described in detail, though a direct link is not provided in the text snippet, the description is sufficient for implementation.
The theory is based on expressivity (universality) and does not address generalization or optimization dynamics. The bounds are not sharp. The findings are specific to the Transolver architecture and may not generalize to other neural operators without further analysis. The paper acknowledges that the cost of removing attention is empirically zero, but the theory only makes it unsurprising, not derived from first principles.
This paper has high impact on the scientific machine learning community by clarifying the fundamental mechanisms of a widely used architecture. It challenges the assumption that self-attention is necessary for global mixing in operator learning, potentially leading to more efficient and simpler models. The efficient kernel implementation will benefit practitioners working with large-scale unstructured meshes. The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
Primary: The Ohio State University
All Institutions: The Ohio State University, RWTH Aachen University
Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
The paper introduces "belief geometry," a unified analytical framework to compare Transformers and State-Space Models (SSMs) in the context of in-context linear regression (ICLR). The methodology is rigorous, moving beyond empirical benchmarks to derive theoretical lower and upper bounds on cumulative Bayes regret. It decomposes sequential learning into three capabilities: evidence assembly, belief maintenance, and addressing. The authors define specific architectural classes (linear/softmax attention, fixed/selective SSMs) and prove sharp separations: SSMs are optimal for stationary belief maintenance (exponential kernels vs. uniform attention), SSMs have a memory advantage for positional assembly (convolution vs. context window), and softmax attention has an exponential width advantage for content addressing (selecting from retained tokens vs. storing candidates in state). The use of cumulative regret rather than terminal loss is a significant methodological improvement for analyzing sequential dynamics.
The experiments validate the theoretical predictions using practical architectures (LLaMA-type Transformers and Mamba-2). The authors test kernel alignment in single-layer models, showing that Transformers learn flatter kernels while Mamba-2 learns exponential ones, matching the theoretical optima. For the routing tasks (positional and content), they demonstrate that performance thresholds align with the theoretical resource requirements (e.g., convolution width for SSMs, context length for Transformers). The experiments effectively bridge the gap between the analytically tractable linear regression testbed and practical non-linear models, confirming that the architectural lessons (e.g., exponential width gap for addressing) hold in practice.
The paper includes a reproducibility statement and provides a link to an anonymous code repository. The experimental protocols, including hyperparameter sweeps and evaluation metrics, are detailed in the appendices. The use of synthetic tasks (Gaussian filtering, Beta-Bernoulli bandits, logistic bandits) ensures that the results are deterministic and reproducible without access to large-scale datasets.
The primary limitation is the reliance on linear regression and conjugate/non-conjugate bandit settings for the theoretical analysis. While the authors argue that the "belief geometry" extends to broader problems, the proofs are specific to these linear/quadratic loss structures. The "exponential width advantage" for attention in content addressing is a strong claim, but it relies on specific definitions of "addressing capacity" and compressed codebooks; real-world semantic addressing may not strictly follow this geometric separation. Additionally, the comparison focuses on representational capability (what the model *can* do) rather than optimization dynamics (what the model *learns* efficiently), though the kernel alignment experiments partially address this.
This paper provides a principled theoretical foundation for the ongoing debate between Transformers and SSMs. By identifying specific architectural mechanisms (softmax vs. selective transitions, context window vs. recurrent state) that confer advantages in different task regimes, it offers actionable insights for architecture design. The framework of "belief geometry" could be extended to other sequential tasks, such as language modeling or control, potentially guiding the development of hybrid architectures that leverage the strengths of both families. The finding that SSMs are inherently better at stationary belief maintenance while attention is better at content addressing helps explain empirical observations in long-context learning and associative recall. Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
Primary: Cornell University
All Institutions: Cornell University, University of Washington
The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
The paper introduces "Bellman distributional certificates," a novel theoretical framework for handling chance-constrained MDPs (CCMDPs). The core innovation is transforming the non-convex, trajectory-dependent chance constraint into a Bellman recursion over a discretized "remaining safety budget" state. This allows the use of standard model-based RL techniques (KL confidence sets) to certify safety with high probability without requiring a union bound over all possible policies or time steps. The method provides matching upper and lower bounds on sample complexity for deterministic policies, proving that the statistical cost of chance constraints is not inherently higher than expected-cost constraints under certain conditions (bounded successor support). Additionally, a model-free variance-reduced policy gradient algorithm is proposed for stochastic policies, offering finite-sample KKT-residual guarantees.
The experimental section is limited compared to the theoretical depth. It evaluates the method on a synthetic CCMDP and an IEEE 14-bus energy storage control benchmark. The results demonstrate that the Bellman-certified selector achieves better safety-performance trade-offs than a Markov-CMDP surrogate, particularly in reducing structural conservatism. However, the experiments are illustrative rather than exhaustive, lacking comparisons with state-of-the-art safe RL baselines (e.g., RCPO, PPO-Lagrangian) on standard continuous control benchmarks (MuJoCo, D4RL).
The paper provides detailed algorithmic descriptions and proofs in the appendix. However, specific hyperparameters, code availability, and implementation details for the "certified planning oracle" are not fully specified in the main text, making independent reproduction difficult without access to the authors' code. The reliance on a "certified planning oracle" as a black-box assumption limits immediate practical applicability.
1) The model-based guarantee is restricted to deterministic policies and assumes a fixed bound on successor support, which may not hold in high-dimensional continuous spaces. 2) The model-free result is local and may return "unresolved" if validation fails, lacking a global convergence guarantee. 3) The experimental validation is narrow, focusing on a single domain (energy storage) and synthetic tasks, without broad empirical validation on standard RL benchmarks. 4) The computational cost of the Bellman table grows with the discretization of the safety budget, which could be prohibitive for tight constraints.
This work provides a rigorous theoretical foundation for safe RL, addressing a critical gap in how probability-level safety constraints are handled. The "Bellman distributional certificate" concept could influence future work on risk-sensitive RL and constrained optimization. However, the strong assumptions (bounded support, deterministic policies for main bound) limit its immediate impact on practical, large-scale safe RL applications. The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.
Primary: Seoul National University
All Institutions: Seoul National University
[One sentence main contribution]. The paper establishes minimax optimality for Transformers in in-context learning on heterogeneous manifold mixtures by connecting oracle tangent local-polynomial estimators to a structure-informed Transformer architecture, providing a rigorous theoretical foundation for the model's ability to exploit local geometric structure.
The paper proposes a theoretical framework for analyzing Transformers in the context of in-context learning (ICL) on data with heterogeneous local geometry (mixtures of manifolds). The core methodological contribution is the derivation of minimax lower bounds for prediction under these complex geometric conditions and the construction of an oracle estimator (tangent local-polynomial) that matches these bounds. Crucially, the authors demonstrate that a specific class of Transformers (two-stage softmax with geometric preconditioners) can approximate this oracle estimator with negligible error, thereby establishing that Transformers can achieve minimax optimality in this setting. The approach is highly theoretical, relying on statistical learning theory and differential geometry rather than empirical model training.
The provided text is primarily theoretical, focusing on proofs and bounds. There is no mention of extensive empirical experiments, benchmarks, or real-world dataset evaluations in the abstract or the visible text fragments. The "experiments" are likely mathematical verifications of the bounds and the approximation capabilities of the proposed Transformer architecture. As a pure theory paper, its value lies in the rigor of the proofs rather than empirical performance metrics.
Reproducibility in the context of this paper refers to the verifiability of the mathematical proofs. The paper is 63 pages long, suggesting detailed appendices with full proofs. However, without access to the code or specific simulation scripts (if any exist to validate the theoretical bounds empirically), reproducibility is limited to the mathematical derivation. The lack of a provided code repository in the extracted information makes empirical validation difficult for external researchers.
The primary limitation is the gap between theory and practice. The conditions required for the minimax optimality (local separation, small-perturbation conditions, specific mixture structures) may be restrictive and not always satisfied by real-world high-dimensional data. Furthermore, the "oracle" nature of the estimator and the specific "structure-informed" Transformer design may not be directly implementable in standard large language model architectures without significant architectural modifications. The paper does not appear to provide empirical evidence that standard Transformers naturally learn these geometric preconditioners.
This paper contributes to the foundational understanding of why Transformers are effective for ICL, extending the theory beyond simple Euclidean or single-manifold settings. It provides a rigorous justification for the use of local polynomial approximations within Transformer layers. While the immediate practical impact on engineering may be limited, it offers valuable insights for the design of future architectures that explicitly account for data geometry, potentially influencing the development of more efficient and robust foundation models. [One sentence main contribution]. The paper establishes minimax optimality for Transformers in in-context learning on heterogeneous manifold mixtures by connecting oracle tangent local-polynomial estimators to a structure-informed Transformer architecture, providing a rigorous theoretical foundation for the model's ability to exploit local geometric structure.
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jointly by gradient descent to minimize the energy without training data. Conceptually, GS-DFT is 3D Gaussian splatting with the renderer replaced by quantum mechanics. We introduce two key solver components: adaptive density fitting with screening for efficient evaluation of two-electron integrals, and a regularized differentiable orthogonalization of the molecular orbitals. Empirically, the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters, converging systematically in energy, density, and nuclear forces. At equal parameter count, it captures the stretched-bond and anion physics that fixed bases only recover with specialized basis augmentation. The resulting solver exhibits quadratic peak memory scaling in the cloud size, allowing us to simulate systems of up to 2,742 atoms (10,406 electrons) without any modifications at triple-zeta scale using a single four-GPU node.
Primary: Unknown
All Institutions: Unknown
The paper introduces GS-DFT, a novel method using 3D Gaussian splatting to represent molecular orbitals in DFT, achieving high accuracy with fewer parameters and enabling large-scale simulations on GPU clusters. This represents a significant methodological advance in computational chemistry, combining differentiable rendering concepts with quantum mechanical solvers to overcome the limitations of fixed basis sets.
The paper proposes GS-DFT, a method that replaces fixed atom-centered basis sets in Density Functional Theory with a cloud of 3D Gaussian splats. The positions, shapes, and mixing coefficients of these Gaussians are optimized via gradient descent to minimize the electronic energy. This is a significant conceptual shift, framing the basis set optimization as a differentiable rendering problem where the "renderer" is the quantum mechanical solver. The introduction of adaptive density fitting with screening for two-electron integrals and a regularized differentiable orthogonalization scheme are critical technical contributions that enable the stability and efficiency of this optimization. The approach is end-to-end differentiable and does not require training data, distinguishing it from standard neural network surrogates for DFT.
The experimental results are impressive, claiming that the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters. The paper reports systematic convergence in energy, density, and nuclear forces. A key highlight is the ability to simulate systems of up to 2,742 atoms (10,406 electrons) on a single four-GPU node, demonstrating quadratic peak memory scaling. The comparison to fixed bases shows that GS-DFT captures stretched-bond and anion physics that typically require specialized basis augmentation in traditional DFT. However, the lack of specific benchmark details (e.g., which molecules, which functionals, comparison to specific state-of-the-art DFT codes like ORCA or Gaussian) in the provided abstract limits the full verification of these claims.
The paper describes the solver components (adaptive density fitting, regularized orthogonalization) which are crucial for reproducibility. However, without access to the full code or detailed hyperparameter settings (e.g., number of Gaussians, optimization schedule, regularization strength), independent reproduction may be challenging. The reliance on a specific hardware setup (four-GPU node) for the large-scale simulations also poses a barrier for smaller groups.
The primary limitation is the computational cost of the gradient descent optimization for the basis set itself, which may be prohibitive for very large systems despite the efficient solver. The method is currently demonstrated on molecular systems; its applicability to periodic boundary conditions (solids) is not discussed. The "quadratic peak memory scaling" is a significant improvement over cubic scaling of traditional DFT, but it still limits the system size compared to linear-scaling DFT methods. The lack of comparison to other learned basis set methods or neural network potentials is a gap.
This work has the potential to significantly impact computational chemistry and materials science by providing a flexible, high-accuracy basis set representation that scales better than traditional methods. It bridges the gap between differentiable rendering techniques (3D Gaussian Splatting) and quantum mechanics, opening new avenues for differentiable physics. The ability to simulate larger systems on standard GPU hardware could democratize high-accuracy DFT calculations. The paper introduces GS-DFT, a novel method using 3D Gaussian splatting to represent molecular orbitals in DFT, achieving high accuracy with fewer parameters and enabling large-scale simulations on GPU clusters. This represents a significant methodological advance in computational chemistry, combining differentiable rendering concepts with quantum mechanical solvers to overcome the limitations of fixed basis sets.
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.
Primary: Tencent
All Institutions: Tencent
KuaFu introduces an item-level, two-axis compression layer for user behavior sequences that enables billion-scale LLM-based profiling with significant cost and latency reductions. The paper demonstrates that by treating individual behavior items as independent, cacheable units and employing a multi-stage training recipe including hallucination-aware RL, it is possible to maintain or improve downstream task performance while drastically reducing computational overhead, validated by a 1.37% GMV lift in a major production environment.
The paper proposes KuaFu, a unified behavior-compression layer that treats individual behavior items as the minimal unit of compression. The core architectural contribution is a two-axis projector that compresses item embeddings along both the token axis (reducing sequence length) and the width axis (reducing dimensionality), allowing for significant storage and compute savings (10KB to 0.5KB per item). The training methodology is sophisticated, involving a four-stage curriculum: (1) reconstruction pre-training with a length-based curriculum to ensure convergence on long sequences, (2) post-training for compressed QA, (3) co-training of the compressor and decoder, and (4) hallucination-aware reinforcement learning (DAPO) specifically designed to penalize fabrication, omission, date misattribution, and broken logic. This multi-stage approach, particularly the explicit handling of hallucination types in the RL reward function, is a strong methodological contribution for LLM-based user modeling.
The experimental evaluation is extensive and convincing, combining offline benchmarks with large-scale online A/B testing. Offline, KuaFu outperforms state-of-the-art compression methods (SAC, EPL, ICAE) on MRQA benchmarks, showing superior fidelity at high compression ratios. On RecBench, the compressed 4B model outperforms the uncompressed 8B model, demonstrating the efficiency gains. The online A/B test on Tencent's platform is the strongest evidence of impact, showing a 1.37% GMV lift and significant throughput improvements (37-350% per-GPU QPM) while saving 190 GPUs. The ablation studies clearly demonstrate the necessity of the curriculum learning and the specific projector design.
Reproducibility is moderate. While the paper provides detailed descriptions of the architecture, training stages, and hyperparameters, the core training data consists of proprietary industrial behavior logs from Tencent, which cannot be released. The authors state that public datasets (MRQA, RecBench) are used for evaluation, allowing partial reproduction of the compression and understanding components. However, the specific industrial gains and the full training pipeline on proprietary data cannot be fully replicated by external researchers.
The primary limitation is the reliance on proprietary data, which limits external verification of the industrial claims. The compression ratio is currently fixed per task family, and the paper acknowledges that adaptation to sparser sequences and overly long items remains an area for improvement. Additionally, the scaling laws along data and parameter size have not been systematically validated.
This work has significant implications for the deployment of LLMs in industrial recommendation and advertising systems. By solving the context length and cost bottlenecks through item-level compression, it enables the use of large language models for real-time, billion-scale user profiling. The framework for evaluating compression fidelity (the layered intermediate evaluation) is also valuable for the broader community working on context compression for LLMs. KuaFu introduces an item-level, two-axis compression layer for user behavior sequences that enables billion-scale LLM-based profiling with significant cost and latency reductions. The paper demonstrates that by treating individual behavior items as independent, cacheable units and employing a multi-stage training recipe including hallucination-aware RL, it is possible to maintain or improve downstream task performance while drastically reducing computational overhead, validated by a 1.37% GMV lift in a major production environment.
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.
Primary: Alibaba Group
All Institutions: Alibaba Group
The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
The paper proposes a "closed-loop AI-for-AI" framework, which is a high-level architectural concept rather than a single novel algorithmic breakthrough. The core technical contributions are the integration of a data flywheel (AI-assisted task generation and curation), a hybrid training pipeline (SFT cold start + RL), and a runtime Harness (Skills/Memory). The most specific technical novelty is CARE (Competence-Aware Reward-and-Advantage Engineering), which adjusts reward shaping based on group success rates to prevent efficiency signals from dominating in saturated success groups. While the components (RL for agents, memory systems, data synthesis) are individually known, the systematic integration into a self-improving loop for mobile agents is a significant engineering and methodological contribution.
The evaluation is conducted on MobilePA-Bench, a large-scale benchmark (1,700+ tasks). The results show Qwen-Planner-Agent (27B) outperforming strong closed-source competitors like GPT-6 Astra and Claude Opus 5, as well as larger open-source models. The ablation studies effectively isolate the contributions of the Planner Model versus the Harness, demonstrating that the runtime context (Skills/Memory) provides substantial gains over the raw model. The efficiency analysis (cost per task) is a valuable addition, showing the agent is not just better but cheaper than frontier commercial APIs.
As a report from a major industry lab (Alibaba), the paper provides high-level architectural details but likely lacks the granular hyperparameters, exact prompt templates, and code releases typically required for full academic reproducibility. The "AI-for-AI" loop implies a complex, proprietary infrastructure for data generation and verification that is difficult to replicate externally. However, the clarity of the framework description allows for conceptual replication.
The primary limitation is the reliance on a proprietary, large-scale infrastructure for the "AI-for-AI" loop, making it difficult for smaller labs to verify the specific benefits of the data flywheel. The "closed-loop" claim is somewhat aspirational; the paper admits that human review is retained for critical decisions, meaning it is not fully autonomous. Additionally, the evaluation is heavily skewed toward the specific mobile planning domain, and while generalization is claimed, the non-mobile benchmarks are less detailed in the provided text.
This paper represents a significant step toward scalable agent development. By formalizing the use of AI to generate training data and diagnose failures for agent systems, it offers a roadmap for reducing the manual effort required to build robust LLM agents. The focus on mobile planning is highly relevant to current industry trends in on-device and cross-app automation. The CARE method offers a useful technique for RL training of agents where success rates vary widely across tasks. The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25\% and 50\% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
Primary: NVIDIA
All Institutions: NVIDIA
RAZOR introduces a training-free expert pruning method for MoE models that scores experts by functional replaceability using consensus residuals, accounting for survivor renormalization and router refill. The paper demonstrates that this geometric approach yields superior downstream performance and predictive fidelity compared to magnitude-based baselines across multiple large-scale MoE architectures, while also revealing nuanced trade-offs in generation behavior that task metrics alone miss.
The paper introduces RAZOR, a training-free expert pruning method for Mixture-of-Experts (MoE) models. The core innovation is the shift from scoring experts based on usage frequency or output magnitude (as in REAP or EAN) to "functional replaceability." This is achieved by calculating "consensus residuals," which measure the deviation of an expert's output from the original weighted mixture of active experts. The method derives an exact single-deletion identity that accounts for two critical dynamic effects often ignored in static pruning: survivor renormalization (how the weights of remaining experts adjust) and router-selected refill (how the router promotes a new expert to replace the pruned one). The derivation is mathematically sound, providing a local surrogate for the global distributional shift. The approach is computationally efficient, requiring only forward passes on calibration data without gradients or recovery training.
The evaluation is extensive, covering four distinct MoE backbones (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at two pruning budgets (25% and 50% expert removal). RAZOR consistently outperforms baselines (Frequency, EAN, REAP) in macro-average downstream performance across all eight model-budget settings. It also demonstrates superior predictive fidelity, measured by lower reverse KL divergence compared to REAP. A notable strength is the inclusion of generation behavior analysis (RQ3), which reveals that while task performance is retained, pruning can still affect diversity and termination patterns, providing a nuanced view of the trade-offs. The experiments are rigorous, using matched calibration sets and detailed ablations on scoring components (RMS vs. Mean, fixed-support vs. refill).
The paper provides high reproducibility. It includes detailed algorithmic pseudocode, specific hyperparameters for evaluation, and clear descriptions of the calibration data composition (Nemotron datasets). The implementation details regarding memory optimization (chunked scoring, layer-wise execution) are well-documented. However, the model checkpoints are not redistributable by the authors, which may limit independent verification for those without access to the specific proprietary or large-scale open models used.
The primary limitation is that the scoring is local and single-deletion based; it does not account for complex interactions when multiple experts are pruned simultaneously, nor does it guarantee that the "refill" candidate remains available if it is also pruned. The paper acknowledges that local output change is a surrogate, not a guarantee of global distributional fidelity. Additionally, the evaluation is limited to four specific model families, and the generalizability to other MoE architectures (e.g., those with different router mechanisms) is not fully established. The lack of measured serving latency or energy savings is a practical gap, as the method's utility is partly defined by compression efficiency.
This work contributes significantly to the field of model compression by providing a principled, geometry-aware method for MoE pruning. It challenges the common heuristic that "less used" experts are less important, showing that "less replaceable" experts are the critical ones to keep. This insight can guide future research in structured pruning and model distillation for sparse models. The finding that task retention does not equate to generation stability is also valuable for practitioners deploying pruned models in production, highlighting the need for multi-faceted evaluation metrics. RAZOR introduces a training-free expert pruning method for MoE models that scores experts by functional replaceability using consensus residuals, accounting for survivor renormalization and router refill. The paper demonstrates that this geometric approach yields superior downstream performance and predictive fidelity compared to magnitude-based baselines across multiple large-scale MoE architectures, while also revealing nuanced trade-offs in generation behavior that task metrics alone miss.
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
Synthetic Hospital introduces an open, verifiable, and physician-validated longitudinal EHR benchmark that resolves privacy and ground-truth barriers in clinical AI evaluation. By grounding synthetic records in standard medical ontologies and public educational material, the paper provides a rigorous, reproducible testbed that reveals significant gaps in current frontier models' ability to perform longitudinal clinical reasoning, offering a crucial resource for advancing reliable clinical AI systems.
The paper proposes a deterministic, five-stage pipeline to construct "Synthetic Hospital," a longitudinal EHR benchmark. The core innovation is the decoupling of clinical ground truth from narrative generation. By first constructing a structured medical knowledge graph grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) from public USMLE-style board questions, and then rendering this graph into realistic clinical narratives using an LLM (Kimi 2.5), the authors ensure that every diagnosis and finding has a verifiable provenance chain. This addresses the two major barriers to open clinical benchmarks: privacy (no PHI) and ground truth ambiguity (real charts reflect documentation, not necessarily patient state). The methodology includes rigorous validation steps, such as a blinded physician study to assess realism and a pre-registered reference set to validate ontology-derived relationships.
The evaluation is comprehensive. It includes a realism study where physicians could not distinguish synthetic from real records (53% accuracy, near chance). It evaluates 10 frontier and open models across four tasks (patient diagnosis, summarization, retrieval, imaging indication). Key findings include that no model approaches ceiling performance, the best model achieves a severity-weighted F1 of 0.73 on diagnosis (matching the mean of seven physicians but below the best), and that agentic multi-turn loops often degrade performance compared to single-turn inference when context is available upfront, except for longitudinal diagnosis tasks. The paper also provides a robustness analysis showing that re-rendering the corpus with a different LLM (GPT-5.3) does not significantly change model rankings, mitigating concerns about generator bias.
High. The paper provides a detailed description of the pipeline, including specific thresholds for ontology mapping, clustering rules, and prompting strategies. Code and data are openly available on GitHub. The use of deterministic functions for benchmark labels and the release of a training split with verifiable rewards enhances reproducibility and utility for reinforcement learning or fine-tuning.
The benchmark is derived from medical education material (USMLE-style questions), which may not fully capture the complexity, noise, and atypical presentations of real-world clinical practice. The case mix is education-derived by design, potentially under-representing rare conditions. The agentic evaluation is limited to three models and one scaffold. The paper acknowledges that it does not explicitly simulate missingness or documentation errors, which are common in real EHRs.
This benchmark has high potential impact on the clinical AI community. By providing an open, verifiable, and realistic longitudinal EHR dataset, it enables the development and evaluation of clinical AI systems without the legal and ethical hurdles of using real patient data. The finding that current frontier models struggle with longitudinal synthesis and that agentic approaches have mixed utility provides actionable insights for system designers. The open nature of the data facilitates broader research and standardization of evaluation metrics in clinical NLP. Synthetic Hospital introduces an open, verifiable, and physician-validated longitudinal EHR benchmark that resolves privacy and ground-truth barriers in clinical AI evaluation. By grounding synthetic records in standard medical ontologies and public educational material, the paper provides a rigorous, reproducible testbed that reveals significant gaps in current frontier models' ability to perform longitudinal clinical reasoning, offering a crucial resource for advancing reliable clinical AI systems.
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
Primary: UC Berkeley
All Institutions: UC Berkeley, Stanford University, University of Washington, Impossible Research, Google DeepMind
The paper introduces OmniTaskonomy, a comprehensive taxonomy and empirical study demonstrating that visual generation tasks can significantly improve visual understanding capabilities through selective and task-dependent transfer. By systematically analyzing 19 generation tasks and 25 understanding capabilities, the authors reveal both intuitive and surprising connections, supported by gradient alignment analysis, providing a rigorous roadmap for leveraging generation supervision in multimodal learning.
The paper proposes a systematic framework, OmniTaskonomy, to analyze the transferability of visual generation tasks to visual understanding tasks. The methodology is rigorous, employing controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that share underlying semantic structures. By constructing a unified taxonomy of 19 generation tasks and 25 understanding capabilities, the authors move beyond anecdotal evidence to a comprehensive empirical study. A key methodological strength is the use of gradient alignment analysis to correlate with downstream performance gains, providing a mechanistic explanation for the observed transfer patterns.
The experimental evaluation is extensive, covering a wide range of tasks and datasets. The paper demonstrates that under specific training recipes, I2I training significantly improves I2T performance, with gains scaling with data volume. The transfer map reveals both intuitive connections (e.g., depth prediction aiding 3D reasoning) and surprising ones (e.g., 2.5D segmentation aiding category recognition). The results are consistent and well-supported by statistical analysis, providing a clear roadmap for curriculum learning in multimodal models.
The paper includes a detailed reproducibility statement, with appendices providing task definitions, training recipes, taxonomy annotations, and complete transfer results. The availability of numerical exports and figure-generation code further enhances reproducibility. The use of standard benchmarks and clear experimental protocols ensures that the findings can be verified by the community.
The study is primarily focused on image-based tasks and may not fully generalize to video or other modalities. The analysis relies on specific model architectures and training setups, which might limit the generalizability of the findings to other model families. Additionally, the computational cost of training and evaluating such a large number of task pairs is significant, which could be a barrier for some researchers.
The findings have significant implications for the design of multimodal learning systems, suggesting that visual generation can serve as a powerful pre-training signal for visual understanding. This could lead to more efficient and effective training curricula for large multimodal models. The paper also provides a valuable resource for the community in the form of the OmniTaskonomy taxonomy, which can be used to guide future research on task transfer and curriculum learning. The paper introduces OmniTaskonomy, a comprehensive taxonomy and empirical study demonstrating that visual generation tasks can significantly improve visual understanding capabilities through selective and task-dependent transfer. By systematically analyzing 19 generation tasks and 25 understanding capabilities, the authors reveal both intuitive and surprising connections, supported by gradient alignment analysis, providing a rigorous roadmap for leveraging generation supervision in multimodal learning.
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Primary: Tsinghua University
All Institutions: Tsinghua University, Shanghai Jiao Tong University, Fudan University, University of Science and Technology of China, Washington University in St. Louis, The Institute of Artificial Intelligence, China Telecom (TeleAI)
V-JEPA Policy demonstrates that frozen predictive visual latents are a sufficient foundation for effective world-action models, outperforming generative and discriminative alternatives under distribution shifts. The paper rigorously validates this approach through extensive simulation benchmarks and real-world manipulation, highlighting the critical role of predictor pretraining on video-instruction pairs for robust generalization.
The paper proposes V-JEPA Policy, a framework that decouples the predictive visual latent space from the generative model. Instead of fine-tuning a large video generator, it freezes the V-JEPA 2.1 encoder and trains a lightweight instruction-conditioned future-latent predictor and a flow-matching action expert from scratch. The key architectural innovation is the use of the predictor's future-informed context key-value states to condition the action generation, effectively creating a world-action model (WAM) that relies on predictive latents rather than pixel-space reconstruction or full video generation. This approach is computationally efficient (0.9B total params, 0.6B trainable) and leverages the robustness of predictive pretraining.
The experiments are extensive, covering LIBERO, LIBERO-Plus, and RoboCasa-GR1 benchmarks. The paper provides rigorous ablations comparing V-JEPA latents against discriminative, reconstructive, and video-understanding-oriented alternatives, showing superior performance under distribution shifts. A significant finding is the benefit of pretraining the predictor on DROID video-instruction pairs without action labels, which yields substantial gains in downstream control and out-of-distribution generalization that cannot be recovered by simply training longer from scratch. The results are competitive with state-of-the-art WAMs and VLA baselines.
The authors provide a public GitHub repository with code. The paper details the training setup, including the separation of frozen and trainable parameters, and the specific pretraining data (DROID). The use of standard benchmarks and open-source models (V-JEPA 2.1) enhances reproducibility.
The method relies heavily on the quality of the V-JEPA 2.1 encoder; if the encoder fails to capture relevant dynamics, the policy will suffer. The flow-matching action expert is trained from scratch, which may require careful tuning. The paper focuses on simulation and limited real-world bimanual manipulation; broader real-world deployment across diverse environments is not fully explored. The reliance on DROID for pretraining limits the immediate applicability to domains similar to DROID.
This work provides a compelling argument for using predictive latents as a foundation for robot learning, potentially reducing the computational cost of training WAMs. It offers a pathway to leverage large-scale video pretraining without the burden of generative model fine-tuning. The findings on the transferability of future-modeling knowledge from in-the-wild videos to control tasks are significant for the robotics community. V-JEPA Policy demonstrates that frozen predictive visual latents are a sufficient foundation for effective world-action models, outperforming generative and discriminative alternatives under distribution shifts. The paper rigorously validates this approach through extensive simulation benchmarks and real-world manipulation, highlighting the critical role of predictor pretraining on video-instruction pairs for robust generalization.
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Primary: University of California, Los Angeles (UCLA)
All Institutions: University of California, Los Angeles (UCLA)
The paper introduces a novel RL framework for unified multimodal models that jointly optimizes reflection and generation, achieving significant improvements in image generation accuracy through self-correction. By leveraging group-relative advantage estimation and graded rewards, the method effectively addresses the sparsity of feedback in compositional generation tasks, demonstrating that unified models can learn to diagnose and repair their own errors without external critics, with gains that transfer robustly to unseen benchmarks.
The paper proposes UMM-Reflection, a framework for applying Reinforcement Learning (RL) to unified multimodal models (specifically BAGEL) to improve image generation via self-reflection. The core methodological contribution is the joint optimization of reflection text (diagnosis) and image generation (repair) within a single model using group-relative advantage estimation. By sampling sibling trajectories from the same initial image, the method avoids the combinatorial explosion of per-round credit assignment and allows credit to flow across the entire reflection loop. A significant technical detail is the "graded reward" mechanism, which addresses the sparsity of binary GenEval scores by providing partial credit for partial constraint satisfaction, thereby stabilizing RL training. The approach is theoretically sound, leveraging the unified nature of the model to eliminate the need for external critics or separate verifier models at inference time.
The experimental evaluation is rigorous and comprehensive. The paper reports a substantial improvement of +12.05 points on GenEval over the SFT baseline. Crucially, the authors demonstrate that these gains transfer to unseen benchmarks (WISE, OneIG-Bench, T2I-CompBench++), suggesting the model has learned generalizable repair strategies rather than overfitting to the training distribution. Ablation studies are thorough, including comparisons against direct T2I-RL (without reflection), external critic pipelines (GPT-5.5), and different reward shaping strategies (penalty vs. bonus for stopping). The analysis of "pass@16" vs "pass@1" highlights that the SFT model already contains the capability for correct repairs, and RL effectively selects and reinforces these high-success paths. The visual pathway stability analysis confirms that RL primarily modifies the decision-making (text) pathway rather than disrupting the underlying visual generation capabilities.
The paper provides high reproducibility standards. It details the specific training hyperparameters (learning rates, batch sizes, number of updates), the composition of the RL prompt pool, and the exact reward calculation formulas. The authors explicitly state that the code and prompt pool will be released. The use of standard benchmarks (GenEval, WISE, etc.) and clear evaluation protocols (50 denoising steps, specific resolutions) allows for easy comparison with other works. The distinction between controlled evaluations and native leaderboard protocols is clearly defined, preventing misinterpretation of results.
The primary limitation is the reliance on the BAGEL model architecture; while the method is general, the specific implementation details (e.g., MoT decoder layers, flow-based revisions) are tied to this unified model structure. The counting category in GenEval shows no improvement, indicating that current reflection strategies are insufficient for precise numerical constraints. Additionally, the method requires significant computational resources (16 H100 GPUs for 33 hours) for the RL phase, which may limit adoption for smaller labs. The evaluation is limited to 512x512 resolution for training and evaluation, whereas the model natively supports 1024x1024, leaving open questions about performance at higher resolutions.
This work has significant implications for the development of autonomous multimodal agents. By demonstrating that a single model can effectively self-correct its generations through RL, it paves the way for more robust and reliable generative systems that do not rely on external, potentially misaligned, critic models. The insight that RL can "select" correct repairs from an existing SFT distribution rather than generating them from scratch is a valuable finding for the broader field of generative AI. The transferability of gains to unseen benchmarks suggests that self-reflection is a generalizable capability, potentially applicable to other modalities or tasks. The paper introduces a novel RL framework for unified multimodal models that jointly optimizes reflection and generation, achieving significant improvements in image generation accuracy through self-correction. By leveraging group-relative advantage estimation and graded rewards, the method effectively addresses the sparsity of feedback in compositional generation tasks, demonstrating that unified models can learn to diagnose and repair their own errors without external critics, with gains that transfer robustly to unseen benchmarks.
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Primary: Tencent
All Institutions: Tencent
The paper systematically characterizes the scaling laws of encoder-free MLLMs, predicting they will match encoder-based models at $10^{22}$ FLOPs. It provides a rigorous comparison using a matched model ladder and offers mechanistic insights into how decoders compensate for the lack of a visual encoder, positioning encoder-free architectures as a promising direction for future multimodal pretraining.
The paper employs a rigorous scaling law framework to compare encoder-free and encoder-based Multimodal Large Language Models (MLLMs). The methodology is sound, utilizing a matched ladder of 11 sparse MoE models (1.1B-44B) to isolate the effect of the visual encoder. The use of IsoFLOP profiles and compute-optimal allocation analysis is standard and well-executed. The introduction of "vision-specific adaptation" probes (attention patterns, layerwise representation evolution, expert routing) to mechanistically explain the scaling differences is a strong methodological contribution that goes beyond simple loss comparison. The derivation of the overtraining loss equation is mathematically sound and provides a useful tool for predicting efficiency gains under non-optimal training regimes.
The experimental setup is robust, covering a wide range of model scales and compute budgets ($10^{19}$ to $10^{21}$ FLOPs). The comparison is controlled by fixing the visual encoder size and data mixture. The results are consistent, showing that while encoder-free models lag at small scales, their loss decreases more rapidly with compute, predicting a crossover at $10^{22}$ FLOPs. The analysis by topic (STEM vs. Perception) provides valuable nuance, showing that the crossover is earlier for language-heavy tasks. The inclusion of downstream benchmark evaluations (CV-Bench, ChartQA, etc.) further validates the scaling trends, although the primary focus remains on validation loss.
The paper provides detailed implementation details in the appendix, including the specific architecture of the front ends, the MoE topology, and the training hyperparameters. The use of standard components (SigLIP 2, Muon optimizer) and clear descriptions of the data mixture enhance reproducibility. However, the specific data sources are not fully detailed, which may limit exact replication. The code is not explicitly linked in the provided text, but the level of detail suggests it is likely available or easily implementable.
The primary limitation is the reliance on extrapolation. The predicted crossover at $10^{22}$ FLOPs is well beyond the measured range ($10^{21}$ FLOPs), introducing uncertainty. The assumption that the visual encoder size remains fixed as the decoder scales is a simplification; in practice, joint scaling of the encoder and decoder might alter the dynamics. Additionally, the study focuses on a specific data mixture (1:1 text/multimodal), and results may vary with different data compositions. The "catch-up" prediction assumes that the irreducible loss floor is shared, which is a reasonable but unproven assumption at these scales.
This paper has significant implications for the design of future multimodal models. By demonstrating that the advantage of pretrained visual encoders diminishes with scale, it encourages the exploration of unified, encoder-free architectures that may be simpler and more efficient in the long run. The insights into how decoders adapt to take over visual encoding (e.g., bidirectional attention, early-layer processing) could inform the design of new decoder architectures specifically optimized for native multimodal learning. This work could shift the field's focus from improving visual encoders to optimizing the language model's ability to process raw visual tokens. The paper systematically characterizes the scaling laws of encoder-free MLLMs, predicting they will match encoder-based models at $10^{22}$ FLOPs. It provides a rigorous comparison using a matched model ladder and offers mechanistic insights into how decoders compensate for the lack of a visual encoder, positioning encoder-free architectures as a promising direction for future multimodal pretraining.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
[One sentence main contribution]. The paper presents a differentiable and GPU-accelerated acoustic simulator for the vocal tract that enables the reconstruction of vocal tract shapes from speech, with applications in linguistics and medical imaging. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant due to the novel frequency-domain formulation and the integration of turbulence modeling. The methodology is well-designed, leveraging neural networks to overcome the non-convexity of the inverse problem. The significance to the field is high, as it bridges the gap between physical simulation and deep learning, enabling new capabilities in speech analysis and medical imaging.
The paper introduces a differentiable acoustic simulator for the vocal tract, which is a significant technical achievement. The core innovation lies in the frequency-domain formulation of fluid dynamics, which is claimed to be 70x more GPU-parallelizable than time-domain finite differences. This is a crucial engineering insight for enabling real-time or near-real-time gradient-based optimization. The integration of a differentiable turbulence model for consonants is a strong addition, as consonants are notoriously difficult to model with simple tube acoustics. The use of neural networks to parameterize the geometry (similar to INRs) is a clever application of existing techniques to a new domain, helping to escape local minima in the non-convex inverse problem. The methodology is sound and well-motivated, combining physical simulation with modern deep learning techniques.
The experiments demonstrate the simulator's ability to reconstruct vocal tract shapes from speech across 11 languages, which is a robust test of generalizability. The self-supervised autoencoding task shows the model can learn meaningful representations without labeled data. The most impressive result is the coupling with a generative MRI model to reconstruct moving vocal tracts from speech alone, without paired data. This is a novel application that bridges speech processing and medical imaging. However, the evaluation of the MRI reconstruction quality is likely limited by the lack of ground truth for moving vocal tracts, and the paper should provide more quantitative metrics on the fidelity of the reconstructed shapes compared to actual MRI scans.
The paper provides a supplementary material link, which is a good sign for reproducibility. The frequency-domain formulation and turbulence model are described in detail, but the specific implementation details for the neural network parameterization and the MRI generative model coupling would be critical for reproduction. The claim of 70x speedup should be backed by detailed benchmarking against standard time-domain solvers.
The acoustic tube model is a simplification of the actual vocal tract, which is a 3D structure. The model may not capture all the nuances of speech production, especially for complex consonants or pathological speech. The MRI reconstruction is a novel application, but the clinical utility of the reconstructed shapes is not evaluated. The paper also does not discuss the computational cost of the MRI generative model coupling, which could be a bottleneck.
This work has significant potential impact in both linguistics and medical imaging. In linguistics, it could enable new insights into speech production and variation across languages. In medical imaging, it could lead to new methods for diagnosing speech disorders or planning surgical interventions. The differentiable simulator could also be used in other domains where inverse problems with physical constraints are important. [One sentence main contribution]. The paper presents a differentiable and GPU-accelerated acoustic simulator for the vocal tract that enables the reconstruction of vocal tract shapes from speech, with applications in linguistics and medical imaging. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant due to the novel frequency-domain formulation and the integration of turbulence modeling. The methodology is well-designed, leveraging neural networks to overcome the non-convexity of the inverse problem. The significance to the field is high, as it bridges the gap between physical simulation and deep learning, enabling new capabilities in speech analysis and medical imaging.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: Unknown (Likely ByteDance based on "YuE" naming convention and technical style, but not explicitly stated in provided text)
All Institutions: Unknown
YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
The paper proposes YuE2, a unified framework that integrates symbolic music generation (ABC notation) with audio generation (flow matching) within a single Mixture-of-Transformers (MoT) architecture. The core methodological contribution is the "symbolic planning" step, where the model first generates a readable score (melody, harmony, form) before expanding it into semantic tokens and acoustic latents. This is supported by two auxiliary models: MERT2, a music representation learner that sets new SOTA on MARBLE benchmarks, and SheetSage2, a full-song transcription model that generates the symbolic supervision signals required for training. The use of an AR-NAR MoT to handle both discrete symbolic/semantic tokens and continuous acoustic latents is a sophisticated architectural choice that allows for bidirectional attention in the acoustic stream while maintaining causal generation for the score.
The evaluation is extensive, covering automatic metrics (SongBench, SongEval, AudioBox) and expert listening tests. YuE2 outperforms public baselines and is competitive with proprietary systems like Suno v4.5/v5. The ablation study on symbolic planning is particularly strong, demonstrating that generating the score first significantly improves perceived musicality and overall quality compared to direct audio generation. The score-editing experiments show that the model can preserve unedited content while modifying specific sections, validating the utility of the symbolic interface.
The paper provides detailed architectural descriptions and training procedures. However, as a technical report from a likely industry lab, the full code and weights may not be immediately available to the public, which limits immediate reproducibility. The reliance on proprietary or large-scale datasets (346,000 hours of music) also poses a barrier for independent replication.
The model is large (3.58B parameters) and computationally expensive. The evaluation against proprietary systems is limited to expert listening and some automatic metrics, as direct API access for rigorous benchmarking is often restricted. The "best-of-8" selection strategy, while effective for benchmarking, may not reflect real-time interactive use cases where latency is critical.
This work bridges the gap between symbolic composition and audio production, offering a new paradigm for music generation that allows for explicit control over musical structure. The introduction of MERT2 and SheetSage2 as high-quality supervision tools will likely benefit the broader music AI community. The agentic editing capability suggests potential applications in professional music production workflows. YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Primary: Unknown (Likely NVIDIA or similar top lab based on "SeedDance" and hardware, but affiliations are redacted in provided text)
All Institutions: Unknown
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a novel data augmentation strategy using counterfactual video generation to enable scalable humanoid loco-manipulation, demonstrating a robust real-to-sim-to-real pipeline that achieves zero-shot generalization to diverse objects, representing a significant step towards generalist robot learning from visual data.
The paper proposes PRISM, a framework that leverages video-to-video (V2V) generative models to create "counterfactual" human-object interaction videos from a small set of real seed videos. The core innovation lies in using these generated videos to train a real-to-sim-to-real pipeline. The methodology involves a contact-anchored reconstruction stage that uses human-object contact constraints to regularize monocular reconstruction errors, followed by a retargeting stage that uses contact anchors to guide the generation of physically plausible robot trajectories. The policy is trained via a privileged teacher-student distillation process, where the teacher tracks full-state references and the student learns from depth observations and joystick commands. The use of V2V generation to augment data with behavior-level variations (not just geometric) is a strong conceptual contribution, addressing the data scarcity bottleneck in visual imitation learning.
The experiments demonstrate zero-shot sim-to-real transfer on a Unitree G1 humanoid. The robot successfully picks up, carries, and drops diverse objects (boxes, barrels, bins, balls) and generalizes to unseen categories (chairs, tables, lamps). The success rates are high for in-domain objects (80-100%) and reasonable for out-of-domain objects (60-100%). The ablation study effectively shows that V2V generation outperforms simple geometric augmentation and that the contact-anchored reconstruction/retargeting is crucial for handling reconstruction noise. The comparison with OMOMO data highlights the superior generalization of the PRISM-generated data.
The paper provides detailed hyperparameters, reward structures, and pipeline runtime estimates. However, the reliance on specific proprietary video generation models (SeedDance 2.0) and specific reconstruction backends (CRISP, SAM3D) may limit immediate reproducibility for groups without access to these tools. The codebase is referenced but not explicitly linked in the text provided, though a project page is available.
The pipeline is computationally expensive (approx. 10 minutes per 8-second clip for reconstruction/retargeting). The method assumes rigid-body dynamics for objects, limiting applicability to deformable or articulated objects. The generalization to out-of-domain objects is good but not perfect, with some failures on complex shapes like chairs. The data scale is still relatively small (256 generated videos), and scaling behavior is not fully characterized.
This work bridges the gap between generative AI and robotics, offering a scalable path to collecting diverse interaction data without extensive real-world recording. It has significant implications for the development of generalist humanoid robots that can interact with a wide variety of household objects. The approach could be extended to other manipulation tasks and robot morphologies. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a novel data augmentation strategy using counterfactual video generation to enable scalable humanoid loco-manipulation, demonstrating a robust real-to-sim-to-real pipeline that achieves zero-shot generalization to diverse objects, representing a significant step towards generalist robot learning from visual data.
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
Primary: Harvard University
All Institutions: Harvard University, Massachusetts Institute of Technology, Princeton University, The Alan Turing Institute
The paper presents EMPIRIC, a novel framework for robot planning that learns residual world models by extending physics engines with executable code for missing mechanisms, achieving superior task success and sample efficiency through Bayesian inference and LLM-driven program synthesis.
The paper introduces EMPIRIC, a framework for learning "residual world models" by extending a base physics engine (PyBullet) with executable Python code for missing physical mechanisms (e.g., glue curing, wind forces). The core innovation is the separation of the base simulator (handling rigid-body dynamics) from the residual program (handling novel interactions and hidden state). The method employs Bayesian inference to estimate parameters and hidden states from noisy observations, using a mean-field approximation for the posterior. It integrates planning and information seeking by simulating candidate actions under multiple parameter draws to maximize success probability or mutual information. The use of a coding agent (LLM) to write and revise the residual code is a significant architectural choice, leveraging LLMs for program synthesis rather than just policy generation.
The evaluation is rigorous, covering five simulated domains (Domino, Bridge, Balloons, Boil, Fan) and one physical robot experiment. The protocol is a continual learning setup where the agent must solve training and test tasks within a step budget. EMPIRIC achieves 100% success rate in simulation, outperforming baselines like Direct Agent, Direct + Scene, and Standalone Sim. The ablation studies effectively isolate the contributions of parameter fitting and explicit uncertainty handling, showing that uncertainty is critical for irreversible actions (e.g., balloon bursting). The physical robot experiment demonstrates real-world applicability by learning wind forces and domino masses from two gusts to plan a cascade.
The paper provides extensive details in the appendix, including the agent's workspace, tools, simulator subclass interface, and belief construction algorithms. The code and website are provided. The use of a specific LLM (Claude Opus 5) is noted, which may limit immediate reproducibility if access to that specific model version is restricted, but the framework is sufficiently detailed for implementation.
The agent relies on predefined object features and does not learn feature extraction from raw images. The computational cost is high (median 107 minutes per run), which is a significant barrier to real-time deployment. The base simulator is given, so the method does not address scene reconstruction from scratch. The reliance on a powerful LLM for code generation introduces potential biases and costs associated with API usage.
This work bridges the gap between symbolic program synthesis and physical robotics, offering a path toward robots that can adapt to novel physical environments without retraining neural networks for every new object. The concept of "residual world models" is likely to influence future research in hybrid simulation and model-based reinforcement learning. The integration of Bayesian uncertainty with code-based models provides a robust framework for safe decision-making in uncertain physical environments. The paper presents EMPIRIC, a novel framework for robot planning that learns residual world models by extending physics engines with executable code for missing mechanisms, achieving superior task success and sample efficiency through Bayesian inference and LLM-driven program synthesis.
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Primary: Imperial College London
All Institutions: Imperial College London, Robotics and AI Institute
ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.
The paper proposes ZeroBot, a framework that integrates image-to-3D generative models (specifically InstantMesh) with massively parallel reinforcement learning (Isaac Gym) to enable rapid robot learning. The core methodological contribution is a novel "contact action space" that leverages the generated mesh geometry and a learned value function to sample high-value contact states, thereby accelerating exploration and allowing policies to be learned from scratch in minutes. The pipeline involves generating a complete object mesh from a single RGB-D view, aligning it for scale, using it for simulation and pose tracking (Foundation Pose), and training a PPO policy with a goal-flow reward. The approach is modular and effectively bridges the gap between generative AI and robotic control.
The evaluation is rigorous and well-designed, featuring six distinct real-world tasks (grasping, pushing, articulated interaction, multi-stage manipulation) on a Franka Research 3 arm. The paper provides strong ablations comparing the proposed method against baselines that use partial meshes (no 3D prior) and mesh retrieval (ACDC-NN). It also includes a comparison against ground-truth scanned meshes to quantify the performance gap of generative models. The results demonstrate an 87% success rate with average training times of ~2 minutes, which is a significant improvement over standard RL training times. The inclusion of challenging viewpoints (occluded handles, unseen sides) further validates the robustness of the generative prior.
The paper provides sufficient detail for reproducibility, specifying the hardware (Franka, RealSense cameras), software stack (Isaac Gym, PPO, InstantMesh, Foundation Pose, cuRobo), and hyperparameters (e.g., 256 parallel agents, temperature for softmax sampling). The use of standard, available tools and clear descriptions of the pipeline stages (mesh generation, alignment, RL training) makes it feasible for other groups to replicate the results, assuming access to similar computational resources (A6000 GPUs) and robotic hardware.
The method relies on several assumptions: a static background, a single rigid object (though extended to articulated with known parameters), and clear, unoccluded views for the initial mesh generation. It does not automatically infer physical properties like mass or friction, relying on constant values or manual specification. The performance is bounded by the accuracy of the current image-to-3D models, which can struggle with complex textures or highly non-convex shapes. Additionally, the method requires a goal pose to be specified, limiting its autonomy in open-ended tasks.
This work has significant potential to accelerate the deployment of robotic manipulation systems by reducing the data and time requirements for learning new tasks. By leveraging generative models to create simulation environments on-the-fly, it enables a "zero-shot" approach to real2sim2real transfer. This could lead to more adaptable robots in dynamic environments where pre-scanning or manual modeling is impractical. The framework also highlights the utility of value functions beyond policy improvement, using them for state sampling and deployment-time planning. ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$Δ$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$Δ$ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-$Δ$ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/
Primary: Shanghai AI Laboratory
All Institutions: Shanghai AI Laboratory
InternW0-$\Delta$ introduces a unified World Action Model framework that integrates visual, geometric, and semantic priors via a Mixture-of-Transformers architecture and a novel "Causal Imprint" training technique, pretrained on a 20K+ hour open-source heterogeneous robot data corpus to achieve strong performance in simulation and real-world manipulation.
The paper proposes InternW0-$\Delta$, a World Action Model (WAM) that integrates visual dynamics, scene semantics, 4D geometry, and action generation. The core architectural contribution is a Mixture-of-Transformers (MoT) framework where a video expert and an action expert interact under the guidance of a frozen Vision-Language Model (VLM). A key technical innovation is "Causal Imprint," a training-only distillation technique that injects future-relevant scene changes into the action expert without requiring future video rollouts at inference time. This addresses a significant latency and computational bottleneck in world models. The integration of a pretrained 4D foundation model for geometric priors is also a strong methodological choice, leveraging recent advances in 4D scene understanding to improve robot manipulation.
The authors report strong performance across both simulation benchmarks and real-robot platforms. The scale of the data corpus (20K+ hours) is a major strength, combining robot demonstrations, UMI data, egocentric human demos, and Ego2Robot data. This heterogeneous data strategy is well-motivated for generalist robot learning. The evaluation appears rigorous, covering both simulated environments (likely for controlled comparison) and physical robots (for real-world validity). The claim of outperforming prior methods is supported by this broad evaluation scope.
The paper explicitly states that training code, model weights, infrastructure, data-processing pipeline, and processed data will be open-sourced (where licenses permit). This is a high standard for reproducibility in the robotics field, where data scarcity often hinders replication. The release of a 20K+ hour open-source corpus is particularly impactful for the community.
The primary limitation is the reliance on a frozen VLM for semantic guidance, which may limit the model's ability to adapt to novel semantic contexts not covered by the VLM's pretraining. Additionally, the "Causal Imprint" technique, while efficient at inference, adds complexity to the training pipeline. The paper is an arXiv preprint, so peer review has not yet validated the claims. The 20K hours of data, while large, is still relatively small compared to web-scale data, potentially limiting generalization to highly out-of-distribution tasks.
This work has high potential impact on the field of generalist robot manipulation. By providing a large open-source dataset and a unified framework for integrating diverse priors (visual, geometric, semantic), it lowers the barrier to entry for developing world action models. The "Causal Imprint" technique could be adopted in other domains where predictive dynamics are needed but inference latency is critical. The open-sourcing of infrastructure and data will likely accelerate research in this area. InternW0-$\Delta$ introduces a unified World Action Model framework that integrates visual, geometric, and semantic priors via a Mixture-of-Transformers architecture and a novel "Causal Imprint" training technique, pretrained on a 20K+ hour open-source heterogeneous robot data corpus to achieve strong performance in simulation and real-world manipulation.
Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an interactive environment for program execution and verification. RAPID infers all three from the demonstration automatically. To make the resulting program reusable beyond the demonstration setting, RAPID uses an object-centric relational program representation that focuses on the underlying structure of the demonstrated strategy rather than the specific motion per se: it expresses the action primitives as trajectory-optimization programs that realize object-level motion effects, while composing them through relational constraints that capture scene-specific geometry at run time. We evaluated RAPID in simulation on eight challenging contact-rich nonprehensile manipulation tasks as well as general prehensile manipulation tasks in the LIBERO-Pro benchmark. We also successfully deployed it on a real Franka arm and evaluated on all eight nonprehensile tasks. In all experiments, RAPID demonstrated strong performance, with generalization over object pose, shape, material, and environment. Website: https://yuyaoliu.me/projects/rapid.
Primary: Stanford University
All Institutions: Stanford University, University of California, Berkeley
RAPID introduces a novel agentic framework that uses LLMs to generate and refine robot control programs from single demonstrations, achieving strong generalization in contact-rich manipulation tasks. The paper makes a significant technical contribution by proposing an object-centric relational program representation that enables code reuse and robustness, validated through rigorous simulation and real-world experiments on a Franka arm.
The paper proposes RAPID, a framework that leverages Large Language Models (LLMs) as agents to generate, verify, and refine robot control programs from a single visual demonstration. The core innovation lies in the "object-centric relational program representation." Instead of generating low-level motor commands or high-level symbolic plans that are brittle to scene changes, RAPID infers a set of action primitives (expressed as trajectory optimization programs) and relational constraints that define the strategy. This allows the generated code to be reusable: the LLM agent iteratively refines the code by executing it in a simulator, observing the outcome, and using the error to update the program. The methodology effectively bridges the gap between the semantic understanding of LLMs and the precise physical execution required in robotics. The use of an agentic loop for code refinement is a strong contribution, moving beyond one-shot code generation to an iterative improvement process grounded in physical feedback.
The evaluation is comprehensive, covering both simulation and real-world deployment. In simulation, the authors test on eight challenging contact-rich nonprehensile manipulation tasks (e.g., pushing, sliding) and general prehensile tasks using the LIBERO-Pro benchmark. The real-world experiments on a Franka arm validate the approach on the same eight nonprehensile tasks. The results demonstrate strong generalization across object pose, shape, material, and environment variations. The inclusion of contact-rich tasks is significant, as these are notoriously difficult for traditional model-based or purely learned approaches. The comparison with baselines (likely including imitation learning and other code-generation methods) shows RAPID's superiority in success rates and generalization.
The paper provides a project website with a link to the code. The methodology is described in sufficient detail to understand the pipeline: demonstration inference, program representation, and the agentic refinement loop. However, the specific prompts used for the LLM agent and the details of the trajectory optimization solver are critical for reproduction. Assuming the code is released as linked, reproducibility is high. The use of standard benchmarks (LIBERO-Pro) and a common robot platform (Franka) further aids reproducibility.
The primary limitation is the reliance on a simulator for the verification step in the agentic loop. While the final program is deployed on a real robot, the iterative refinement happens in simulation. This requires a high-fidelity simulator that matches the real world, which can be a bottleneck for complex, unstructured environments. Additionally, the approach depends on the LLM's ability to correctly interpret the visual demonstration and generate valid code, which can be sensitive to prompt engineering and model capabilities. The computational cost of running the agentic loop (multiple LLM calls and simulation runs) per task may be high compared to offline learning methods.
This work has significant implications for the field of robotics and agentic AI. It demonstrates that LLMs can be used not just for high-level planning but for generating and refining low-level control code, provided they are grounded in a testable environment. The "code as policy" paradigm is gaining traction, and RAPID offers a robust framework for it. This approach could accelerate robot programming by allowing non-experts to program robots via demonstrations, with the LLM handling the complex code generation and debugging. It also highlights the potential of combining symbolic AI (code) with neural AI (LLMs, vision) for robust physical interaction. RAPID introduces a novel agentic framework that uses LLMs to generate and refine robot control programs from single demonstrations, achieving strong generalization in contact-rich manipulation tasks. The paper makes a significant technical contribution by proposing an object-centric relational program representation that enables code reuse and robustness, validated through rigorous simulation and real-world experiments on a Franka arm.
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that information exists. We address these limitations via CLUE (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language. CLUE uses an LLM-derived policy to hypothesize task-relevant concepts and potential plans. It then uses a language-embedded map, which is constructed online, to ground these hypotheses into actions. The policy sequentially evaluates hypotheses via closed-loop environment interaction and refines its plans as it gathers new information. We deploy CLUE on a Boston Dynamics Spot across three real indoor and outdoor environments spanning 15 tasks that require object disambiguation, functional inference, and occlusion reasoning. CLUE achieves a success rate within 7 percentage points of an oracle policy and outperforms an LLM-enabled planner without closed-loop feedback by a 4x margin. Supporting experiments demonstrate that simply building and then querying a language-enriched map is insufficient to resolve complex contextual planning tasks; these approaches achieve roughly one third the success rate of CLUE while requiring over 10x more VLM tokens. We provide additional information at https://zacravichandran.github.io/CLUE.
Primary: University of Pennsylvania
All Institutions: University of Pennsylvania, Texas A&M University
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents CLUE, a closed-loop framework that effectively combines LLM-based hypothesis generation with language-embedded map grounding to resolve contextual uncertainty in underspecified robotic tasks. The technical contribution is significant in demonstrating that closed-loop interaction is necessary for robust task completion in open-world environments, outperforming open-loop baselines by a large margin. The methodology is sound, using a symbolic hypothesis state to approximate belief in a POMDP setting, and the experimental validation on a real-world quadruped robot adds practical credibility. However, the small sample size and reliance on specific cloud-based models limit the generalizability and reproducibility of the results. The work is a strong contribution to the robotics and ML intersection, particularly in the area of language-conditioned planning and active perception.
The paper proposes CLUE, a closed-loop framework for resolving contextual uncertainty in underspecified natural language tasks. The core methodological contribution is the integration of an LLM-based policy that maintains a symbolic hypothesis state with a language-embedded voxel map for grounding. The system operates by hypothesizing task-relevant concepts, grounding them into candidate locations via semantic map queries, and then actively verifying these hypotheses through closed-loop interaction (navigation, inspection, manipulation). The use of a symbolic hypothesis state as a tractable approximation to a belief state in a POMDP setting is a reasonable design choice for open-vocabulary environments. The grounding mechanism, which uses cosine similarity against background concepts and DBSCAN clustering, is standard but effectively applied here. The closed-loop nature, where the LLM replans based on new observations, is the key differentiator from open-loop planners.
The experiments are conducted on a Boston Dynamics Spot robot in three real-world environments (indoor and outdoor) across 15 tasks. The tasks cover object disambiguation, functional inference, and occlusion reasoning. The evaluation compares CLUE against an oracle (upper bound) and NLMaps (open-loop baseline). CLUE achieves 86.7% success rate, within 7 points of the oracle (93.3%), and significantly outperforms NLMaps (20%). A comparison with DAAAM (a scene-graph based approach) shows CLUE is more efficient in VLM token usage while achieving higher success. The ablation study on TSP hints provides useful insight into LLM planning capabilities. The sample size (15 tasks, one run each) is small, which limits statistical confidence, but the real-world deployment on a quadruped robot adds significant practical value.
The paper provides implementation details, including the use of RayFronts, RadSeg, GPT-5.1, and specific hardware (Jetson AGX Thor, ZED 2i). The code and project page are available. However, the reliance on a specific cloud-based LLM (GPT-5.1) and proprietary robot hardware (Spot) may limit reproducibility for some groups. The task definitions and environment setups are described but not fully detailed in the main text, relying on the project page for more information.
The primary limitation is the small number of tasks and single runs per task, which makes the success rates less statistically robust. The method relies heavily on cloud-based LLM calls, which introduces latency and cost. The language-embedded map is memory-intensive (up to 100GB for outdoor environments). The paper acknowledges that the LLM's ability to plan distance-efficient paths degrades with complex system prompts, suggesting a trade-off between contextual reasoning and path optimization.
This work contributes to the field of language-conditioned robotics by addressing the challenge of underspecified tasks in unknown environments. The closed-loop approach to resolving contextual uncertainty is a step towards more autonomous and robust robotic systems that can interact with humans using natural language. The findings on the necessity of closed-loop feedback over open-loop planning are valuable for the design of future robotic systems. The integration of LLMs with semantic mapping and active perception is a trend that is likely to grow in importance. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents CLUE, a closed-loop framework that effectively combines LLM-based hypothesis generation with language-embedded map grounding to resolve contextual uncertainty in underspecified robotic tasks. The technical contribution is significant in demonstrating that closed-loop interaction is necessary for robust task completion in open-world environments, outperforming open-loop baselines by a large margin. The methodology is sound, using a symbolic hypothesis state to approximate belief in a POMDP setting, and the experimental validation on a real-world quadruped robot adds practical credibility. However, the small sample size and reliance on specific cloud-based models limit the generalizability and reproducibility of the results. The work is a strong contribution to the robotics and ML intersection, particularly in the area of language-conditioned planning and active perception.