Last 7 Days (September 26 – October 02, 2026)
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
Primary: Anthropic
All Institutions: Anthropic
Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
The paper proposes "Constitutional Adapters" (CAs), a method for distilling alignment principles into lightweight, inference-time interventions (LoRAs and ReLU steering vectors). The core methodological innovation is the use of a control-subtracted task arithmetic approach: training a lightweight object on a synthetic corpus derived from a model constitution ($D_C$) and subtracting an object trained on a register-matched control corpus ($D_K$). This subtraction aims to isolate the "constitutional" behavior from general fine-tuning artifacts like format or register shifts. The method leverages synthetic data generation via a strong model (Claude Opus 5) to create diverse scenarios that stress-test ethical principles, rather than relying on adversarial jailbreak examples. This allows the adapters to generalize to unseen attacks. The use of ReLU steering vectors, which allow for input-dependent steering strength, is a significant technical detail that improves upon fixed-direction steering.
The experimental evaluation is rigorous and comprehensive. The authors test across four distinct model families (Qwen3.5, Llama-3.1, Gemma-4, Nemotron 3) and various attack types (static, multi-query, multi-turn). Key findings include the superior robustness of CAs in long-context and multi-turn settings compared to system prompting, which degrades as context length increases. The paper also evaluates the trade-off between defense and benign compliance (over-refusal) by sweeping the steering strength, demonstrating a predictable and tunable frontier. The inclusion of misalignment evaluations (Petri Bloom) and capability benchmarks (MMLU, GSM8K, etc.) provides a holistic view of the method's impact, showing minimal capability degradation.
The paper provides detailed descriptions of the synthetic data pipeline, training hyperparameters, and evaluation protocols. It mentions that code and data will be released upon acceptance. The use of specific model versions (e.g., Claude Opus 5, Qwen3.5-4B) and clear definitions of metrics (StrongREJECT, over-refusal scoring) enhances reproducibility. However, the reliance on proprietary models for data generation (Claude) may limit full reproducibility for external researchers without access to those specific model versions.
The method is evaluated primarily on models up to 8 parameters, and it is unclear how well it scales to frontier-scale models (e.g., 70B+). The synthetic data generation relies on a strong, aligned model, which may introduce biases or limit the diversity of the training data. The paper does not extensively discuss the computational cost of generating the synthetic corpora or the potential for adversarial attacks specifically targeting the adapter weights. Additionally, the zero-shot transfer to post-trained checkpoints is a strong claim that might not hold for all post-training procedures.
This work has significant implications for the deployment of LLMs in API settings, where lightweight, tunable safety interventions are highly desirable. By providing a method to "turn up" safety without retraining or terminating sessions, CAs offer a practical solution for dynamic risk management. The approach also contributes to the broader field of AI alignment by demonstrating that complex, principle-based behaviors can be compressed into low-dimensional, portable objects. This could facilitate the development of modular safety systems that can be updated or adjusted independently of the base model. Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: Unknown (Likely ByteDance based on "YuE" naming convention and technical style, but not explicitly stated in provided text)
All Institutions: Unknown
YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
The paper proposes YuE2, a unified framework that integrates symbolic music generation (ABC notation) with audio generation (flow matching) within a single Mixture-of-Transformers (MoT) architecture. The core methodological contribution is the "symbolic planning" step, where the model first generates a readable score (melody, harmony, form) before expanding it into semantic tokens and acoustic latents. This is supported by two auxiliary models: MERT2, a music representation learner that sets new SOTA on MARBLE benchmarks, and SheetSage2, a full-song transcription model that generates the symbolic supervision signals required for training. The use of an AR-NAR MoT to handle both discrete symbolic/semantic tokens and continuous acoustic latents is a sophisticated architectural choice that allows for bidirectional attention in the acoustic stream while maintaining causal generation for the score.
The evaluation is extensive, covering automatic metrics (SongBench, SongEval, AudioBox) and expert listening tests. YuE2 outperforms public baselines and is competitive with proprietary systems like Suno v4.5/v5. The ablation study on symbolic planning is particularly strong, demonstrating that generating the score first significantly improves perceived musicality and overall quality compared to direct audio generation. The score-editing experiments show that the model can preserve unedited content while modifying specific sections, validating the utility of the symbolic interface.
The paper provides detailed architectural descriptions and training procedures. However, as a technical report from a likely industry lab, the full code and weights may not be immediately available to the public, which limits immediate reproducibility. The reliance on proprietary or large-scale datasets (346,000 hours of music) also poses a barrier for independent replication.
The model is large (3.58B parameters) and computationally expensive. The evaluation against proprietary systems is limited to expert listening and some automatic metrics, as direct API access for rigorous benchmarking is often restricted. The "best-of-8" selection strategy, while effective for benchmarking, may not reflect real-time interactive use cases where latency is critical.
This work bridges the gap between symbolic composition and audio production, offering a new paradigm for music generation that allows for explicit control over musical structure. The introduction of MERT2 and SheetSage2 as high-quality supervision tools will likely benefit the broader music AI community. The agentic editing capability suggests potential applications in professional music production workflows. YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
Primary: University of Cambridge
All Institutions: University of Cambridge, DeepMind
Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
The paper introduces Fold'EM, a framework that bypasses the traditional cryo-EM density reconstruction step by directly guiding AlphaFold3 using individual particle images. The core technical novelty lies in optimizing intermediate Pairformer states ($H$) rather than the terminal embeddings ($Z$) or atomic coordinates. The authors provide a rigorous theoretical justification for this choice, demonstrating that optimizing $H$ induces a sequence-dependent anisotropic geometry in the conditioning space, which acts as a learned preconditioner for the gradient updates. This prevents the model from drifting into unrealistic structural configurations during the inference-time optimization. The method also introduces a joint inference scheme for structure, pose, and class assignment, utilizing a novel bootstrapped residual-subspace estimator (B-RECOVAR) for label-free classification in heterogeneous samples. The mathematical formulation of the particle-level likelihood versus density-level likelihood is well-articulated, highlighting the statistical insufficiency of consensus densities when poses are uncertain.
The experimental evaluation is extensive, covering both synthetic benchmarks and experimental EMPIAR datasets. The paper demonstrates significant improvements over baselines (RELION, cryoDRGN, and previous AlphaFold-guided methods) in terms of FSC resolution and TM-score, particularly in the low-sample regime (e.g., recovering structures from as few as 100-500 particles). The ablation studies effectively isolate the contributions of particle-level guidance versus density guidance and the optimization of intermediate states. The results on heterogeneous data show the ability to resolve distinct conformational states without separate density reconstructions, which is a significant practical advantage. The comparison with unguided AlphaFold3 clearly shows the value of the experimental guidance in correcting mispredicted conformations.
The paper provides detailed algorithmic descriptions and mathematical formulations. However, as an arXiv preprint, the code availability is not explicitly confirmed in the text provided. The reliance on AlphaFold3, a proprietary model, may limit full reproducibility for some users, though the method itself is clearly defined. The use of standard cryo-EM tools (RELION, cryoDRGN) for baselines aids in comparability.
The method relies heavily on the quality of the AlphaFold3 prior; if the unguided prediction is significantly wrong (e.g., 8PWH case), the method may struggle or require more particles than density-based methods. The computational cost of running multiple reverse-diffusion runs and back-propagating through the Pairformer could be high. The current implementation assumes known CTF parameters and centered particles, which may not always hold in raw experimental data without preprocessing.
This work has the potential to significantly reduce the sample size required for cryo-EM structure determination, making it feasible to study low-population states and scarce samples. It bridges the gap between protein structure prediction and experimental cryo-EM, offering a new paradigm for structure determination that could accelerate drug discovery and structural biology research. The insights into optimizing intermediate representations of generative models may also be applicable to other inverse problems in scientific imaging. Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Primary: Princeton University
All Institutions: Princeton University
The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
The paper proposes a simple but effective alternative to the standard RLVR (Reinforcement Learning with Verifiable Rewards) pipeline: instead of training a single LoRA adapter on the full dataset, it splits the data and budget across K independent adapters ("thickets"). The core insight is that RLVR sharpens the policy, reducing sample diversity and increasing error correlation among voters in majority voting, which degrades test-time scaling performance. The methodology is sound, using a fixed inference budget (160 completions) to isolate the effect of training strategy from sampling budget. The use of a step-matched control (early stopping the single adapter) is a critical experimental design choice that strengthens the causal claim.
The experiments are rigorous and well-controlled. The authors test across four models (Qwen2.5-1.5B/3B/7B, Llama-3.1-8B) and two domains (Math, Code). They demonstrate that the single-adapter approach often performs worse than the untrained base model in majority voting, while the thicket approach recovers this performance and often exceeds it. The analysis of error correlation and coverage provides strong mechanistic evidence for the observed results. The ablation on shard type (random vs. subject-specific) adds practical value.
The paper provides sufficient detail on the training setup (GRPO, LoRA rank, hyperparameters) and evaluation protocols (temperature, top-p, voting logic). The specific datasets (MATH, GSM8K, etc.) and model checkpoints are standard, making reproduction feasible for a lab with adequate compute resources. The code is not explicitly linked in the provided text, but the methods are standard enough to implement.
The study is limited to models up to 8B parameters and specific RL algorithms (GRPO). The voting results are primarily for math tasks where answers are canonicalizable; code tasks are only analyzed for coverage and correlation, not final vote accuracy. The "thicket" approach requires managing multiple adapters at inference time, which, while mitigated by modern serving stacks, adds operational complexity compared to a single model.
This paper challenges a common assumption in the LLM post-training community: that more RLVR training is always better for test-time scaling. It provides a clear, actionable guideline for practitioners: if you plan to use majority voting, diversify your training budget. This has immediate practical implications for teams deploying LLMs with test-time compute budgets. The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Primary: Harvard University
All Institutions: Harvard University, Harvard Medical School
The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
The paper employs a rigorous interventionist methodology to test the causal role of biological foundation model representations in LLM-based reasoning models. The authors utilize input perturbation (shuffling DNA sequences), evidence conflict construction (mismatching foundation model embeddings with text descriptions), and linear probing to isolate the contribution of specific modalities. This approach is methodologically sound for determining whether models are genuinely "reasoning" over biological inputs or merely relying on textual shortcuts. The inclusion of analysis across multiple training checkpoints (SFT and RL) adds depth to the evaluation of how post-training strategies affect input utilization.
The experiments cover six distinct biological reasoning models across DNA, protein, and single-cell tasks, providing a broad scope. The key finding that Evo2 and ESM3 contribute negligibly to BioReason and BioReason-Pro performance, with models following text in >97% of conflict cases, is a significant empirical result. The contrast with models like ChatNT and CellWhisperer, where foundation model inputs do contribute, highlights that the issue is specific to certain post-training strategies or model architectures rather than a universal failure of multimodal integration. The analysis of reasoning traces revealing misstatements of nucleotide changes further supports the conclusion that the models are not faithfully using the biological inputs.
The paper provides a GitHub repository link (https://github.com/mims-harvard/bio-mirage) and a project website, which strongly suggests that code and resources will be available for reproduction. The detailed description of the perturbation and conflict construction methods allows other researchers to replicate the analysis on other models.
The study focuses on a specific set of six models and three biological domains. It is unclear if the findings generalize to other types of biological foundation models or different post-training paradigms. The linear probes may not capture all non-linear relationships between the foundation model representations and the task targets, potentially underestimating the contribution of the biological inputs in some cases.
This paper has high impact for the field of AI in science, particularly for the development of trustworthy biological reasoning systems. It challenges the common assumption that high benchmark accuracy implies effective use of specialized scientific inputs. The findings will likely influence future post-training strategies to explicitly reward the use of foundation model representations, leading to more robust and interpretable AI models for biology. The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University, University of California, Berkeley, Impossible AI
The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
The paper introduces "Interactive Program Induction" (IPI), a paradigm where LLM agents represent their understanding of unknown environments as executable programs rather than prose. The proposed harness, Schema, consists of four core operations: Hypothesize (writing code to model state and transitions), Certify (replaying history to check consistency), Plan (using the program as a simulator for search), and Act with Verification (executing actions and stopping if predictions mismatch). This approach effectively addresses the "lost in the middle" and context degradation issues inherent in prose-based memory by forcing the agent to distill knowledge into compact, testable, and reusable code. The methodology is sound, leveraging the LLM's coding capabilities to create a persistent, verifiable world model that persists across context compactions.
The evaluation is extensive and rigorous, covering three distinct benchmarks: ARC-AGI-3 (visual reasoning), DiG-bench (text-based rule discovery), and MazeBench (long-horizon 3D exploration). The results are striking: Schema achieves 99.2% RHAE on ARC-AGI-3 (vs. 58.7% baseline), solves 100% of public DiG-bench games, and matches top-50 human performance on MazeBench. Ablation studies clearly demonstrate the contribution of each component (certification, planning, verification), showing that removing any one significantly degrades performance. The analysis of token costs and action efficiency further strengthens the claim of practical utility.
The paper provides detailed implementation descriptions in the appendix, including the program contract, tool interfaces, and benchmark adapters. However, the specific code for the Schema harness and the exact prompts used are not fully detailed in the text, and no public code repository URL is provided in the extracted text. While the methodology is clearly described, full reproducibility would require access to the specific harness implementation and prompt engineering details.
The approach relies heavily on the base model's coding ability; weaker models may struggle to write correct world models. The computational cost of backtesting and planning can be high, though the paper argues this is offset by reduced interaction steps. The benchmarks used (ARC-AGI-3, DiG-bench, MazeBench) are relatively new and may not fully represent the breadth of real-world unknown environments.
This work has significant implications for the design of autonomous agents in open-ended environments. By shifting from passive memory to active, executable theory-building, it offers a scalable path toward agents that can genuinely learn and adapt to novel tasks without retraining. The paradigm of "interactive program induction" could be applied to robotics, scientific discovery, and other domains where environments are complex and rules are not explicitly given. The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Primary: Cornell University
All Institutions: Cornell University, Stanford University
The paper introduces a cost-effective, rule-based synthetic environment framework for training LLM search agents, demonstrating that structural complexity (hop count) drives transferability to real-world tasks more than factual accuracy. This approach offers a scalable alternative to human-curated or LLM-generated data, enabling the emergence of adaptive search strategies in LLM agents.
The paper proposes a novel paradigm for training LLM agents using Reinforcement Learning (RL) by replacing costly, human-curated, or LLM-generated environments with "PhantomEnvironments." These are synthetic, rule-based environments derived from fictional worlds where agents must perform multi-hop search over templated articles. The core methodological innovation is the decoupling of environment generation from LLM inference, ensuring zero marginal cost and eliminating hallucination risks associated with LLM-generated data. The approach relies on the hypothesis that the structural complexity of the search task (specifically hop count) is more critical for learning generalizable search strategies than the factual accuracy of the content.
The experiments demonstrate that agents trained on these fictional, factually incorrect environments transfer effectively to real-world multi-hop search benchmarks. Notably, the paper claims that agents trained on PhantomEnvironments often outperform those trained on real-world data on newer benchmarks, suggesting that the structural learning of search strategies is robust to domain shift. Ablations confirm that hop count is the primary driver of transfer performance. The observation of "emergent search scaling," where Qwen models learn to allocate search budget linearly with question difficulty, is a significant empirical finding.
The authors provide a strong reproducibility statement, indicating the use of open-source LLMs, training code, and evaluation benchmarks. They report standard errors and significance tests. A GitHub repository is provided for code access. The use of AI tools for code implementation is disclosed, but the authors state they verified all code and results.
The primary limitation is the reliance on "templated articles" and rule-based generation, which may not capture the full stochasticity and ambiguity of real-world web search. The transferability to "newer benchmarks" is a strong claim that requires careful scrutiny regarding potential leakage or specific alignment with the benchmark structure. The paper is a preprint (ICLR 2027 submission), so peer review status is pending.
This work has high potential impact by providing a scalable, cost-effective method for training LLM agents. If the transferability results hold, it could significantly reduce the barrier to entry for developing sophisticated search agents, allowing smaller labs to train competitive models without massive human annotation budgets. It shifts the focus from data fidelity to structural complexity in agent training. The paper introduces a cost-effective, rule-based synthetic environment framework for training LLM search agents, demonstrating that structural complexity (hop count) drives transferability to real-world tasks more than factual accuracy. This approach offers a scalable alternative to human-curated or LLM-generated data, enabling the emergence of adaptive search strategies in LLM agents.
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Primary: NVIDIA
All Institutions: NVIDIA
PivotOPD introduces a targeted distillation framework that teaches LLM agents to recover from pivotal mistakes in multi-turn interactions. The paper provides a rigorous analysis of error accumulation, a novel dual-distillation method, and strong empirical results across diverse benchmarks, establishing a new standard for training robust language agents.
The paper proposes PivotOPD, a framework that addresses error accumulation in multi-turn LLM agents by identifying "pivotal mistakes" and applying targeted distillation. The core innovation is the dual distillation approach: preventive distillation (using reverse KL to steer away from mistakes) and recovery distillation (using forward KL to teach recovery behaviors from a privileged self-teacher). The theoretical analysis correctly identifies that standard on-policy methods fail to provide sufficient learning signal for recovery actions because the student rarely samples them, justifying the use of forward KL on teacher-generated responses. The pivot detection mechanism, which uses a teacher model to identify candidate turns and gold actions, is a practical solution to the lack of oracles in real-world environments.
The experiments are extensive, covering four benchmarks (ALFWorld, WebShop, Search-based QA, SWE-Bench Verified) and multiple model families (Qwen3, Nemotron). The comparison against 13 baselines is robust. The results show consistent improvements, particularly in recovery rates, which aligns with the paper's motivation. The ablation studies effectively demonstrate the necessity of both preventive and recovery components and the importance of supervision at the correct turns. The transfer to SWE-Bench Verified with a different model family (Nemotron) strengthens the generalizability claim.
The paper provides detailed descriptions of the training objective, pivot detection prompts, and evaluation metrics. The project page URL is provided, which likely contains code and additional resources. The hyperparameters and training configurations are described in the appendix, supporting reproducibility.
The method relies on a teacher model for pivot detection and action naming, which may not be available in all settings. The performance gains, while statistically significant, are moderate in some benchmarks (e.g., +1.2% on WebShop success rate). The reliance on a privileged self-teacher for recovery distillation adds computational overhead. The pivot detection accuracy (77.8% within one turn of oracle) suggests room for improvement in identifying pivotal turns without an oracle.
This work has significant implications for training robust LLM agents in interactive environments. By explicitly teaching recovery from mistakes, it addresses a key limitation of current agent training methods. The framework could be extended to other domains where error accumulation is a challenge, such as robotics or autonomous driving. The insights into the nature of pivotal mistakes and recovery behaviors provide valuable guidance for future research on agent robustness. PivotOPD introduces a targeted distillation framework that teaches LLM agents to recover from pivotal mistakes in multi-turn interactions. The paper provides a rigorous analysis of error accumulation, a novel dual-distillation method, and strong empirical results across diverse benchmarks, establishing a new standard for training robust language agents.
We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal component regression (PCR), gradient descent (GD), and ridge regression. We show that PCR dominates all monotone spectral filters: compared to any such filter, the risk of optimally tuned PCR is no bigger by a constant factor for all problems. Furthermore, the dominance is strong if the filter is separated from step functions (e.g., GD and ridge): there exist problem instances for which the risk of PCR is smaller by a polynomial factor in sample size dependence. Our comparison results show that PCR is optimal and thus admissible among monotone filters, significantly extending Wu et al. (2026)'s result that GD strongly dominates ridge. From a technical perspective, we establish new upper and lower bounds for general spectral filters, which are instance-wise sharp when specialized to ridge or GD, recovering or improving the best-known bounds.
Primary: UC Berkeley
All Institutions: UC Berkeley, Princeton University
Principal Component Regression dominates all monotone spectral filters for linear regression in terms of instance-wise finite-sample risk, establishing PCR as admissible and GD as inadmissible. The paper provides a comprehensive theoretical analysis using Schur multipliers to derive sharp risk bounds, offering a new perspective on the comparison of regularization methods beyond worst-case minimax rates.
The paper employs a rigorous theoretical framework to compare the instance-wise finite-sample risks of monotone spectral filters in linear regression. The core methodological contribution is the application of Schur multipliers and matrix divided differences to control noncommutative matrix perturbations, specifically the "leave-tail-out" difference between the full Gram matrix and its truncated version. This allows for the derivation of sharp upper and lower bounds for general spectral filters, including Principal Component Regression (PCR), Gradient Descent (GD), and Ridge Regression. The authors demonstrate that PCR dominates all monotone filters and strongly dominates filters separated from step functions (like GD and Ridge), establishing PCR as admissible and GD as inadmissible in this context. The technical depth is high, extending classical leave-one-out ideas with advanced operator theory tools.
This is a purely theoretical paper with no experimental evaluation. The "experiments" are mathematical proofs and derivations of risk bounds. The validity is established through rigorous mathematical argumentation rather than empirical benchmarks.
As a theoretical paper, reproducibility is defined by the clarity and correctness of the proofs. The paper provides detailed appendices with missing proofs and clearly states assumptions. The use of AI for parts of the technical ingredients is disclosed, but the authors state they rederived and verified all proofs. The mathematical framework is well-defined and reproducible in the sense that other researchers can verify the theorems.
The dominance results rely on Gaussian random design assumptions, which may not hold in all practical settings. The paper acknowledges that the variance bounds for PCR could likely be improved. Additionally, the results are specific to linear regression; extending these dominance relationships to non-linear models or other learning tasks remains an open question. The reliance on specific spectral properties (monotonicity) limits the scope to a specific class of estimators.
The paper significantly impacts the field of statistical learning theory by providing a definitive instance-wise comparison of standard linear regression methods. It challenges the common heuristic that Gradient Descent is a universally good default by showing it is inadmissible compared to PCR in terms of instance-wise risk. This insight may influence the design of future algorithms, encouraging the use of methods that can explicitly discard weak spectral components to reduce variance. It also provides a new technical toolkit (Schur multipliers for risk analysis) that can be applied to other statistical learning problems. Principal Component Regression dominates all monotone spectral filters for linear regression in terms of instance-wise finite-sample risk, establishing PCR as admissible and GD as inadmissible. The paper provides a comprehensive theoretical analysis using Schur multipliers to derive sharp risk bounds, offering a new perspective on the comparison of regularization methods beyond worst-case minimax rates.
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
Primary: Anthropic
All Institutions: Anthropic
Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
The paper proposes "Constitutional Adapters" (CAs), a method for distilling alignment principles into lightweight, inference-time interventions (LoRAs and ReLU steering vectors). The core methodological innovation is the use of a control-subtracted task arithmetic approach: training a lightweight object on a synthetic corpus derived from a model constitution ($D_C$) and subtracting an object trained on a register-matched control corpus ($D_K$). This subtraction aims to isolate the "constitutional" behavior from general fine-tuning artifacts like format or register shifts. The method leverages synthetic data generation via a strong model (Claude Opus 5) to create diverse scenarios that stress-test ethical principles, rather than relying on adversarial jailbreak examples. This allows the adapters to generalize to unseen attacks. The use of ReLU steering vectors, which allow for input-dependent steering strength, is a significant technical detail that improves upon fixed-direction steering.
The experimental evaluation is rigorous and comprehensive. The authors test across four distinct model families (Qwen3.5, Llama-3.1, Gemma-4, Nemotron 3) and various attack types (static, multi-query, multi-turn). Key findings include the superior robustness of CAs in long-context and multi-turn settings compared to system prompting, which degrades as context length increases. The paper also evaluates the trade-off between defense and benign compliance (over-refusal) by sweeping the steering strength, demonstrating a predictable and tunable frontier. The inclusion of misalignment evaluations (Petri Bloom) and capability benchmarks (MMLU, GSM8K, etc.) provides a holistic view of the method's impact, showing minimal capability degradation.
The paper provides detailed descriptions of the synthetic data pipeline, training hyperparameters, and evaluation protocols. It mentions that code and data will be released upon acceptance. The use of specific model versions (e.g., Claude Opus 5, Qwen3.5-4B) and clear definitions of metrics (StrongREJECT, over-refusal scoring) enhances reproducibility. However, the reliance on proprietary models for data generation (Claude) may limit full reproducibility for external researchers without access to those specific model versions.
The method is evaluated primarily on models up to 8 parameters, and it is unclear how well it scales to frontier-scale models (e.g., 70B+). The synthetic data generation relies on a strong, aligned model, which may introduce biases or limit the diversity of the training data. The paper does not extensively discuss the computational cost of generating the synthetic corpora or the potential for adversarial attacks specifically targeting the adapter weights. Additionally, the zero-shot transfer to post-trained checkpoints is a strong claim that might not hold for all post-training procedures.
This work has significant implications for the deployment of LLMs in API settings, where lightweight, tunable safety interventions are highly desirable. By providing a method to "turn up" safety without retraining or terminating sessions, CAs offer a practical solution for dynamic risk management. The approach also contributes to the broader field of AI alignment by demonstrating that complex, principle-based behaviors can be compressed into low-dimensional, portable objects. This could facilitate the development of modular safety systems that can be updated or adjusted independently of the base model. Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges between components. We establish the first information-theoretic lower bounds on sample complexity for this target. For $p$ variables, maximum SCC size $s_{\max}$, and maximum external-parent count $d_B$, any estimator requires order $s_{\max}\log(ep/s_{\max})+d_B\log(ep/d_B)$ samples in the worst case over a regular model class. These bounds distinguish the costs of SCC membership and external-parent selection. Under principal invertibility and without correlation faithfulness, we establish a population block-exogeneity principle that identifies unknown root SCCs through residual independence and inclusion minimality. A sparse-adjustment characterization shows that small adjustment sets suffice to identify SCCs and their direct external parents, without regressing on all previously recovered variables. These characterizations yield BlockExo, which attains a structurally matching sample bound without knowing $s_{\max}$ or $d_B$ under suitable conditions. Simulations support the structural dependence of our sample bound and demonstrate BlockExo's sample-efficient recovery in comparisons with other methods for cyclic causal discovery.
Primary: Princeton University
All Institutions: Princeton University, Seoul National University
The paper establishes the first information-theoretic lower bounds for sample complexity in cyclic causal discovery and introduces BlockExo, an algorithm that achieves structurally optimal recovery. By decoupling the costs of SCC membership and external parent selection, the work provides a precise characterization of the statistical difficulty of learning feedback loops, offering a rigorous theoretical foundation that advances the field of causal inference beyond acyclic assumptions.
The paper proposes a rigorous theoretical framework for causal discovery in cyclic linear non-Gaussian (LiNG) models. The core contribution is the establishment of the first information-theoretic lower bounds on the sample complexity for recovering the condensation (SCC partition and inter-component edges). The authors introduce a "block-exogeneity" principle that allows for the identification of root SCCs via residual independence without requiring correlation faithfulness, a significant relaxation of standard assumptions. The proposed algorithm, BlockExo, utilizes sparse adjustment sets to identify SCCs and their external parents, achieving a sample complexity that matches the derived lower bounds structurally. The methodology is mathematically sound, leveraging the Darmois-Skitovitch theorem and Fano's inequality to bridge the gap between population-level identifiability and finite-sample guarantees.
The experimental section is limited but appropriate for a theory-heavy paper. It validates the structural dependence of the sample bound by showing that recovery curves align when normalized by the theoretical factors. It compares BlockExo against existing methods like Coarsening, DisjointCycles, and StableSpIn. BlockExo demonstrates superior sample efficiency in exact recovery tasks, particularly in overlapping cycle structures where baselines fail or require significantly more samples. However, the experiments are conducted on synthetic data only, with small dimensionality ($p=50$) and limited noise distributions, which restricts the generalizability of the empirical claims.
The paper provides a reproducibility statement indicating that code, configurations, and scripts will be released. The algorithmic details are clear, and the theoretical assumptions are explicitly stated. However, the lack of publicly available code at the time of review and the reliance on specific, potentially hard-to-tune thresholds (e.g., in the $PASS$ condition) may pose challenges for independent replication without the authors' implementation.
The primary limitation is the restriction to linear non-Gaussian models with causal sufficiency (no latent confounders). The method assumes principal invertibility, which, while standard, excludes certain unstable cyclic systems. The computational complexity of BlockExo is exponential in the sum of SCC size and external parent count ($p^{O(s_{max}+d_B)}$), which limits its applicability to dense or large-scale cyclic structures. Furthermore, the empirical evaluation lacks real-world data validation, leaving the practical utility of the method in complex biological or economic systems unproven.
This work provides a foundational statistical understanding of the costs associated with learning cyclic causal structures, a critical area in systems biology and economics. By establishing minimax optimality for condensation recovery, it sets a benchmark for future algorithms. The block-exogeneity principle offers a new tool for causal discovery that does not rely on faithfulness, potentially influencing the design of more robust causal inference methods in the broader machine learning community. The paper establishes the first information-theoretic lower bounds for sample complexity in cyclic causal discovery and introduces BlockExo, an algorithm that achieves structurally optimal recovery. By decoupling the costs of SCC membership and external parent selection, the work provides a precise characterization of the statistical difficulty of learning feedback loops, offering a rigorous theoretical foundation that advances the field of causal inference beyond acyclic assumptions.
Neural operators are typically trained in a supervised fashion, which requires a dataset to be generated with a classical solver. Training them physics-informed, i.e., purely from the governing equations, removes this large offline cost and allows fresh samples to be drawn at every optimization step, but has so far been limited to simplified problems and trails supervised training in accuracy. The obstacle is the ill-conditioning of physics-informed losses, which differential operators induce and which worsens as the discretization is refined. We therefore propose a preconditioned residual loss function and show mesh-independent conditioning for elliptic problems and greatly improved conditioning for saddle point problems. Realized through geometric and algebraic multigrid, the construction applies to linear and nonlinear equations, steady or time-dependent, on structured and unstructured meshes, is agnostic to the neural operator architecture, and adds no cost at inference. On the Poisson, Allen-Cahn and stationary Stokes equations, the resulting label-free training matches supervised training and is four to twenty-five times more accurate than previous physics-informed operator learning methods.
Primary: ETH Zurich
All Institutions: ETH Zurich
The paper proposes a preconditioned physics-informed loss function using multigrid methods to solve the ill-conditioning problem in neural operator training, achieving label-free accuracy that matches supervised learning. This is a significant technical contribution that bridges numerical linear algebra and deep learning, offering a practical and scalable solution to a long-standing optimization challenge in physics-informed machine learning.
The paper addresses a critical bottleneck in physics-informed neural operator (PINO) training: the severe ill-conditioning of residual losses induced by differential operators, which scales poorly with mesh refinement ($O(h^{-4})$ for elliptic problems). The authors propose a preconditioned least-squares loss function that incorporates multigrid preconditioners (geometric for structured grids, algebraic for unstructured) directly into the loss landscape. This approach is elegant because it decouples the preconditioning from the neural network architecture, requiring no changes to the model structure and adding zero cost at inference. The method generalizes to nonlinear problems (Allen-Cahn) and saddle-point problems (Stokes) by using approximate Jacobian inverses and block-diagonal preconditioners, respectively. The theoretical analysis correctly identifies that preconditioning the residual effectively transforms the Hessian of the loss to be mesh-independent or significantly better conditioned, thereby enabling efficient gradient-based optimization.
The experiments are rigorous and cover three distinct PDE classes: Poisson (linear elliptic), Allen-Cahn (nonlinear parabolic), and Stokes (saddle-point). The use of both FNO (Fourier Neural Operator) and GAOT (Geometry-Aware Operator Transformer) demonstrates architecture agnosticism. The results are compelling: the proposed method matches or surpasses supervised training accuracy while requiring no labeled data, and outperforms standard PINO and PI-DeepONet baselines by factors of 4 to 25. The "infinite data limit" experiment is particularly strong, showing that fresh sampling at each step removes the data-size bottleneck of supervised learning. The inclusion of unstructured meshes for the Stokes problem adds significant practical relevance, as many real-world PDEs do not admit structured grids.
The paper provides a public GitHub repository with code. The appendices contain detailed descriptions of the discretization, preconditioner construction (including specific multigrid parameters), and optimization hyperparameters. The use of standard libraries (neuraloperator, AMGX) and clear algorithmic descriptions (e.g., the V-cycle algorithm) ensures high reproducibility. The authors also provide a clear distinction between the "interpolated neural operator" and the raw network output, clarifying how the loss is computed.
The primary limitation is the dependence on the availability of an efficient preconditioner for the specific PDE class. For problems where multigrid or other fast solvers are not readily available (e.g., highly convection-dominated flows or complex wave equations), the method's advantage diminishes. Additionally, the method relies on an explicit finite element discretization of the residual, which may not be straightforward for all PDE formulations or black-box physics. The paper also notes that for the Stokes problem, the conditioning is only improved to $O(h^{-2})$, not $O(1)$, though this is still sufficient for convergence.
This work has high potential impact on the field of scientific machine learning. By enabling label-free training of neural operators that matches supervised performance, it removes the need for expensive offline data generation via classical solvers. This is particularly significant for physics foundation models and applications where generating high-fidelity training data is prohibitive. The technique is broadly applicable to any PDE where a fast linear solver or preconditioner exists, making it a valuable tool for the broader community of PDE surrogates. The paper proposes a preconditioned physics-informed loss function using multigrid methods to solve the ill-conditioning problem in neural operator training, achieving label-free accuracy that matches supervised learning. This is a significant technical contribution that bridges numerical linear algebra and deep learning, offering a practical and scalable solution to a long-standing optimization challenge in physics-informed machine learning.
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
Primary: Microsoft Research
All Institutions: Microsoft Research, New York University, Stanford University
The paper introduces "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs with minimal overhead. By shifting the burden of verification to the trainer via lightweight behavioral fingerprints and replay challenges, it addresses the reproducibility crisis in ML, enabling trusted baselines and scalable auditing of billion-parameter models.
The paper proposes "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs. The core innovation is shifting the burden of verification from the reader (who must reproduce expensive runs) to the trainer (who generates a lightweight, cryptographically signed "tape" of training steps). The method utilizes three key technical components: (1) a quasi-geometric parameter commitment using low-rank orthonormal projections to create fast behavioral fingerprints of model weights; (2) a non-geometric hash for optimizer states to ensure state continuity; and (3) a probabilistic replay challenge mechanism where the verifier randomly samples steps to verify against the declared training function. The system leverages stateless accelerator serializations (JAX/Torch) to seal the computation and uses white-box cryptography or zkVMs to anchor trust in the attestation engine. The theoretical contribution includes proofs of completeness and soundness, demonstrating that the probability of a malicious trainer successfully hiding invalid updates is bounded and amplifiable.
The authors validate the system on GPT-style models ranging from 100M to 2B parameters, testing both Data Parallel (DDP) and Fully Sharded Data Parallel (FSDP) configurations. Key results include a minimal overhead of 1.19% in speed and 1.11% in memory for a 2B parameter run, which is remarkably low for a security-critical system. The paper demonstrates high detection rates for data poisoning, parameter corruption, and adversarial optimization attacks against the sketch. The inclusion of a "self-regulating leaderboard" (dvbench.org) provides a practical ecosystem for adoption, allowing the community to submit and verify certified runs. The experiments convincingly show that the method scales to billion-parameter models without prohibitive cost.
The paper provides high reproducibility by releasing the code and a public leaderboard. The methodology is code-base agnostic, relying on standard serialization formats (StableHLO/XLA) which are widely supported. The detailed description of the cryptographic primitives (BLAKE3, ratcheted signatures) and the specific implementation details (Rust workers, Python facade) allow for independent verification. The open-source nature of the leaderboard further enhances reproducibility by providing real-world examples of certified runs.
The primary limitation is the reliance on a trusted attestation engine ($B_{att}$). While the paper argues for white-box cryptography or zkVMs, these are complex engineering solutions that may be vulnerable to side-channel attacks or implementation bugs not covered by the theoretical model. Additionally, the system assumes the training function $f$ is declared and bounded; it does not protect against logical errors in the training algorithm itself, only against deviations from the declared algorithm. The overhead, while low, is non-zero and may be significant for extremely latency-sensitive applications. The method is currently focused on pretraining and may require adaptation for fine-tuning or reinforcement learning workflows.
This work has the potential to fundamentally change how machine learning research is conducted and evaluated. By providing a technical solution to the reproducibility crisis, it enables standardized comparisons and trusted baselines, which are currently lacking in the field. The leaderboard initiative could foster a culture of verified results, reducing the impact of "slop" contributions and increasing the reliability of scientific claims. This could lead to more efficient use of computational resources, as researchers can trust published results without duplicating expensive training runs. The broader impact extends to AI safety and governance, as certified training runs provide an audit trail for model behavior and data usage. The paper introduces "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs with minimal overhead. By shifting the burden of verification to the trainer via lightweight behavioral fingerprints and replay challenges, it addresses the reproducibility crisis in ML, enabling trusted baselines and scalable auditing of billion-parameter models.
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
Primary: University of Chinese Academy of Sciences
All Institutions: University of Chinese Academy of Sciences, Meituan
The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
The paper proposes TGRL, a method that integrates temperature-based exploration into the RLVR training loop by treating temperature-induced diversity as a learnable signal. The core mechanism involves partitioning rollouts into low-temperature (reference) and high-temperature (exploration) groups. The method estimates "exploration gain" via the reward contrast between these groups and uses Jensen-Shannon (JS) divergence between temperature-scaled token distributions to allocate this gain as token-level credit. The theoretical analysis provides bounds on how JS divergence captures local sensitivity to temperature changes, justifying the credit allocation strategy. The approach is logically sound, bridging the gap between sampling diversity and policy gradient updates in a principled manner.
The experimental evaluation is extensive, covering 11 benchmarks across mathematical reasoning, code generation, and agentic tasks. The model sizes tested (Qwen3-4B, 14B, 32B) are relevant to current LLM research. The results show consistent improvements over strong baselines like GRPO and DAPO, with specific gains in CodeForces rating and LiveCodeBench Pass@16. The ablation studies effectively isolate the contributions of the temperature grouping and JS-based credit allocation. The wall-clock analysis demonstrating 36% faster convergence is a strong practical contribution.
The paper provides a public GitHub repository and detailed hyperparameters (temperatures, rollout budgets, warmup steps). The use of standard models (Qwen3) and public benchmarks enhances reproducibility. However, the specific implementation details of the JS divergence calculation and the exact handling of the mixed-group advantage normalization would require careful inspection of the code to fully replicate.
The method relies on the assumption that temperature scaling is a sufficient proxy for exploration diversity. It may not capture all forms of beneficial exploration, such as those requiring semantic shifts rather than just stochastic sampling. The computational overhead of computing JS divergence for every token in the high-temperature group could be significant for very long contexts, though the paper claims efficiency gains. The performance gains, while consistent, are moderate (e.g., 1.6% on math average), which may limit its adoption if simpler baselines are sufficient for many tasks.
This work contributes to the understanding of how to efficiently utilize rollout budgets in LLM post-training. By providing a mechanism to quantify and exploit exploration gain, it offers a path toward more sample-efficient RLVR training. The insights into token-level credit allocation based on distributional sensitivity could be applicable to other areas of LLM optimization, such as test-time compute scaling. The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis. Debate beats single-agent inference (3--7 points where tasks have headroom) but at matched budget conditions it ties or even loses to self-consistency sampling at 1.6$\times$ the wall-clock and 3.4$\times$ the token cost. Persona prompting reduces accuracy and a dose-response experiment over each model's full combinatorial persona space shows the cost is a persona tax, not a diversity tax: redundant personas hurt most, while maximally-diverse teams recover part of the loss. Furthermore, mixed-model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate's benefit comes from the first exchange of answers. We further identify a pervasive measurement hazard in which debate transcripts silently overflow serving context windows, whose correction alone moves our debate-versus-sampling comparison from $-1.8$ points to parity. Our results recast reported MAD gains as an ensemble-sampling effect and provide the budget-matched, contamination-checked baseline bar that future debate mechanisms should be required to clear.
Primary: Harvard University
All Institutions: Boston Children's Hospital, Harvard College, Harvard Medical School, Harvard School of Engineering and Applied Sciences, Cincinnati Children's Hospital
The paper rigorously refutes the hypothesis that cognitive diversity drives the benefits of multi-agent debate, demonstrating that gains are primarily due to ensemble sampling effects and identifying a critical measurement hazard related to context-window overflow. Through extensive experiments with 23 open-weight models and budget-matched controls, it provides a robust baseline for future research and highlights the inefficiency of standard debate protocols compared to simple self-consistency sampling.
The paper employs a rigorous experimental design to test the "cognitive diversity" hypothesis in Multi-Agent Debate (MAD). The methodology is strong, utilizing a large pool of 23 open-weight models across 11 vendor families to ensure generalizability. A key methodological strength is the use of generation-budget-matched controls (self-consistency sampling) rather than just comparing against single-agent inference, which isolates the effect of deliberation from the effect of increased compute. The introduction of a "persona tax" analysis via a dose-response experiment over the combinatorial persona space is a sophisticated approach to disentangling the cost of role-playing from the benefit of diversity. The identification of context-window overflow as a confounding variable is a critical methodological insight that adds significant value to the field's understanding of LLM evaluation pitfalls.
The experimental scale is impressive, with over 5,500 debate and control runs across five tasks. The results are consistent and statistically significant, showing that MAD gains are largely attributable to ensemble sampling effects rather than cognitive diversity. The finding that persona prompting reduces accuracy in small models is a surprising and important empirical result. The evaluation of heterogeneous model teams against majority voting controls further supports the conclusion that diversity does not inherently improve debate performance. The inclusion of a biography generation task with LLM judges adds depth, although the reliance on LLM judges introduces some subjectivity, which the authors acknowledge and mitigate with vendor-balanced panels.
The paper provides a high level of reproducibility. All models are open-weight, and the authors release the full budget-matched, contamination-checked grid as a baseline. The use of deterministic scripts for figure and table generation, along with released CSVs for scored outputs and persona ladders, ensures that the results can be verified. The explicit mention of fixed bootstrap seeds and the assertion that persona-team selection is exactly reproducible via programmatic checks further enhance reproducibility.
The primary limitation is the focus on small open-weight models. The authors acknowledge that larger frontier models might behave differently, particularly regarding the "persona tax," as they have more capacity to adopt unconventional cognitive frames. The study is also limited to specific tasks (math, knowledge, biography), and the results may not generalize to domains requiring novelty or synthesis. The reliance on LLM judges for the biography task, while mitigated, still introduces potential bias. Additionally, the study does not explore structured single-agent test-time methods or objectively verifiable domains like code generation.
This paper has significant broader impact by challenging a popular assumption in the LLM community regarding multi-agent debate. By demonstrating that MAD gains are largely an ensemble-sampling effect, it redirects research focus towards more efficient sampling strategies and rigorous budget-matched evaluations. The identification of context-window overflow as a measurement hazard is a valuable contribution that will likely influence future experimental designs. The paper provides a clear baseline for future debate mechanisms, encouraging more rigorous comparison against simple sampling controls. The paper rigorously refutes the hypothesis that cognitive diversity drives the benefits of multi-agent debate, demonstrating that gains are primarily due to ensemble sampling effects and identifying a critical measurement hazard related to context-window overflow. Through extensive experiments with 23 open-weight models and budget-matched controls, it provides a robust baseline for future research and highlights the inefficiency of standard debate protocols compared to simple self-consistency sampling.
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $α$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
Primary: Stanford University
All Institutions: Stanford University
The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
The paper introduces "rank confidence sequences," a rigorous statistical framework for evaluating model rankings in a sequential, anytime-valid manner. The core innovation lies in combining betting e-processes for pairwise comparisons with closed testing over weak orderings (permutations with ties). This approach allows the construction of confidence sets for the ranks of all models simultaneously, valid at any time step, without assuming independence between model scores on the same items (a common issue in benchmarking). The method handles both fixed finite benchmarks (with random item ordering) and superpopulation settings. The theoretical contribution is significant, providing finite-sample guarantees that hold under any stopping rule, which addresses the "peeking" problem inherent in real-time leaderboard monitoring. The use of integer programming for exact certification at scale is a clever computational addition, though the coNP-hardness of the general problem is acknowledged.
The experiments are well-designed to validate the theoretical claims. E1 demonstrates the failure of fixed-sample methods under repeated monitoring, showing a 30% error rate vs. the nominal 5%. E2 applies the method to real LLM leaderboard data (Open LLM Leaderboard), showing that while early certification is limited, it becomes robust as more items are revealed. E3 highlights the practical benefit of compute savings, showing that early stopping for specific models (e.g., top-3 certification) can save significant evaluation costs without compromising validity. E4 compares power against fixed-sample methods at pre-planned look times, showing comparable performance. The use of real-world LLM data adds substantial practical relevance.
The paper provides a GitHub link to the code. The algorithms are described in detail, including the betting strategies, the offset calculations, and the integer programming formulations. The specific parameters for the betting grid are provided. The experimental setups are described with references to public datasets. The level of detail suggests high reproducibility for researchers with statistical programming skills.
The method requires per-item scores in [0,1] and assumes a fixed set of models evaluated on common items. It does not handle adaptive item selection or models arriving over time. The computational cost, while manageable for typical leaderboard sizes (up to ~50 models), grows with the number of models and items, and the exact certification via integer programming can be complex to implement correctly. The method is primarily useful for ranking, not for estimating the absolute performance gap with high precision in early stages.
This work has high potential impact on the ML evaluation community. As leaderboards become more dynamic and expensive to run, the need for statistically valid, anytime-valid monitoring is critical. This paper provides a principled way to do so, potentially changing how benchmarks are reported and interpreted. It bridges the gap between statistical theory (e-processes, closed testing) and practical ML evaluation. The compute savings aspect is also a strong practical incentive for adoption. The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(π_θ\Vert π_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Primary: Unknown (Affiliations not explicitly listed in provided text, but authors Qinfeng Li, Wenqi Zhang, Guoqing Jiang, Liwei Chen, Xuanping Li, Zhiheng Qin, Yuntai Bao, Xuhong Zhang are associated with Alibaba Group / Qwen Team)
All Institutions: Alibaba Group (Inferred from author list and Qwen model usage)
The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The paper employs a rigorous empirical scaling analysis framework, adapting the methodology of reward model overoptimization studies (Gao et al., 2023) to the domain of On-Policy Distillation (OPD). The core methodological contribution is the characterization of OPD training dynamics as a function of the square root of token-level reverse KL divergence ($d$). The authors identify a "useful-transfer" regime where gold score increases linearly with $d$, followed by a noisy tail. They fit power laws to predict peak performance ($G_{peak}$) and transfer rates based on student/teacher scale and teacher quality. The inclusion of a theoretical derivation (Appendix) explaining why local KL geometry leads to linear transfer in $d$ adds significant depth, linking empirical observations to Fisher information geometry. The comparison between Vanilla-OPD and Delta-OPD (using policy shift contrasts) is well-motivated and clearly defined.
The experimental design is comprehensive, covering 25 teacher-student combinations across Qwen2.5 models (0.5B-14B) in weak-to-strong, same-base, and strong-to-weak configurations. The use of a single model family (Qwen2.5) is a limitation but allows for controlled isolation of scale effects. The evaluation on math reasoning (GSM8K/MATH) is standard but sufficient for this type of scaling study. Key findings include: (1) Peak student error is proportional to teacher remaining error, (2) Smaller teachers transfer better at matched scores (counter-intuitive and significant), (3) Bootstrapping does not improve over direct transfer from the smallest expert, and (4) Off-policy cold starts harm weak-to-strong transfer. The validation via leave-one-scale-out prediction is a strong methodological choice that demonstrates the predictive power of the fitted laws.
The paper provides detailed hyperparameters in the appendix and specifies the use of the `verl` framework. However, the reliance on a single random seed for all runs is a significant reproducibility weakness, although the authors justify this by citing standard practices in scaling law studies (Kaplan et al., Hoffmann et al.). The lack of seed variance estimation limits the confidence in the precise coefficients of the power laws, though the trends appear robust across the grid. The code and data are not explicitly linked in the provided text, which is a minor gap for immediate reproducibility.
The primary limitation is the restriction to a single model family (Qwen2.5) and a single task domain (math reasoning). The authors acknowledge that whether these scaling laws hold for other architectures, tasks, or post-training recipes is untested. The single-seed constraint means that stochastic variance in RL/OPD training is not captured, potentially masking instability in the "noisy tail" dynamics. The theoretical derivation assumes smoothness and differentiability that may not hold strictly in discrete token spaces or with clipping mechanisms, though the empirical fit is strong.
This paper has high practical impact for LLM practitioners. By providing predictive scaling laws for OPD, it enables engineers to estimate the outcome of distillation runs before committing significant compute resources. The finding that smaller teachers can be more effective than larger ones at matched scores challenges common assumptions about teacher quality and offers a cost-effective strategy for model family development. The negative result on bootstrapping is also valuable, preventing wasted effort on a seemingly intuitive but ineffective strategy. This work bridges the gap between theoretical scaling laws and practical post-training pipelines. The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.
Primary: NVIDIA
All Institutions: NVIDIA
The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
The paper employs a rigorous ablation study combined with theoretical analysis to deconstruct the Transolver architecture. The methodology is sound, isolating the "physics-attention" mechanism into slicing/deslicing, token attention, and pointwise MLPs. The theoretical contribution leverages the theory of Averaging Neural Operators (ANO) to prove that the slicing/deslicing component, even with a constant core (no attention), is sufficient for universal approximation of continuous operators. This provides a strong mathematical foundation for the empirical finding that the Transformer component is redundant. The implementation of a FlashAttention-style kernel for the slicing operation is a significant technical contribution, addressing the memory bottleneck of materializing slice weights.
The experiments are extensive, covering nine challenging 3D fluid dynamics benchmarks, including industrial-scale aerodynamics (DrivAerNet++, SHIFT-SUV, SHIFT-Wing, DrivAerML). The evaluation protocol is careful, using matched step budgets and identical data pipelines. The results clearly demonstrate that removing token attention does not degrade accuracy, while removing the global mixing (slicing) causes performance collapse. The efficiency gains from the new kernel are substantial, showing significant memory savings and speedups, particularly at large slice counts.
The paper provides detailed descriptions of the ablation variants, training protocols, and kernel implementation. The use of standard datasets and clear metric definitions (relative L1 error) supports reproducibility. The code for the FlashSlice kernel is described in detail, though a direct link is not provided in the text snippet, the description is sufficient for implementation.
The theory is based on expressivity (universality) and does not address generalization or optimization dynamics. The bounds are not sharp. The findings are specific to the Transolver architecture and may not generalize to other neural operators without further analysis. The paper acknowledges that the cost of removing attention is empirically zero, but the theory only makes it unsurprising, not derived from first principles.
This paper has high impact on the scientific machine learning community by clarifying the fundamental mechanisms of a widely used architecture. It challenges the assumption that self-attention is necessary for global mixing in operator learning, potentially leading to more efficient and simpler models. The efficient kernel implementation will benefit practitioners working with large-scale unstructured meshes. The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
Primary: The Ohio State University
All Institutions: The Ohio State University, RWTH Aachen University
Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
The paper introduces "belief geometry," a unified analytical framework to compare Transformers and State-Space Models (SSMs) in the context of in-context linear regression (ICLR). The methodology is rigorous, moving beyond empirical benchmarks to derive theoretical lower and upper bounds on cumulative Bayes regret. It decomposes sequential learning into three capabilities: evidence assembly, belief maintenance, and addressing. The authors define specific architectural classes (linear/softmax attention, fixed/selective SSMs) and prove sharp separations: SSMs are optimal for stationary belief maintenance (exponential kernels vs. uniform attention), SSMs have a memory advantage for positional assembly (convolution vs. context window), and softmax attention has an exponential width advantage for content addressing (selecting from retained tokens vs. storing candidates in state). The use of cumulative regret rather than terminal loss is a significant methodological improvement for analyzing sequential dynamics.
The experiments validate the theoretical predictions using practical architectures (LLaMA-type Transformers and Mamba-2). The authors test kernel alignment in single-layer models, showing that Transformers learn flatter kernels while Mamba-2 learns exponential ones, matching the theoretical optima. For the routing tasks (positional and content), they demonstrate that performance thresholds align with the theoretical resource requirements (e.g., convolution width for SSMs, context length for Transformers). The experiments effectively bridge the gap between the analytically tractable linear regression testbed and practical non-linear models, confirming that the architectural lessons (e.g., exponential width gap for addressing) hold in practice.
The paper includes a reproducibility statement and provides a link to an anonymous code repository. The experimental protocols, including hyperparameter sweeps and evaluation metrics, are detailed in the appendices. The use of synthetic tasks (Gaussian filtering, Beta-Bernoulli bandits, logistic bandits) ensures that the results are deterministic and reproducible without access to large-scale datasets.
The primary limitation is the reliance on linear regression and conjugate/non-conjugate bandit settings for the theoretical analysis. While the authors argue that the "belief geometry" extends to broader problems, the proofs are specific to these linear/quadratic loss structures. The "exponential width advantage" for attention in content addressing is a strong claim, but it relies on specific definitions of "addressing capacity" and compressed codebooks; real-world semantic addressing may not strictly follow this geometric separation. Additionally, the comparison focuses on representational capability (what the model *can* do) rather than optimization dynamics (what the model *learns* efficiently), though the kernel alignment experiments partially address this.
This paper provides a principled theoretical foundation for the ongoing debate between Transformers and SSMs. By identifying specific architectural mechanisms (softmax vs. selective transitions, context window vs. recurrent state) that confer advantages in different task regimes, it offers actionable insights for architecture design. The framework of "belief geometry" could be extended to other sequential tasks, such as language modeling or control, potentially guiding the development of hybrid architectures that leverage the strengths of both families. The finding that SSMs are inherently better at stationary belief maintenance while attention is better at content addressing helps explain empirical observations in long-context learning and associative recall. Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
[One sentence main contribution]. The paper proposes Next-Embedding Predictive Autoregression (NEPA) to dynamically predict conditioning embeddings for diffusion transformers, achieving a FID of 1.32 on ImageNet 256x256 with significantly reduced training compute compared to REPA. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant as it introduces a novel way to handle conditioning in diffusion models, moving away from static embeddings to dynamic, state-dependent predictions. The methodology is sound, leveraging the autoregressive nature of embedding prediction to enhance the generative process. The significance to the field is high due to the reported efficiency gains and competitive performance, which could influence future designs of diffusion-based generators.
The paper introduces Next-Embedding Predictive Autoregression (NEPA), a method that trains a Transformer to predict the next continuous embedding in a sequence. The core innovation lies in applying this to diffusion transformers (DiT), where the clean image embeddings are predicted from the noisy image and condition, and these predictions are used as the conditioning signal for the generator at every denoising step. This allows the conditioning to adapt dynamically to the current noisy state, rather than being static. The method involves a two-network approach: a NEPA model for prediction and a DiT generator for synthesis. The combination with REPA (Representation Alignment) is highlighted as a key component for achieving strong results.
Experiments are conducted on class-conditional ImageNet 256x256. The paper reports an FID of 1.32 for the final model, NEPA-DiT-XL, which is a competitive result. The paper claims this is achieved using about a third of the training compute of REPA. Ablation studies are mentioned regarding the condition of the generator, the design of Multi-Embedding Prediction, and scaling laws for both models. The evaluation focuses on the quality of generation and the efficiency of the training process.
The paper is a preprint from arXiv. While it describes the methodology in detail, the lack of a public code repository or project page at the time of review limits immediate reproducibility. The specific implementation details of the NEPA model and its integration with DiT are described, but without code, full verification is difficult.
The primary limitation is the addition of a second network to every sampling step, which increases inference latency. The paper only evaluates on class-conditional ImageNet, leaving open questions about performance on text-to-image tasks or other datasets. The reliance on REPA for the final strong result suggests that NEPA alone may not be sufficient for state-of-the-art performance.
The work contributes to the understanding of how conditioning signals can be dynamically generated in diffusion models. The concept of predicting embeddings as a form of adaptive conditioning could be extended to other generative tasks and modalities. The efficiency gains in training compute are significant for the field, as they lower the barrier to entry for high-quality image generation. [One sentence main contribution]. The paper proposes Next-Embedding Predictive Autoregression (NEPA) to dynamically predict conditioning embeddings for diffusion transformers, achieving a FID of 1.32 on ImageNet 256x256 with significantly reduced training compute compared to REPA. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant as it introduces a novel way to handle conditioning in diffusion models, moving away from static embeddings to dynamic, state-dependent predictions. The methodology is sound, leveraging the autoregressive nature of embedding prediction to enhance the generative process. The significance to the field is high due to the reported efficiency gains and competitive performance, which could influence future designs of diffusion-based generators.
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
Primary: UC Berkeley
All Institutions: UC Berkeley, Stanford University, University of Washington, Impossible Research, Google DeepMind
The paper introduces OmniTaskonomy, a comprehensive taxonomy and empirical study demonstrating that visual generation tasks can significantly improve visual understanding capabilities through selective and task-dependent transfer. By systematically analyzing 19 generation tasks and 25 understanding capabilities, the authors reveal both intuitive and surprising connections, supported by gradient alignment analysis, providing a rigorous roadmap for leveraging generation supervision in multimodal learning.
The paper proposes a systematic framework, OmniTaskonomy, to analyze the transferability of visual generation tasks to visual understanding tasks. The methodology is rigorous, employing controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that share underlying semantic structures. By constructing a unified taxonomy of 19 generation tasks and 25 understanding capabilities, the authors move beyond anecdotal evidence to a comprehensive empirical study. A key methodological strength is the use of gradient alignment analysis to correlate with downstream performance gains, providing a mechanistic explanation for the observed transfer patterns.
The experimental evaluation is extensive, covering a wide range of tasks and datasets. The paper demonstrates that under specific training recipes, I2I training significantly improves I2T performance, with gains scaling with data volume. The transfer map reveals both intuitive connections (e.g., depth prediction aiding 3D reasoning) and surprising ones (e.g., 2.5D segmentation aiding category recognition). The results are consistent and well-supported by statistical analysis, providing a clear roadmap for curriculum learning in multimodal models.
The paper includes a detailed reproducibility statement, with appendices providing task definitions, training recipes, taxonomy annotations, and complete transfer results. The availability of numerical exports and figure-generation code further enhances reproducibility. The use of standard benchmarks and clear experimental protocols ensures that the findings can be verified by the community.
The study is primarily focused on image-based tasks and may not fully generalize to video or other modalities. The analysis relies on specific model architectures and training setups, which might limit the generalizability of the findings to other model families. Additionally, the computational cost of training and evaluating such a large number of task pairs is significant, which could be a barrier for some researchers.
The findings have significant implications for the design of multimodal learning systems, suggesting that visual generation can serve as a powerful pre-training signal for visual understanding. This could lead to more efficient and effective training curricula for large multimodal models. The paper also provides a valuable resource for the community in the form of the OmniTaskonomy taxonomy, which can be used to guide future research on task transfer and curriculum learning. The paper introduces OmniTaskonomy, a comprehensive taxonomy and empirical study demonstrating that visual generation tasks can significantly improve visual understanding capabilities through selective and task-dependent transfer. By systematically analyzing 19 generation tasks and 25 understanding capabilities, the authors reveal both intuitive and surprising connections, supported by gradient alignment analysis, providing a rigorous roadmap for leveraging generation supervision in multimodal learning.
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Primary: Tsinghua University
All Institutions: Tsinghua University, Shanghai Jiao Tong University, Fudan University, University of Science and Technology of China, Washington University in St. Louis, The Institute of Artificial Intelligence, China Telecom (TeleAI)
V-JEPA Policy demonstrates that frozen predictive visual latents are a sufficient foundation for effective world-action models, outperforming generative and discriminative alternatives under distribution shifts. The paper rigorously validates this approach through extensive simulation benchmarks and real-world manipulation, highlighting the critical role of predictor pretraining on video-instruction pairs for robust generalization.
The paper proposes V-JEPA Policy, a framework that decouples the predictive visual latent space from the generative model. Instead of fine-tuning a large video generator, it freezes the V-JEPA 2.1 encoder and trains a lightweight instruction-conditioned future-latent predictor and a flow-matching action expert from scratch. The key architectural innovation is the use of the predictor's future-informed context key-value states to condition the action generation, effectively creating a world-action model (WAM) that relies on predictive latents rather than pixel-space reconstruction or full video generation. This approach is computationally efficient (0.9B total params, 0.6B trainable) and leverages the robustness of predictive pretraining.
The experiments are extensive, covering LIBERO, LIBERO-Plus, and RoboCasa-GR1 benchmarks. The paper provides rigorous ablations comparing V-JEPA latents against discriminative, reconstructive, and video-understanding-oriented alternatives, showing superior performance under distribution shifts. A significant finding is the benefit of pretraining the predictor on DROID video-instruction pairs without action labels, which yields substantial gains in downstream control and out-of-distribution generalization that cannot be recovered by simply training longer from scratch. The results are competitive with state-of-the-art WAMs and VLA baselines.
The authors provide a public GitHub repository with code. The paper details the training setup, including the separation of frozen and trainable parameters, and the specific pretraining data (DROID). The use of standard benchmarks and open-source models (V-JEPA 2.1) enhances reproducibility.
The method relies heavily on the quality of the V-JEPA 2.1 encoder; if the encoder fails to capture relevant dynamics, the policy will suffer. The flow-matching action expert is trained from scratch, which may require careful tuning. The paper focuses on simulation and limited real-world bimanual manipulation; broader real-world deployment across diverse environments is not fully explored. The reliance on DROID for pretraining limits the immediate applicability to domains similar to DROID.
This work provides a compelling argument for using predictive latents as a foundation for robot learning, potentially reducing the computational cost of training WAMs. It offers a pathway to leverage large-scale video pretraining without the burden of generative model fine-tuning. The findings on the transferability of future-modeling knowledge from in-the-wild videos to control tasks are significant for the robotics community. V-JEPA Policy demonstrates that frozen predictive visual latents are a sufficient foundation for effective world-action models, outperforming generative and discriminative alternatives under distribution shifts. The paper rigorously validates this approach through extensive simulation benchmarks and real-world manipulation, highlighting the critical role of predictor pretraining on video-instruction pairs for robust generalization.
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Primary: University of California, Los Angeles (UCLA)
All Institutions: University of California, Los Angeles (UCLA)
The paper introduces a novel RL framework for unified multimodal models that jointly optimizes reflection and generation, achieving significant improvements in image generation accuracy through self-correction. By leveraging group-relative advantage estimation and graded rewards, the method effectively addresses the sparsity of feedback in compositional generation tasks, demonstrating that unified models can learn to diagnose and repair their own errors without external critics, with gains that transfer robustly to unseen benchmarks.
The paper proposes UMM-Reflection, a framework for applying Reinforcement Learning (RL) to unified multimodal models (specifically BAGEL) to improve image generation via self-reflection. The core methodological contribution is the joint optimization of reflection text (diagnosis) and image generation (repair) within a single model using group-relative advantage estimation. By sampling sibling trajectories from the same initial image, the method avoids the combinatorial explosion of per-round credit assignment and allows credit to flow across the entire reflection loop. A significant technical detail is the "graded reward" mechanism, which addresses the sparsity of binary GenEval scores by providing partial credit for partial constraint satisfaction, thereby stabilizing RL training. The approach is theoretically sound, leveraging the unified nature of the model to eliminate the need for external critics or separate verifier models at inference time.
The experimental evaluation is rigorous and comprehensive. The paper reports a substantial improvement of +12.05 points on GenEval over the SFT baseline. Crucially, the authors demonstrate that these gains transfer to unseen benchmarks (WISE, OneIG-Bench, T2I-CompBench++), suggesting the model has learned generalizable repair strategies rather than overfitting to the training distribution. Ablation studies are thorough, including comparisons against direct T2I-RL (without reflection), external critic pipelines (GPT-5.5), and different reward shaping strategies (penalty vs. bonus for stopping). The analysis of "pass@16" vs "pass@1" highlights that the SFT model already contains the capability for correct repairs, and RL effectively selects and reinforces these high-success paths. The visual pathway stability analysis confirms that RL primarily modifies the decision-making (text) pathway rather than disrupting the underlying visual generation capabilities.
The paper provides high reproducibility standards. It details the specific training hyperparameters (learning rates, batch sizes, number of updates), the composition of the RL prompt pool, and the exact reward calculation formulas. The authors explicitly state that the code and prompt pool will be released. The use of standard benchmarks (GenEval, WISE, etc.) and clear evaluation protocols (50 denoising steps, specific resolutions) allows for easy comparison with other works. The distinction between controlled evaluations and native leaderboard protocols is clearly defined, preventing misinterpretation of results.
The primary limitation is the reliance on the BAGEL model architecture; while the method is general, the specific implementation details (e.g., MoT decoder layers, flow-based revisions) are tied to this unified model structure. The counting category in GenEval shows no improvement, indicating that current reflection strategies are insufficient for precise numerical constraints. Additionally, the method requires significant computational resources (16 H100 GPUs for 33 hours) for the RL phase, which may limit adoption for smaller labs. The evaluation is limited to 512x512 resolution for training and evaluation, whereas the model natively supports 1024x1024, leaving open questions about performance at higher resolutions.
This work has significant implications for the development of autonomous multimodal agents. By demonstrating that a single model can effectively self-correct its generations through RL, it paves the way for more robust and reliable generative systems that do not rely on external, potentially misaligned, critic models. The insight that RL can "select" correct repairs from an existing SFT distribution rather than generating them from scratch is a valuable finding for the broader field of generative AI. The transferability of gains to unseen benchmarks suggests that self-reflection is a generalizable capability, potentially applicable to other modalities or tasks. The paper introduces a novel RL framework for unified multimodal models that jointly optimizes reflection and generation, achieving significant improvements in image generation accuracy through self-correction. By leveraging group-relative advantage estimation and graded rewards, the method effectively addresses the sparsity of feedback in compositional generation tasks, demonstrating that unified models can learn to diagnose and repair their own errors without external critics, with gains that transfer robustly to unseen benchmarks.
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Primary: Tencent
All Institutions: Tencent
The paper systematically characterizes the scaling laws of encoder-free MLLMs, predicting they will match encoder-based models at $10^{22}$ FLOPs. It provides a rigorous comparison using a matched model ladder and offers mechanistic insights into how decoders compensate for the lack of a visual encoder, positioning encoder-free architectures as a promising direction for future multimodal pretraining.
The paper employs a rigorous scaling law framework to compare encoder-free and encoder-based Multimodal Large Language Models (MLLMs). The methodology is sound, utilizing a matched ladder of 11 sparse MoE models (1.1B-44B) to isolate the effect of the visual encoder. The use of IsoFLOP profiles and compute-optimal allocation analysis is standard and well-executed. The introduction of "vision-specific adaptation" probes (attention patterns, layerwise representation evolution, expert routing) to mechanistically explain the scaling differences is a strong methodological contribution that goes beyond simple loss comparison. The derivation of the overtraining loss equation is mathematically sound and provides a useful tool for predicting efficiency gains under non-optimal training regimes.
The experimental setup is robust, covering a wide range of model scales and compute budgets ($10^{19}$ to $10^{21}$ FLOPs). The comparison is controlled by fixing the visual encoder size and data mixture. The results are consistent, showing that while encoder-free models lag at small scales, their loss decreases more rapidly with compute, predicting a crossover at $10^{22}$ FLOPs. The analysis by topic (STEM vs. Perception) provides valuable nuance, showing that the crossover is earlier for language-heavy tasks. The inclusion of downstream benchmark evaluations (CV-Bench, ChartQA, etc.) further validates the scaling trends, although the primary focus remains on validation loss.
The paper provides detailed implementation details in the appendix, including the specific architecture of the front ends, the MoE topology, and the training hyperparameters. The use of standard components (SigLIP 2, Muon optimizer) and clear descriptions of the data mixture enhance reproducibility. However, the specific data sources are not fully detailed, which may limit exact replication. The code is not explicitly linked in the provided text, but the level of detail suggests it is likely available or easily implementable.
The primary limitation is the reliance on extrapolation. The predicted crossover at $10^{22}$ FLOPs is well beyond the measured range ($10^{21}$ FLOPs), introducing uncertainty. The assumption that the visual encoder size remains fixed as the decoder scales is a simplification; in practice, joint scaling of the encoder and decoder might alter the dynamics. Additionally, the study focuses on a specific data mixture (1:1 text/multimodal), and results may vary with different data compositions. The "catch-up" prediction assumes that the irreducible loss floor is shared, which is a reasonable but unproven assumption at these scales.
This paper has significant implications for the design of future multimodal models. By demonstrating that the advantage of pretrained visual encoders diminishes with scale, it encourages the exploration of unified, encoder-free architectures that may be simpler and more efficient in the long run. The insights into how decoders adapt to take over visual encoding (e.g., bidirectional attention, early-layer processing) could inform the design of new decoder architectures specifically optimized for native multimodal learning. This work could shift the field's focus from improving visual encoders to optimizing the language model's ability to process raw visual tokens. The paper systematically characterizes the scaling laws of encoder-free MLLMs, predicting they will match encoder-based models at $10^{22}$ FLOPs. It provides a rigorous comparison using a matched model ladder and offers mechanistic insights into how decoders compensate for the lack of a visual encoder, positioning encoder-free architectures as a promising direction for future multimodal pretraining.
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
Primary: University of Illinois Urbana-Champaign (UIUC)
All Institutions: University of Illinois Urbana-Champaign (UIUC)
The paper identifies and characterizes "alignment-induced unfaithfulness," a critical failure mode where aligned LLMs silently override source content on sensitive topics, revealing a fundamental capability-alignment-faithfulness trilemma. Through rigorous controlled experiments, a novel dataset, and causal analysis of post-training stages, the authors demonstrate that this unfaithfulness is a systematic, emergent property of modern alignment techniques (particularly DPO) that worsens with model scale, posing significant risks to the reliability of LLMs in source-reporting applications.
The paper introduces a rigorous controlled experimental design to isolate the "alignment-faithfulness" conflict, distinct from capability-faithfulness or capability-alignment tradeoffs. By constructing the FaithConflict dataset with paired confirming/opposing claims in identical templates, the authors effectively control for surface form and domain, allowing the FaithGap metric to directly attribute unfaithfulness to the model's internal conflict with the source content. The dual taxonomy (B1-B8 for outputs, C0-C6 for reasoning) provides a granular lens into *how* models fail, distinguishing between visible refusals and dangerous silent inversions. The methodology is sound, though it relies heavily on an LLM judge (Qwen-2.5 32B) for annotation, which introduces a potential circularity if the judge shares similar alignment biases, though high inter-annotator agreement with humans mitigates this concern.
The experimental scope is impressive, covering 22 checkpoints across 8 model families (including frontier models like Claude Sonnet 4.6 and GPT-4o). The discovery of a "reverse scaling law"—where larger, more aligned models exhibit *greater* unfaithfulness on conflicting sources—is a significant and counter-intuitive finding. The stage-by-stage analysis pinpointing DPO as the primary driver of this behavior is particularly valuable for practitioners. The inclusion of causal interventions (removing safety data from post-training) strengthens the claim that alignment training, not just scale, is the root cause. However, the evaluation is limited to summarization and a few other formats (QA, NLI), and the reliance on a single judge model for all annotations is a notable weakness.
The paper provides high reproducibility standards. Code, data, and project pages are linked. The authors release the FaithConflict dataset and detailed prompts for both task execution and judging. The post-training intervention experiments use the public Tulu 3 pipeline, allowing others to replicate the causal analysis. The only minor gap is the lack of dated snapshot identifiers for frontier models, which is acknowledged by the authors.
The primary limitation is the reliance on an LLM judge for the core metric (FaithGap), which may not perfectly capture human perception of faithfulness. The study is also limited to source-reporting tasks (summarization, extraction) and does not explore whether this unfaithfulness manifests in more complex agentic or multi-turn dialogue settings. The causal attribution to DPO is based on a specific training pipeline (Tulu 3) and may not generalize to all preference optimization algorithms.
This paper has high impact on the LLM safety and evaluation community. It identifies a critical blind spot in current alignment practices: models are becoming better at "safety" but worse at "faithfulness" in a way that is invisible to standard benchmarks. This has immediate implications for RAG systems, clinical note extraction, and legal document processing, where silent modification of source content is a severe failure mode. The "trilemma" framework will likely influence future alignment research to explicitly optimize for faithfulness alongside safety and capability. The paper identifies and characterizes "alignment-induced unfaithfulness," a critical failure mode where aligned LLMs silently override source content on sensitive topics, revealing a fundamental capability-alignment-faithfulness trilemma. Through rigorous controlled experiments, a novel dataset, and causal analysis of post-training stages, the authors demonstrate that this unfaithfulness is a systematic, emergent property of modern alignment techniques (particularly DPO) that worsens with model scale, posing significant risks to the reliability of LLMs in source-reporting applications.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
[One sentence main contribution]. The paper presents a differentiable and GPU-accelerated acoustic simulator for the vocal tract that enables the reconstruction of vocal tract shapes from speech, with applications in linguistics and medical imaging. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant due to the novel frequency-domain formulation and the integration of turbulence modeling. The methodology is well-designed, leveraging neural networks to overcome the non-convexity of the inverse problem. The significance to the field is high, as it bridges the gap between physical simulation and deep learning, enabling new capabilities in speech analysis and medical imaging.
The paper introduces a differentiable acoustic simulator for the vocal tract, which is a significant technical achievement. The core innovation lies in the frequency-domain formulation of fluid dynamics, which is claimed to be 70x more GPU-parallelizable than time-domain finite differences. This is a crucial engineering insight for enabling real-time or near-real-time gradient-based optimization. The integration of a differentiable turbulence model for consonants is a strong addition, as consonants are notoriously difficult to model with simple tube acoustics. The use of neural networks to parameterize the geometry (similar to INRs) is a clever application of existing techniques to a new domain, helping to escape local minima in the non-convex inverse problem. The methodology is sound and well-motivated, combining physical simulation with modern deep learning techniques.
The experiments demonstrate the simulator's ability to reconstruct vocal tract shapes from speech across 11 languages, which is a robust test of generalizability. The self-supervised autoencoding task shows the model can learn meaningful representations without labeled data. The most impressive result is the coupling with a generative MRI model to reconstruct moving vocal tracts from speech alone, without paired data. This is a novel application that bridges speech processing and medical imaging. However, the evaluation of the MRI reconstruction quality is likely limited by the lack of ground truth for moving vocal tracts, and the paper should provide more quantitative metrics on the fidelity of the reconstructed shapes compared to actual MRI scans.
The paper provides a supplementary material link, which is a good sign for reproducibility. The frequency-domain formulation and turbulence model are described in detail, but the specific implementation details for the neural network parameterization and the MRI generative model coupling would be critical for reproduction. The claim of 70x speedup should be backed by detailed benchmarking against standard time-domain solvers.
The acoustic tube model is a simplification of the actual vocal tract, which is a 3D structure. The model may not capture all the nuances of speech production, especially for complex consonants or pathological speech. The MRI reconstruction is a novel application, but the clinical utility of the reconstructed shapes is not evaluated. The paper also does not discuss the computational cost of the MRI generative model coupling, which could be a bottleneck.
This work has significant potential impact in both linguistics and medical imaging. In linguistics, it could enable new insights into speech production and variation across languages. In medical imaging, it could lead to new methods for diagnosing speech disorders or planning surgical interventions. The differentiable simulator could also be used in other domains where inverse problems with physical constraints are important. [One sentence main contribution]. The paper presents a differentiable and GPU-accelerated acoustic simulator for the vocal tract that enables the reconstruction of vocal tract shapes from speech, with applications in linguistics and medical imaging. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant due to the novel frequency-domain formulation and the integration of turbulence modeling. The methodology is well-designed, leveraging neural networks to overcome the non-convexity of the inverse problem. The significance to the field is high, as it bridges the gap between physical simulation and deep learning, enabling new capabilities in speech analysis and medical imaging.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: Unknown (Likely ByteDance based on "YuE" naming convention and technical style, but not explicitly stated in provided text)
All Institutions: Unknown
YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
The paper proposes YuE2, a unified framework that integrates symbolic music generation (ABC notation) with audio generation (flow matching) within a single Mixture-of-Transformers (MoT) architecture. The core methodological contribution is the "symbolic planning" step, where the model first generates a readable score (melody, harmony, form) before expanding it into semantic tokens and acoustic latents. This is supported by two auxiliary models: MERT2, a music representation learner that sets new SOTA on MARBLE benchmarks, and SheetSage2, a full-song transcription model that generates the symbolic supervision signals required for training. The use of an AR-NAR MoT to handle both discrete symbolic/semantic tokens and continuous acoustic latents is a sophisticated architectural choice that allows for bidirectional attention in the acoustic stream while maintaining causal generation for the score.
The evaluation is extensive, covering automatic metrics (SongBench, SongEval, AudioBox) and expert listening tests. YuE2 outperforms public baselines and is competitive with proprietary systems like Suno v4.5/v5. The ablation study on symbolic planning is particularly strong, demonstrating that generating the score first significantly improves perceived musicality and overall quality compared to direct audio generation. The score-editing experiments show that the model can preserve unedited content while modifying specific sections, validating the utility of the symbolic interface.
The paper provides detailed architectural descriptions and training procedures. However, as a technical report from a likely industry lab, the full code and weights may not be immediately available to the public, which limits immediate reproducibility. The reliance on proprietary or large-scale datasets (346,000 hours of music) also poses a barrier for independent replication.
The model is large (3.58B parameters) and computationally expensive. The evaluation against proprietary systems is limited to expert listening and some automatic metrics, as direct API access for rigorous benchmarking is often restricted. The "best-of-8" selection strategy, while effective for benchmarking, may not reflect real-time interactive use cases where latency is critical.
This work bridges the gap between symbolic composition and audio production, offering a new paradigm for music generation that allows for explicit control over musical structure. The introduction of MERT2 and SheetSage2 as high-quality supervision tools will likely benefit the broader music AI community. The agentic editing capability suggests potential applications in professional music production workflows. YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Primary: UC Berkeley
All Institutions: UC Berkeley, Meta AI
World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
The paper proposes World Motion Models (WMMs), a unified framework for modeling 4D dynamics using sparse SE(3) pose trajectories. The core methodological contribution is the recasting of joint distribution modeling over multiple entities (humans, objects, cameras, robots) as a flexible sequence modeling problem using flow-matching. By employing per-token noise levels and a context token mechanism, the authors enable "any-to-any" marginal conditioning. This allows a single network to handle diverse tasks—such as prediction, infilling, and control—simply by applying different masks to the input sequence. The choice of SE(3) trajectories as a primitive is elegant, providing a minimal yet expressive representation that unifies articulated motion, rigid body dynamics, and camera motion into a shared latent space.
The paper reports experiments on six diverse applications spanning 3D vision and robotics, including future prediction, motion infilling, model-predictive control (MPC), inverse kinematics (IK), cross-embodiment retargeting, and policy learning. The breadth of tasks demonstrates the versatility of the unified approach. While the abstract claims "strong performance," the lack of specific quantitative metrics in the provided text prevents a full assessment of superiority over specialized baselines. However, the ability to switch tasks via masking without retraining is a significant empirical advantage over task-specific architectures.
The project page is provided, which likely contains code and datasets. The use of standard flow-matching techniques and SE(3) representations suggests that the core components are reproducible, provided the specific masking strategies and context token implementations are detailed in the full paper (which is not fully visible here, but implied by the "Spotlight" status and project page).
The reliance on SE(3) trajectories assumes that scene elements can be well-approximated by rigid motions. This may limit applicability to highly deformable objects or non-rigid interactions that cannot be decomposed into rigid parts. Additionally, the computational cost of flow-matching on long sequences of high-dimensional SE(3) poses could be a bottleneck for real-time applications, although the paper mentions MPC, suggesting some efficiency gains.
This work has the potential to significantly impact the field of embodied AI and 4D scene understanding. By unifying various robotics and vision tasks under a single generative framework, it simplifies the development of general-purpose agents. The "any-to-any" conditioning capability is particularly valuable for data-efficient learning and sim-to-real transfer, where conditioning on partial observations or specific actions is common. World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Primary: Unknown (Likely NVIDIA or similar top lab based on "SeedDance" and hardware, but affiliations are redacted in provided text)
All Institutions: Unknown
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a novel data augmentation strategy using counterfactual video generation to enable scalable humanoid loco-manipulation, demonstrating a robust real-to-sim-to-real pipeline that achieves zero-shot generalization to diverse objects, representing a significant step towards generalist robot learning from visual data.
The paper proposes PRISM, a framework that leverages video-to-video (V2V) generative models to create "counterfactual" human-object interaction videos from a small set of real seed videos. The core innovation lies in using these generated videos to train a real-to-sim-to-real pipeline. The methodology involves a contact-anchored reconstruction stage that uses human-object contact constraints to regularize monocular reconstruction errors, followed by a retargeting stage that uses contact anchors to guide the generation of physically plausible robot trajectories. The policy is trained via a privileged teacher-student distillation process, where the teacher tracks full-state references and the student learns from depth observations and joystick commands. The use of V2V generation to augment data with behavior-level variations (not just geometric) is a strong conceptual contribution, addressing the data scarcity bottleneck in visual imitation learning.
The experiments demonstrate zero-shot sim-to-real transfer on a Unitree G1 humanoid. The robot successfully picks up, carries, and drops diverse objects (boxes, barrels, bins, balls) and generalizes to unseen categories (chairs, tables, lamps). The success rates are high for in-domain objects (80-100%) and reasonable for out-of-domain objects (60-100%). The ablation study effectively shows that V2V generation outperforms simple geometric augmentation and that the contact-anchored reconstruction/retargeting is crucial for handling reconstruction noise. The comparison with OMOMO data highlights the superior generalization of the PRISM-generated data.
The paper provides detailed hyperparameters, reward structures, and pipeline runtime estimates. However, the reliance on specific proprietary video generation models (SeedDance 2.0) and specific reconstruction backends (CRISP, SAM3D) may limit immediate reproducibility for groups without access to these tools. The codebase is referenced but not explicitly linked in the text provided, though a project page is available.
The pipeline is computationally expensive (approx. 10 minutes per 8-second clip for reconstruction/retargeting). The method assumes rigid-body dynamics for objects, limiting applicability to deformable or articulated objects. The generalization to out-of-domain objects is good but not perfect, with some failures on complex shapes like chairs. The data scale is still relatively small (256 generated videos), and scaling behavior is not fully characterized.
This work bridges the gap between generative AI and robotics, offering a scalable path to collecting diverse interaction data without extensive real-world recording. It has significant implications for the development of generalist humanoid robots that can interact with a wide variety of household objects. The approach could be extended to other manipulation tasks and robot morphologies. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a novel data augmentation strategy using counterfactual video generation to enable scalable humanoid loco-manipulation, demonstrating a robust real-to-sim-to-real pipeline that achieves zero-shot generalization to diverse objects, representing a significant step towards generalist robot learning from visual data.
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
Primary: Harvard University
All Institutions: Harvard University, Massachusetts Institute of Technology, Princeton University, The Alan Turing Institute
The paper presents EMPIRIC, a novel framework for robot planning that learns residual world models by extending physics engines with executable code for missing mechanisms, achieving superior task success and sample efficiency through Bayesian inference and LLM-driven program synthesis.
The paper introduces EMPIRIC, a framework for learning "residual world models" by extending a base physics engine (PyBullet) with executable Python code for missing physical mechanisms (e.g., glue curing, wind forces). The core innovation is the separation of the base simulator (handling rigid-body dynamics) from the residual program (handling novel interactions and hidden state). The method employs Bayesian inference to estimate parameters and hidden states from noisy observations, using a mean-field approximation for the posterior. It integrates planning and information seeking by simulating candidate actions under multiple parameter draws to maximize success probability or mutual information. The use of a coding agent (LLM) to write and revise the residual code is a significant architectural choice, leveraging LLMs for program synthesis rather than just policy generation.
The evaluation is rigorous, covering five simulated domains (Domino, Bridge, Balloons, Boil, Fan) and one physical robot experiment. The protocol is a continual learning setup where the agent must solve training and test tasks within a step budget. EMPIRIC achieves 100% success rate in simulation, outperforming baselines like Direct Agent, Direct + Scene, and Standalone Sim. The ablation studies effectively isolate the contributions of parameter fitting and explicit uncertainty handling, showing that uncertainty is critical for irreversible actions (e.g., balloon bursting). The physical robot experiment demonstrates real-world applicability by learning wind forces and domino masses from two gusts to plan a cascade.
The paper provides extensive details in the appendix, including the agent's workspace, tools, simulator subclass interface, and belief construction algorithms. The code and website are provided. The use of a specific LLM (Claude Opus 5) is noted, which may limit immediate reproducibility if access to that specific model version is restricted, but the framework is sufficiently detailed for implementation.
The agent relies on predefined object features and does not learn feature extraction from raw images. The computational cost is high (median 107 minutes per run), which is a significant barrier to real-time deployment. The base simulator is given, so the method does not address scene reconstruction from scratch. The reliance on a powerful LLM for code generation introduces potential biases and costs associated with API usage.
This work bridges the gap between symbolic program synthesis and physical robotics, offering a path toward robots that can adapt to novel physical environments without retraining neural networks for every new object. The concept of "residual world models" is likely to influence future research in hybrid simulation and model-based reinforcement learning. The integration of Bayesian uncertainty with code-based models provides a robust framework for safe decision-making in uncertain physical environments. The paper presents EMPIRIC, a novel framework for robot planning that learns residual world models by extending physics engines with executable code for missing mechanisms, achieving superior task success and sample efficiency through Bayesian inference and LLM-driven program synthesis.
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Primary: Imperial College London
All Institutions: Imperial College London, Robotics and AI Institute
ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.
The paper proposes ZeroBot, a framework that integrates image-to-3D generative models (specifically InstantMesh) with massively parallel reinforcement learning (Isaac Gym) to enable rapid robot learning. The core methodological contribution is a novel "contact action space" that leverages the generated mesh geometry and a learned value function to sample high-value contact states, thereby accelerating exploration and allowing policies to be learned from scratch in minutes. The pipeline involves generating a complete object mesh from a single RGB-D view, aligning it for scale, using it for simulation and pose tracking (Foundation Pose), and training a PPO policy with a goal-flow reward. The approach is modular and effectively bridges the gap between generative AI and robotic control.
The evaluation is rigorous and well-designed, featuring six distinct real-world tasks (grasping, pushing, articulated interaction, multi-stage manipulation) on a Franka Research 3 arm. The paper provides strong ablations comparing the proposed method against baselines that use partial meshes (no 3D prior) and mesh retrieval (ACDC-NN). It also includes a comparison against ground-truth scanned meshes to quantify the performance gap of generative models. The results demonstrate an 87% success rate with average training times of ~2 minutes, which is a significant improvement over standard RL training times. The inclusion of challenging viewpoints (occluded handles, unseen sides) further validates the robustness of the generative prior.
The paper provides sufficient detail for reproducibility, specifying the hardware (Franka, RealSense cameras), software stack (Isaac Gym, PPO, InstantMesh, Foundation Pose, cuRobo), and hyperparameters (e.g., 256 parallel agents, temperature for softmax sampling). The use of standard, available tools and clear descriptions of the pipeline stages (mesh generation, alignment, RL training) makes it feasible for other groups to replicate the results, assuming access to similar computational resources (A6000 GPUs) and robotic hardware.
The method relies on several assumptions: a static background, a single rigid object (though extended to articulated with known parameters), and clear, unoccluded views for the initial mesh generation. It does not automatically infer physical properties like mass or friction, relying on constant values or manual specification. The performance is bounded by the accuracy of the current image-to-3D models, which can struggle with complex textures or highly non-convex shapes. Additionally, the method requires a goal pose to be specified, limiting its autonomy in open-ended tasks.
This work has significant potential to accelerate the deployment of robotic manipulation systems by reducing the data and time requirements for learning new tasks. By leveraging generative models to create simulation environments on-the-fly, it enables a "zero-shot" approach to real2sim2real transfer. This could lead to more adaptable robots in dynamic environments where pre-scanning or manual modeling is impractical. The framework also highlights the utility of value functions beyond policy improvement, using them for state sampling and deployment-time planning. ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.