High-fidelity simulation with aligned robot control β 15 perturbation factors, 7 manipulation skills, 4,000+ objects β that predicts how VLA policies behave in the real world.
REALM is a large-scale realistic simulation environment and benchmark for generalization in robotic manipulation. It supports 7 distinct manipulation skills and stress-tests them against 15 perturbations. Through empirical validation, we show that evaluation results in simulation are strongly correlated with real-world performance. Setup instructions, task definitions and the evaluation API are documented in the repository wiki.

Evaluating Ο0, Ο0-FAST and GR00T N1.5 under systematic perturbation.
Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to generalize beyond the specific environments and conditions they were trained on, which is presently difficult and expensive to evaluate in the real world. To address this gap, we present REALM, a new simulation environment and benchmark designed to evaluate the generalization capabilities of VLA models, with a specific emphasis on establishing a strong correlation between simulated and real-world performance through high-fidelity visuals and aligned robot control. Our environment offers a suite of 15 perturbation factors, 7 manipulation skills, and more than 4,000 objects. Finally, we establish two task sets that form our benchmark and evaluate the Ο0, Ο0-FAST, and GR00T N1.5 VLA models, showing that generalization and robustness remain an open challenge. More broadly, we also show that simulation gives us a valuable proxy for the real world and allows us to systematically probe for and quantify the weaknesses and failure modes of VLAs.
High-fidelity simulation with aligned robot control can serve as a valuable proxy for real-world performance, mitigating the issue of saturated simulation benchmarks.
Despite VLM backbones pretrained on Internet-scale data, there is a noticeable drop in performance from purely semantic perturbations.
There is still a noticeable sensitivity to camera view for all models, despite the unusually high diversity of viewpoints in the DROID dataset.
Behavioral generalization across objects and their properties is the most challenging for all tested models.
Conversely, all tested models seem to generalize well across known skills when the manipulated object remains the same.
Reliability and robustness under perturbations is still highly challenging β models exhibit very low success rates on many basic manipulation tasks.
While we recognize the tremendous progress that enabled VLAs to start performing manipulation tasks in unseen settings, including many scenes in our simulation, we believe these results indicate that current models still lack the capabilities for autonomous real-world deployment.
To assess the robustness of VLAs under variable conditions, we implemented perturbations that change the visual, behavioral, and semantic properties of the tasks and environments. We adopt 14 of the 22 perturbations from the β-Gen taxonomy and introduce a separate 15th V-LIGHT perturbation for scene illumination.
Hover a segment to see how the scene changes.
A detailed evaluation of Ο0, Ο0-FAST and GR00T N1.5 under 15 distinct perturbations on the 8 tasks that comprise the REALM-base task set.
Evaluation of three VLA models under 15 perturbations across the REALM-base and REALM-articulated task sets. Left: a visualization of the task. Right: task progression on the nominal default setting (black) and under individual perturbations (colored curves).
We are extending the evaluation beyond Ο0, Ο0-FAST and GR00T N1.5 to the current generation of open VLA policies, and will publish the full per-perturbation breakdown as a continuously updated leaderboard. Submissions of your own policy will be open alongside it.
Results are being collected now β rankings publish with the camera-ready release.
Watch the repo for updatesBibTeX
@article{sedlacek2025realm,
title={REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation},
author={Martin Sedlacek and Pavlo Yefanov and Georgy Ponimatkin and Jai Bardhan and
Simon Pilc and Mederic Fourmy and Evangelos Kazakos and Cees G. M. Snoek and
Josef Sivic and Vladimir Petrik},
journal={arXiv preprint arXiv:2512.19562},
year={2025}
}