Why benchmark physical drift?
A manipulation policy deployed in the real world almost never meets the physics it was trained with. Tabletops get waxed or wet, so grasps that used to hold start to slip. A box may be loaded or empty; a bottle may be half-filled, shifting its center of mass mid-grasp. Hinges and drawers age and resist. Workbenches are not perfectly level; camera mounts get bumped during maintenance. None of these changes alter what the scene looks like — yet each of them can decide whether a grasp holds, a placement stays, or a door opens.
Today's benchmarks certify robustness to appearance, not to physics. RoboTwin 2.0's official suite varies backgrounds, lighting, clutter, object poses, and camera viewpoints — while the underlying physical parameters (mass, friction, restitution, joint damping, geometry) are kept at fixed nominal values. A policy can thus ace the official evaluation and still fail the first time the world's physics drifts.
Physics is not a constant. Drag the sliders below — same scene, same action, only one physical factor changes each time — and watch outcomes diverge. If physics matters this much, it deserves its own benchmark dimension: physical conditions that are continuously sampled, exactly reproducible, verified feasible, and labeled with ground-truth values.
Each card is one physical factor; each slider runs three levels from the default outward.
Factor selection is grounded in deployment reality. Each of the 13 factors maps to a routine, well-documented deployment event — a waxed tabletop, a half-filled bottle, an aging hinge, a bumped camera mount — rather than an arbitrary parameter sweep. The factors group into five families: object properties (mass, center of mass, geometry scale), contact properties (friction, restitution), damping properties (object, arm, and gripper joint damping), task environment (table tilt, table height, external force), and camera configuration (camera distance, camera angle).
Continuous, episode-level sampling. The 13 attributes are sampled independently at episode initialization from their designated continuous ranges and remain fixed during the rollout; the episode seed determines the sampling stream, so any instantiated condition can be reproduced exactly. Physical sampling is independent of the original visual and layout randomization — physical variation can be studied in isolation, with visual variation retained, or with both enabled simultaneously.
Task-aware physical validity. A single global range is not appropriate for every task. Nine physics-sensitive tasks — including dump_bin_bigbin, grab_roller, beat_block_hammer, put_bottles_dustbin, open_microwave, turn_switch, put_object_cabinet, scan_object, and place_bread_skillet — are automatically routed to dedicated, empirically calibrated configurations, so every benchmark instance remains a physically plausible instance of its task rather than a numerically valid but meaningless simulator state.
Expert-verified feasibility. A sampled physical condition is retained as a benchmark instance only when the expert planner can complete the task under it. Failures therefore reflect a model's inability to cope with physical variation — not the absence of a valid solution under the sampled environment.
13 injectable physical factors
| Factor | Physical meaning | Real-world counterpart | Typical failure it induces |
|---|---|---|---|
| 🧊 Friction μ | Contact friction coefficient | Waxed / wet tabletop | Grasp slip, slide-off |
| 🏋️ Mass | Object mass scaling (0.125–4×) | Loaded vs. empty box | Drop, unstable placement |
| ⚖️ Center of mass | CoM offset ±1.5 cm per axis | Half-filled bottle | In-grasp rotation, topple |
| 🎾 Restitution | Contact bounciness (0–0.5) | Rubber vs. ceramic parts | Bounce-out of container |
| 🚪 Object joint damping | Articulated-joint resistance (0–25) | Aging hinges | Door / lid won't open |
| 🦾 Arm joint damping | Robot arm PD damping (0.5–3×) | Aging actuators | Sloppy / overdamped motion |
| 🤏 Gripper damping | Gripper PD damping (0.5–3×) | Worn gripper drive | Unstable grasp force |
| 📐 Table tilt | Non-level workspace (0–2°) | Unleveled table | Roll-away, slow topple |
| 🌬️ External force | Persistent force up to 0.1 mg | Airflow, conveyor vibration | Drift during transport |
| 📦 Geometry scale | Uniform object scaling (0.85–1.05×) | Different product batches | Grasp-pose mismatch |
| 🪑 Table height | Workspace height (−8–0 cm) | Non-standard furniture | Out-of-workspace reach |
| 🎥 Camera distance | Head-camera displacement: orbit 0–10 cm or shift 0–3 cm | Bumped camera mount | Viewpoint generalization gap |
| 📸 Camera angle | Head-camera rotation 0–10° | Tilted camera mount | Mis-estimated object poses |
Failure cases current benchmarks never produce
Dataset
Along with the benchmark we release RoboTwin-Phys Training Set v1: more than 5,000 expert episodes spanning the physical-condition distribution defined by the benchmark. Demonstrations are generated with the established expert pipeline — physics-sensitive tasks are routed to their dedicated configurations, visual and layout randomization remains available during collection, and every retained episode is verified executable under its instantiated physical condition. The data remain fully compatible with the official RoboTwin format, so existing WAM and VLA pipelines can consume them without any change to their trajectory interfaces.
| Component | Content | Format |
|---|---|---|
| Trajectories | Actions, states, and camera streams for every episode | HDF5 (official RoboTwin format) |
| Videos | Corresponding MP4 recordings | MP4 |
| phys_meta | The 13 physical attributes instantiated for the episode | JSON, per episode |
| Manifest | Mapping between tasks, benchmark configurations, and episode counts | Included in the release |
Evaluation Protocol
Three environment protocols. All tasks and all methods are evaluated under a single shared setup with three conditions: Clean — all benchmark randomization disabled, corresponding to the nominal environment configuration; Official Random — the original visual and layout randomization of RoboTwin 2.0 enabled, with physical parameters at their nominal values; Physical Random — the original randomization retained, and the 13 physical attributes additionally sampled according to the RoboTwin-Phys protocol. This keeps every reported number comparable across methods and isolates physical robustness from visual robustness.
Task-aware configurations. Nine physics-sensitive tasks are routed to dedicated, empirically calibrated physical domains — the exact task-specific configurations are released with the benchmark — while all other tasks share the global configuration. This ensures the benchmark measures robustness to physical diversity rather than artificially induced failure cases.
Feasibility filtering. The expert planner is run in advance on every instantiated condition; episodes it cannot complete are excluded from the evaluation pool. Reported failures therefore reflect the evaluated model's ability to cope with physical variation, not the absence of a valid solution under the sampled environment.
Reporting. Each task is evaluated with 100 rollout episodes. We report per-task and aggregate success rates under each protocol; Clean and Official Random numbers are the models' reported results from the corresponding public benchmark evaluations, while Physical Random is evaluated by us. The complete 50-task breakdown for Fast-WAM, Motus, and FACT is in the report (Appendix D).
Leaderboard
| Method | Clean | Official Random | Physical Random |
|---|---|---|---|
| Fast-WAM | 91.88 | 91.78 | 44.24 |
| Motus | 88.66 | 87.02 | 39.60 |
| FACT | 88.40 | 86.60 | 39.14 |
| π0.5 | 82.74 | 76.76 | 31.60 |
| GalaxeaVLA (G0.5) | 93.70 | 92.80 | 37.83 |
Aggregate success rates (%) on 50 tasks; Physical Random uses 100 rollout episodes per task. Clean and Official Random are reported results from public benchmark evaluations; Physical Random is evaluated by us.
Citation
@article{zhang2026robotwinphys,
title = {RoboTwin-Phys: Do WAMs and VLAs Understand the Physical World?},
author = {Zhang, Jiaqi and Ye, Feng and Yang, Mingjia and Chen, Zhihong and Xiang, Mingkang and Yao, Xinglin and Li, Yanbin and Ma, Siwei and Jia, Chuanmin},
journal = {arXiv preprint arXiv:2609.26292},
year = {2026}
}