RoboTwin-Phys

RoboTwin-Phys

Jiaqi Zhang1,2,3,4, Feng Ye1,†, Mingjia Yang1, Zhihong Chen1, Mingkang Xiang1, Xinglin Yao1, Yanbin Li1, Siwei Ma1, Chuanmin Jia1,‡
†Project Lead    ‡Corresponding Author
1Peking University    2Advanced Institute of Information Technology, Peking University    3Memo    4DEMX
Peking University and AIIT logo Memo logo DEMX logo

Do WAMs and VLAs Understand the Physical World?

RoboTwin-Phys: a physically grounded simulation benchmark — from visual diversity to physical diversity
Abstract
Manipulation benchmarks increasingly randomize what the world looks like — appearance, layout, lighting, viewpoints — but keep what the world is like fixed: mass, friction, restitution, joint damping, and geometry stay at nominal values. RoboTwin-Phys introduces physical-condition diversity as an explicit dimension of robot-manipulation evaluation on the 50-task bimanual suite of RoboTwin 2.0: 13 physical attributes are continuously sampled at the episode level within physically plausible, task-aware ranges, and every condition is verified executable by an expert planner before it enters the benchmark. Together with the benchmark we release a training set of more than 5,000 expert demonstrations, each annotated with the 13-dimensional ground-truth physical parameters in effect during its episode — enabling physical-attribute estimation, condition-aware modeling, and physics-conditioned policy training. Evaluations of representative WAMs and VLAs — Fast-WAM, Motus, FACT, π0.5, and GalaxeaVLA — reveal a substantial robustness gap: models that remain effective under existing visual and layout randomization degrade markedly when physical conditions vary. RoboTwin-Phys provides the benchmark, data, and evaluation protocol needed to systematically measure and improve robustness to physical-condition diversity.
Motivation

Why benchmark physical drift?

A manipulation policy deployed in the real world almost never meets the physics it was trained with. Tabletops get waxed or wet, so grasps that used to hold start to slip. A box may be loaded or empty; a bottle may be half-filled, shifting its center of mass mid-grasp. Hinges and drawers age and resist. Workbenches are not perfectly level; camera mounts get bumped during maintenance. None of these changes alter what the scene looks like — yet each of them can decide whether a grasp holds, a placement stays, or a door opens.

Today's benchmarks certify robustness to appearance, not to physics. RoboTwin 2.0's official suite varies backgrounds, lighting, clutter, object poses, and camera viewpoints — while the underlying physical parameters (mass, friction, restitution, joint damping, geometry) are kept at fixed nominal values. A policy can thus ace the official evaluation and still fail the first time the world's physics drifts.

Physics is not a constant. Drag the sliders below — same scene, same action, only one physical factor changes each time — and watch outcomes diverge. If physics matters this much, it deserves its own benchmark dimension: physical conditions that are continuously sampled, exactly reproducible, verified feasible, and labeled with ground-truth values.

Each card is one physical factor; each slider runs three levels from the default outward.

Benchmark design

Factor selection is grounded in deployment reality. Each of the 13 factors maps to a routine, well-documented deployment event — a waxed tabletop, a half-filled bottle, an aging hinge, a bumped camera mount — rather than an arbitrary parameter sweep. The factors group into five families: object properties (mass, center of mass, geometry scale), contact properties (friction, restitution), damping properties (object, arm, and gripper joint damping), task environment (table tilt, table height, external force), and camera configuration (camera distance, camera angle).

Continuous, episode-level sampling. The 13 attributes are sampled independently at episode initialization from their designated continuous ranges and remain fixed during the rollout; the episode seed determines the sampling stream, so any instantiated condition can be reproduced exactly. Physical sampling is independent of the original visual and layout randomization — physical variation can be studied in isolation, with visual variation retained, or with both enabled simultaneously.

Task-aware physical validity. A single global range is not appropriate for every task. Nine physics-sensitive tasks — including dump_bin_bigbin, grab_roller, beat_block_hammer, put_bottles_dustbin, open_microwave, turn_switch, put_object_cabinet, scan_object, and place_bread_skillet — are automatically routed to dedicated, empirically calibrated configurations, so every benchmark instance remains a physically plausible instance of its task rather than a numerically valid but meaningless simulator state.

Expert-verified feasibility. A sampled physical condition is retained as a benchmark instance only when the expert planner can complete the task under it. Failures therefore reflect a model's inability to cope with physical variation — not the absence of a valid solution under the sampled environment.

13 injectable physical factors

FactorPhysical meaningReal-world counterpartTypical failure it induces
🧊 Friction μContact friction coefficientWaxed / wet tabletopGrasp slip, slide-off
🏋️ MassObject mass scaling (0.125–4×)Loaded vs. empty boxDrop, unstable placement
⚖️ Center of massCoM offset ±1.5 cm per axisHalf-filled bottleIn-grasp rotation, topple
🎾 RestitutionContact bounciness (0–0.5)Rubber vs. ceramic partsBounce-out of container
🚪 Object joint dampingArticulated-joint resistance (0–25)Aging hingesDoor / lid won't open
🦾 Arm joint dampingRobot arm PD damping (0.5–3×)Aging actuatorsSloppy / overdamped motion
🤏 Gripper dampingGripper PD damping (0.5–3×)Worn gripper driveUnstable grasp force
📐 Table tiltNon-level workspace (0–2°)Unleveled tableRoll-away, slow topple
🌬️ External forcePersistent force up to 0.1 mgAirflow, conveyor vibrationDrift during transport
📦 Geometry scaleUniform object scaling (0.85–1.05×)Different product batchesGrasp-pose mismatch
🪑 Table heightWorkspace height (−8–0 cm)Non-standard furnitureOut-of-workspace reach
🎥 Camera distanceHead-camera displacement: orbit 0–10 cm or shift 0–3 cmBumped camera mountViewpoint generalization gap
📸 Camera angleHead-camera rotation 0–10°Tilted camera mountMis-estimated object poses

Failure cases current benchmarks never produce

⚖️ Off-center CoM — the rolling pin swings down mid-grasp.
🎾 High restitution — the ball bounces out of the bowl.
🚪 Aged joint damping — the drawer barely opens.

Dataset

Along with the benchmark we release RoboTwin-Phys Training Set v1: more than 5,000 expert episodes spanning the physical-condition distribution defined by the benchmark. Demonstrations are generated with the established expert pipeline — physics-sensitive tasks are routed to their dedicated configurations, visual and layout randomization remains available during collection, and every retained episode is verified executable under its instantiated physical condition. The data remain fully compatible with the official RoboTwin format, so existing WAM and VLA pipelines can consume them without any change to their trajectory interfaces.

ComponentContentFormat
TrajectoriesActions, states, and camera streams for every episodeHDF5 (official RoboTwin format)
VideosCorresponding MP4 recordingsMP4
phys_metaThe 13 physical attributes instantiated for the episodeJSON, per episode
ManifestMapping between tasks, benchmark configurations, and episode countsIncluded in the release

Evaluation Protocol

Three environment protocols. All tasks and all methods are evaluated under a single shared setup with three conditions: Clean — all benchmark randomization disabled, corresponding to the nominal environment configuration; Official Random — the original visual and layout randomization of RoboTwin 2.0 enabled, with physical parameters at their nominal values; Physical Random — the original randomization retained, and the 13 physical attributes additionally sampled according to the RoboTwin-Phys protocol. This keeps every reported number comparable across methods and isolates physical robustness from visual robustness.

Task-aware configurations. Nine physics-sensitive tasks are routed to dedicated, empirically calibrated physical domains — the exact task-specific configurations are released with the benchmark — while all other tasks share the global configuration. This ensures the benchmark measures robustness to physical diversity rather than artificially induced failure cases.

Feasibility filtering. The expert planner is run in advance on every instantiated condition; episodes it cannot complete are excluded from the evaluation pool. Reported failures therefore reflect the evaluated model's ability to cope with physical variation, not the absence of a valid solution under the sampled environment.

Reporting. Each task is evaluated with 100 rollout episodes. We report per-task and aggregate success rates under each protocol; Clean and Official Random numbers are the models' reported results from the corresponding public benchmark evaluations, while Physical Random is evaluated by us. The complete 50-task breakdown for Fast-WAM, Motus, and FACT is in the report (Appendix D).

Leaderboard

MethodCleanOfficial RandomPhysical Random
Fast-WAM91.8891.7844.24
Motus88.6687.0239.60
FACT88.4086.6039.14
π0.582.7476.7631.60
GalaxeaVLA (G0.5)93.7092.8037.83

Aggregate success rates (%) on 50 tasks; Physical Random uses 100 rollout episodes per task. Clean and Official Random are reported results from public benchmark evaluations; Physical Random is evaluated by us.

Citation

@article{zhang2026robotwinphys,
  title   = {RoboTwin-Phys: Do WAMs and VLAs Understand the Physical World?},
  author  = {Zhang, Jiaqi and Ye, Feng and Yang, Mingjia and Chen, Zhihong and Xiang, Mingkang and Yao, Xinglin and Li, Yanbin and Ma, Siwei and Jia, Chuanmin},
  journal = {arXiv preprint arXiv:2609.26292},
  year    = {2026}
}