Inverse Dynamics Model · Observation + future latents → ActionsPlaying
Brain Latent world-action model VLM ... ... Vision Expert ... ... Action Expert ... ... Multi-Modal Self-Attention ×LMulti-modal self-attention: L layers Image Prompt Open the washing-machine door. VisionencoderVisionencoder LatentLatent ...... Proprioception-0.10.10.40.2-0.20.0-0.4... Noisy action -0.2 0.1 0.3 0.3 -0.1 0.0 ... -0.2 Motion Command
Controller Whole-body control Robot Unitree G1
Door opening Teleoperation Δp Human Intervention Visual and proprioceptive feedback Execution Optional human input Feedback (from hardware) + Clean future images

Swipe or scroll horizontally to explore the diagram

ConditioningPredictionLoss
Condition: future latents · Action loss
Training and inference use the full diagram at the same scale. Vision Expert and Action Expert titles stay visible in every mode. In inverse dynamics, the unused vision output tokens and arrow are hidden; in forward dynamics, the unused action output tokens and arrow are hidden. Lower conditioning inputs remain visible. All four training modes show current images and four future images below the Vision Encoder. Inverse Dynamics Model uses clean future images; Forward Dynamics Model, Visual Planning and Policy use the supplied noisy future images. Current and future images join at the encoder input. The action input is clean in forward dynamics and noisy in other modes, with proprioception shown separately. Visual Planning keeps the noisy values but grays out the action branch. Policy and inference gray out the future visual branch. Images fade in together while the prompt types in. Policy retains the shared training encoder layout with noisy future images and grays out this branch. Inference retains its original image and encoder layout without future images. Input tokens remain pale; predicted output tokens appear dark. Training ends with the relevant action loss or latent reconstruction loss. Execution is gray and static during training. In inference, human and G1 motion clips begin when their incoming information flows arrive. Human Intervention starts moving together with the outgoing Motion Command and intervention arrows; its reference arm stays gray and its moving arm is green. Pause stops all motion and information flow. Solid, dashed, and dotted lines indicate execution, optional human input, and feedback. Timing is illustrative. The multi-modal self-attention block, marked ×L, is stacked L layers deep.