Inverse Dynamics Model · Observation + future latents → ActionsPlaying
Swipe or scroll horizontally to explore the diagram ↔
Conditioning→Prediction→Loss
Condition: future latents · Action loss
01Inputs
02Action Expert
03Motion command
04Whole-body control
05Feedback
Training and inference use the full diagram at the same scale. Vision Expert and Action Expert titles stay visible in every mode. In inverse dynamics, the unused vision output tokens and arrow are hidden; in forward dynamics, the unused action output tokens and arrow are hidden. Lower conditioning inputs remain visible. All four training modes show current images and four future images below the Vision Encoder. Inverse Dynamics Model uses clean future images; Forward Dynamics Model, Visual Planning and Policy use the supplied noisy future images. Current and future images join at the encoder input. The action input is clean in forward dynamics and noisy in other modes, with proprioception shown separately. Visual Planning keeps the noisy values but grays out the action branch. Policy and inference gray out the future visual branch. Images fade in together while the prompt types in. Policy retains the shared training encoder layout with noisy future images and grays out this branch. Inference retains its original image and encoder layout without future images. Input tokens remain pale; predicted output tokens appear dark. Training ends with the relevant action loss or latent reconstruction loss. Execution is gray and static during training. In inference, human and G1 motion clips begin when their incoming information flows arrive. Human Intervention starts moving together with the outgoing Motion Command and intervention arrows; its reference arm stays gray and its moving arm is green. Pause stops all motion and information flow. Solid, dashed, and dotted lines indicate execution, optional human input, and feedback. Timing is illustrative. The multi-modal self-attention block, marked ×L, is stacked L layers deep.