Introduction

The humanoid form offers the possibility of working wherever people do. Humanoid intelligence is what turns that possibility into useful and reliable action.

Introducing Δ₀, Delta Intelligence’s humanoid foundation model (HFM) for whole-body loco-manipulation. Δ₀ learns from human motion and robot interaction to coordinate perception, locomotion, and manipulation through a learned whole-body controller.

Our homes, workplaces, tools, and appliances are organized around human reach, movement, and dexterity. A humanoid body is suited to these environments and offers a connection to human experience: how people move and handle objects across tasks and settings. Robot interaction grounds that experience in the robot’s own capabilities.

Whole-body loco-manipulation couples movement with object interaction while maintaining balance and managing contact. In long-horizon tasks, every action changes the conditions for what comes next. Completing the task means preserving balance, maintaining useful grasps, and arriving in positions that allow the robot to continue.

The household demonstration brings this pursuit into focus: making a bed, picking up objects from the ground, opening a dishwasher, and operating a step trash can each require coordinated movement and manipulation, connected across a sequence of everyday work.

Loco-manipulation: the most challenging unsolved problem

At Delta Intelligence, we view general, reliable whole-body loco-manipulation as the core problem of humanoid intelligence—and the most challenging unsolved problem on the path to useful general-purpose robots. It brings locomotion, whole-body control, and manipulation together: moving through an environment, maintaining balance under contact, and acting on objects. Each capability continuously constrains the others.

The robot must control a moving body while perceiving and manipulating a changing world. A humanoid has a floating base: its position and orientation change as it moves, and its balance depends on environmental contacts. Stepping and posture adjustments affect both hand position and, with body-mounted cameras, visual perspective. Dexterous manipulation makes this coupling especially demanding. Small body movements can disturb hand–object alignment, so precise interaction requires coordinated control of the hands, torso, and legs.

Long-horizon tasks compound these difficulties. Carrying a cup to a dishwasher requires a grasp that survives transport, walking that keeps the cup stable, and an arrival stance that supports placement. Each action determines the starting conditions for the next. Small errors can accumulate, and a successful intermediate action may leave the robot poorly positioned to continue. Useful execution must preserve balance and control across these transitions while maintaining a practical pace.

The same task must succeed across changing viewpoints, body configurations, support contacts, and object poses. Learning requires experience that covers these variations and a controller capable of executing the resulting actions. Human activity provides diverse coordination patterns; robot demonstrations ground them in the robot’s capabilities. Corrections and physical interaction help address deviations, while repeated sequence evaluation—through simulation and hardware—tests whether improvements hold across the complete task.

Δ₀: six steps toward humanoid intelligence

Long-horizon loco-manipulation demands coordination, practical speed, transferable experience, executable control, corrective learning, and scalable evaluation together. Δ₀ brings these requirements into one learning system.

What it does

  • Unlocks a humanoid’s full dexterity across 69 DoFs, advancing toward human-level loco-manipulation and coordination. Δ₀ brings the hands, arms, torso, and legs into a shared action space, using posture, stepping, and body weight to support manipulation. This enables coordinated actions such as reaching while repositioning and balancing while handling objects. The turntable task illustrates this coordination through smooth, dexterous interaction.
  • A single generalist policy performs diverse everyday tasks at near-human speed. Human-level speed is essential for robots to be useful in everyday life: they must complete tasks in practical timeframes and keep pace with the people around them. Δ₀ connects perception, navigation, and physical interaction in one continuous loop, using cameras, language, and the robot’s own state to coordinate movement and manipulation.

How it’s built

  • The brain and controller are co-designed to scale humanoid intelligence. The brain translates visual and language inputs into whole-body goals; the controller turns them into coordinated motion and motor outputs. Trained on a vast, diverse motion dataset, Δ₀’s controller outperforms all evaluated state-of-the-art controllers, combining a rich motion repertoire, accurate tracking, and compliant interaction. The same controller supports teleoperation and autonomous execution, connecting what the robot can demonstrate with what its policy can learn and perform. Their shared representation enables both components to evolve together, with scaling measured through motion-tracking accuracy and task success.

How it learns and improves

  • Pretrained on large-scale diverse full-body human data. Learning directly from people is a fundamental advantage of humanoid robots. Human activity offers a rich source of examples of how to move and interact with the world. Pairing egocentric observations with full-body motion connects what a person sees with how their body responds. Δ₀ learns from these paired data at scale, building a foundation for coordinated behavior and transfer to downstream tasks.
  • Real-to-sim-to-real accelerates HFM evaluation and enables humanoid agents. Δ₀ incorporates a real-to-sim-to-real loop into policy learning and evaluation. Reconstructed environments let humanoid agents navigate, manipulate, and retry autonomously, supporting repeatable testing and identifying failures to guide policy improvement. Testing the same tasks on hardware closes the loop, checking how well simulation results transfer to the real world.
  • Delta-action control enables human-in-the-loop learning and real-world reinforcement learning. Δ₀’s controller is designed around delta-action inputs, providing a shared interface for policy actions and human corrections during whole-body loco-manipulation. Through corrective feedback and physical interaction, the policy can improve beyond its initial demonstrations toward higher success rates and more robust execution.

Humanoid intelligence in everyday tasks

Everyday life involves a wide range of physical tasks: putting on a record, making the bed, picking up a toy, opening the dishwasher, using a step trash can, and sitting down on a sofa. These familiar activities require careful handling, coordinated movement, and balance as the body interacts with objects and furniture. Across the six tasks, Δ₀ adjusts its reach, posture, and support to meet these different demands, bringing whole-body coordination into practical household activity.

Operating a turntable

Operating a turntable: dexterous handling and precise placement. Δ₀ carefully grasps and aligns the record for placement on the turntable, coordinating its hands, arms, and torso to keep it steady.

Making the bed

Making the bed: deformable-object manipulation while moving. Δ₀ lifts and spreads the comforter while stepping around the bed, maintaining its grip as the fabric folds and shifts.

Picking up objects from the ground

Picking up objects from the ground: stable kneeling and standing. Δ₀ kneels to grasp the toy, then pushes up with one foot to stand, maintaining balance and holding the toy through the change in height.

Opening the dishwasher

Opening the dishwasher: articulated-object manipulation with whole-body force. Δ₀ braces its left hand against the shelf for support while its right hand pulls the heavy door open. It adjusts its arm motion and stance as the door rotates along its hinge.

Operating a step trash can

Operating a step trash can: precise interaction with the foot. Δ₀ shifts its weight onto one leg and aligns the other foot with the pedal, then presses down with controlled force to open the lid.

Sitting on a sofa

Sitting on a sofa: controlled lowering and weight transfer. Δ₀ positions itself in front of the sofa and lowers onto the seat, coordinating its hips, knees, and torso as support shifts from its feet to the cushions.

Two halves of one mind: co-designing the brain and controller

Δ₀ pairs a brain—a latent world-action model—with a learned whole-body controller. The brain predicts future latent world states and produces motion commands; the controller turns those motion commands into coordinated physical action.

We co-design the brain and controller so that the brain can express the motions a task requires and the controller can reliably execute them.

Observations, language, and robot state enter the brain; it predicts future latent world states and produces motion commands; the controller converts the motion commands into 69-DoF joint targets. Teleoperation enters at the motion-command interface; delta-action corrections enter at the controller.

The learned whole-body controller: a foundation for humanoid intelligence

The human cerebellum is estimated to contain about 80% of the brain’s neurons (Azevedo et al. (2009)). For humanoid robots, the analogy motivates a design principle: The whole-body controller should be learned, continually improved, and scaled alongside the brain.

Δ₀’s whole-body controller learns from diverse motion data to coordinate 69 degrees of freedom. It turns the brain’s motion commands into joint-level movement while maintaining balance and responding to contact.

Why whole-body control matters for humanoid intelligence

  • Infrastructure for brain learning. The brain learns to act through the controller. Controller quality therefore shapes the demonstrations we collect, the actions the brain can learn, and their execution on the robot.
  • Making the body’s capabilities available to the policy. A humanoid can change posture, step, shift support, lower its center of mass, kneel, and coordinate its whole body. The controller makes these possibilities reliably accessible to the task model, expanding the tasks and conditions that can be demonstrated and executed.

Design principles for the whole-body controller

  • An interface for intent and intervention. The brain expresses task intent as motion commands, and the controller turns them into coordinated whole-body movement. Because the interface also accepts delta-action corrections, a human can adjust the intended motion during execution without taking over low-level control or switching to a separate control stack.
  • Whole-body capability. The controller must track motion through changes in posture and support, including transitions between standing, stepping, kneeling, and sitting, while maintaining balance and managing contact. These capabilities support the everyday tasks above.
  • Continuous reliability. A single loss of footing or overheating event can interrupt a deployment run and waste an entire data-collection episode. Our controller is designed to sustain stable operation within hardware and thermal limits. In our data-collection setup, it experiences fewer control-related interruptions than the other controllers we evaluated, enabling more efficient data collection.
  • Execution fidelity and responsiveness. Low tracking error and low latency let a teleoperator focus on the task rather than compensate for the robot, so collected trajectories more closely preserve the operator’s intent. The tracking comparison below measures the fidelity side of this principle: global root motion and local whole-body position accuracy across everyday-task motions.
  • Compliance. The controller combines precise motion tracking with the ability to yield to contact, accommodating external forces while maintaining balance and whole-body coordination. When opening a dishwasher, compliance lets the robot follow the door’s constrained path and adapt to changing resistance while maintaining its grasp.

Execution fidelity on everyday-task motions

Zero-shot for all controllers

Global Root Error 0.000.050.100.150.20 Trajectory error (m) Δ₀controllerHEFTMimicLitev1.1SONICv1.1ScaleBFMXL Body Position Error (MPJPE) 0102030 Local position error (mm) Δ₀controllerHEFTMimicLitev1.1SONICv1.1ScaleBFMXL
Evaluation protocol and metric definitions

We evaluate all controllers zero-shot on a motion dataset drawn from our everyday tasks, totaling approximately 3 h 31 min. The evaluation follows the workflow and metric definitions of the Motion Tracking Leaderboard, using its accompanying codebase to run the baselines: HEFT, MimicLite v1.1, SONIC v1.1, and ScaleBFM XL.

Global Root Error (m). Mean per-frame 3D pelvis displacement error relative to the motion start, averaged within each motion and then equally across motions.

Body Position Error / MPJPE (mm). Mean 3D position error over the selected body/link origins after aligning robot and reference pelvis horizontal position and yaw at each frame, averaged equally across motions.

The brain: a latent world-action model

Jointly learning future visual states and motion commands connects the robot’s actions with how the scene is expected to change. As shown in the architecture figure above, the mixture-of-transformers (MoT) architecture gives visual dynamics and whole-body motion dedicated modeling capacity while sharing information from visual observations, language, and the robot’s own state.

The architecture comprises three specialized branches. The vision-language branch processes current multi-view observations and the task instruction. The vision branch predicts future visual states in a semantic feature space provided by DINO, while the action branch models whole-body motion conditioned on proprioception. Each branch retains its own attention projections and feed-forward layers, while shared multimodal self-attention allows information to flow across the three streams.

We train the model with four complementary modes by varying the conditioning inputs and denoising targets. Forward dynamics predicts future visual features given actions, while inverse dynamics predicts actions given future visual features. Visual planning predicts future visual features without action conditioning, while policy-only training predicts actions without conditioning on future visual targets. Current observations and task context provide the common grounding across these modes. Together, these objectives train a single model to capture the relationships between observations, visual futures, and actions.

Long-horizon execution through stage-by-stage instructions. Δ₀ carries out extended tasks through stage-level instructions supplied by a higher-level model, agent, or a human. Connecting tasks is itself a whole-body challenge: completing one task does not necessarily leave the robot ready for the next. The instructions specify the next objective; Δ₀ must learn how to reposition its body, adjust its stance, and preserve or release a grasp so it can reach the next object and interact with it while staying balanced. Our focus is on these physical transitions, rather than a dedicated high-level task planner.

Scaling reliable humanoid intelligence in the real world

Scaling humanoid intelligence requires expanding what robots can learn, how thoroughly their capabilities can be evaluated, and how they improve through interaction. Human activity provides a rich source of coordinated behavior. Reconstructed environments make repeated evaluation across tasks and conditions more practical. Real-world experience provides feedback for refining skills and learning from failures. These three pathways—data, evaluation, and experience—connect broader learning with increasingly reliable physical execution. Δ₀ brings them together through human-data pretraining, real-to-sim-to-real evaluation, and real-world reinforcement learning.

Scaling data: learning from human data

Human activity offers a rich source of whole-body skills: people coordinate their hands, posture, and movement to interact with the world. The similarity between human and humanoid body structure makes this experience a useful starting point for learning both how to move and how to act in a scene.

We train the controller on large-scale high-quality human motion data, expanding its repertoire of coordinated, executable movements. We measure this foundation through tracking accuracy as the motion dataset grows: how faithfully the controller executes a requested movement. This measures execution quality, while the brain’s success is measured by whether its actions accomplish a task.

The brain pretraining utilizes over 10,000 hours of paired human data. Paired egocentric observations and whole-body motion connect what a person sees with the actions they take. These examples capture how people establish support and reach—stepping closer, lowering the torso, or changing posture—as they handle objects. Pretraining on this experience exposes the policy to varied scenes, intermediate postures, and object configurations.

We propose a shared action representation for Δ₀ that enables the brain to learn from different data sources and embodiments. Its 154 dimensions cover arm and hand actions, root commands, and motion commands, padded to a common 180-dimensional layout. Each data type contributes the channels it provides.

Many embodiments. One action space.

180D

Select a data type to highlight its action channels.

Left 54D
Arm joints7D
EEF pose6D
Gripper1D
Fingertips15D
Hand joints25D
Right 54D
Arm joints7D
EEF pose6D
Gripper1D
Fingertips15D
Hand joints25D
Body 46D
Root command4D
Motion command42D
Wholebody motionHands, root and motion commands
126Dactive channels
Left40/54
Right40/54
Body46/46
Padding26

Image credits

Illustrative imagery: DROID (single arm), ALOHA (dual arm), EgoDex, Hoque et al. (egocentric hands), UMI, Humanoid Everyday (humanoid robot), and BONES-SEED, Bones Studio (motion). Wholebody views are from our human-motion recordings. Images illustrate data types; the channel mapping shows our shared representation.

The examples below showcase our own collected egocentric data, pairing video with continuous whole-body motion and detailed hand reconstruction. The synchronized views connect the person’s perspective with body movement and finger motion during everyday activities.

We evaluate controller scaling by measuring high-precision tracking coverage on everyday-task motions at increasing motion-data budgets. As motion data scales, coverage rises steadily, reflecting clear gains in controller capability.

High-precision Tracking Coverage vs. Training Data In zero-shot evaluation on a subset of everyday-task motions, 29.2%, 48.0%, and 72.6% meet both tracking thresholds at motion training data settings labeled 10%, 40%, and 100%, respectively. This evaluation emphasizes end-effector positional accuracy, measured at the wrists, while also constraining global root error. Each motion must have mean wrist position error at most 2 cm and mean global root position error at most 15 cm. Data budgets are shown as discrete settings. High-precision Tracking Coveragevs. Training Data 020406080100 10%40%100%Motion Training Data UsageHigh-precision tracking coverage (%)
Coverage: share of everyday-task motions tracked zero-shot with mean wrist error ≤2 cm and mean root error ≤15 cm.
Policy Success vs. Training Data 020406080100 0%10%100%Pre-training Data UsageSuccess rate (%) Loco-manipulation tasks Bimanual dexterous tasks Whole-body dexterous tasks
Observed success rates. Pretraining data usage is shown at three discrete settings: 0%, 10%, and 100%.

We evaluate policy scaling by measuring task success at increasing pretraining data budgets. In our internal evaluations, task success improves with more pretraining data across three task categories: basic loco-manipulation that a mobile manipulator could also perform; bimanual tasks requiring fine manipulation; and whole-body tasks involving contact beyond the hands, such as closing a cabinet with a leg.

We also examine how the policy responds to unfamiliar conditions and disruptions through the following demonstrations. While loading a cup into the dishwasher, Δ₀ operates under shifting colored light it never saw in training. The paired views show the scene and the robot’s onboard observations.

Lighting variation — Third-person view

Lighting variation — Onboard camera views

The recovery clips show a missed grasp, a new cup placed on the counter after the first is loaded, and a person pushing the trash can lid shut. In these demonstrations, Δ₀ responds autonomously and completes the task without teleoperation or a manual reset.

Autonomous recovery — Dishwasher

Autonomous recovery — Step trash can

These demonstrations show adaptation within individual tasks. Establishing how consistently a policy handles such variations requires systematic evaluation across repeated trials and conditions.

Scaling evaluation: closing the real-to-sim-to-real loop

Scaling evaluation is a bottleneck to accelerating humanoid model development. As training produces more candidate models, progress depends on quickly determining which changes improve performance and where failures remain. Testing every checkpoint on a physical humanoid is slow, costly, and limited in coverage, delaying the feedback needed to guide the next training cycle. Real-to-sim-to-real evaluation addresses this bottleneck by reconstructing real tasks in simulation, comparing policies through repeatable trials, and returning selected candidates to hardware for validation. This shortens the feedback cycle between training, evaluation, and model refinement.

We reconstruct deployment environments and task-relevant objects as calibrated simulation assets. Our pipeline combines 3D reconstruction methods with agentic frameworks built on cutting-edge large language models (LLMs) such as GPT-6 Astra, which automatically turn video captures of a scene into simulation-ready assets with high visual, geometric, and physical fidelity, covering object geometry, materials, articulated components and hinges, contact surfaces, friction, the robot and its sensors, and the support beneath it. Editable properties allow controlled changes to lighting, appearance, friction, and object configuration while retaining the structure of the real task.

Deployment environment (left) and its reconstructed simulation (right), used for policy training and evaluation.

Within these environments, Δ₀ runs autonomous rollouts that connect navigation, manipulation, and retries across task stages. Repeated evaluation on the same task suite and controlled variations supports checkpoint comparison and reveals failures between stages that isolated trials may miss.

Dishwasher task evaluation in simulation: Δ₀ runs rollouts to assess policy behaviour before real-world validation.

On the dishwasher task, simulated and real-world success rates follow the same overall improvement trend as training data increases. This provides an early signal for comparing training changes before hardware validation.

Dishwasher · Real & Sim Evaluation 020406080100 20%40%60%80%100% Training Data Usage Success rate (%) Real-world success rate Simulation success rate

Dishwasher policy success in simulation and on hardware at increasing training data budgets.

Selected policies return to the real robot to test performance under physical contact, sensing, actuation, and balance demands. Hardware trials also check whether the relative ordering of candidates in simulation holds in reality. Discrepancies inform the next iteration of reconstruction and evaluation, closing the loop between repeatable simulated tests and physical execution.

Scaling experience: humanoid real-world RL post-training

Real-world reinforcement learning extends policy training beyond demonstrations by learning from autonomous attempts and targeted human guidance. Δ₀’s delta-action controller gives policy outputs and human-in-the-loop (HIL) corrections a shared execution interface.

A value model estimates task progress from visual observations, recent history, and robot state. During RL post-training, action chunks with larger gains in predicted progress receive a positive conditioning signal; those with smaller gains or setbacks receive a negative signal.

The predicted value represents the estimated cost-to-go until success. Values closer to zero indicate progress; a failed grasp moves the value away from zero.

Human corrections supply recovery examples where the autonomous policy struggles. HIL correction chunks receive a positive conditioning signal directly, allowing targeted guidance to contribute alongside the model’s assessment of autonomous attempts.

Human intervention corrects an erroneous grasp by an early-stage policy.

Dishwasher · OOD generalization

Success rate (%) 100500 20% 65% SFTAfter rollouts
SFT stands for supervised fine-tuning. SFT: 4/20 · After rollouts: 13/20.

On the dishwasher task, we further explore whether HIL corrections can help in out-of-distribution (OOD) settings. We evaluate in a mirrored layout: the kitchen is rearranged so the dishwasher and sink swap sides relative to the original, a spatial distribution shift that can expose the policy to OOD states. The policy trained on teleoperation data from the original scene initially achieved 20% success (4/20) in this layout. After collecting rollouts with targeted HIL corrections there and applying RL post-training, success reached 65% (13/20) under the same evaluation conditions. Together with the pretraining scaling results above, this suggests that pretraining followed by real-world RL post-training offers an effective path toward reliable humanoid loco-manipulation in everyday tasks.

What comes in the future

Scaling humanoid intelligence. A central question is how larger models and more varied training data translate into broader humanoid capabilities. Human activity, robot demonstrations, and interaction with the world offer complementary sources of experience, with the potential to support richer coordination and more adaptable behavior.

Toward reliable deployment. These capabilities become useful when robots can perform them dependably in everyday settings. Long-horizon tasks and changing environments bring questions of consistency, recovery, and practical execution speed into focus. Closing the distance between demonstrated skill and dependable everyday behavior remains central to the broader pursuit of humanoid intelligence.

Citation

Please cite this work as:

Delta Intelligence Team, “Delta-0 (Δ₀): A New Chapter in Humanoid Intelligence”, Delta Intelligence Blog, September 2026.

Or use the BibTeX citation:

@article{delta0,
  author = {{Delta Intelligence Team}},
  title = {Delta-0 (Δ₀): A New Chapter in Humanoid Intelligence},
  journal = {Delta Intelligence Blog},
  year = {2026},
  url = {https://deltai.com/en/blog/delta-0},
}

We’re hiring people with a passion for humanoid intelligence. Join us to build robots that can move, learn, and do useful work in the real world.

Follow @delta_intelli on X for future releases and updates.

Open roles Careers: delta-hr@deltai.com