Back to Open Positions
Experienced Roles

Robot Post-Training Engineer (VLA/WAM & Real-world RL)

Beijing
Full-time
Real robot deployment experience
Relevant background

Job Description

  • As the core driver of post-training and real-device deployment, you will promote embodied foundational models from 'digital brains' to 'physical entities,' directly participating in the training, deployment, and closed-loop optimization of models on real robots.

Job Responsibilities

  • Post-training for VLA and WAM: Responsible for the SFT, RLHF, and DPO processes of Vision-Language-Action (VLA) and World/Action Model (WAM); designed efficient fine-tuning paradigms for high-dimensional continuous action spaces to reduce action hallucinations and temporal inconsistencies under multimodal inputs.
  • Real-world RL: Designing and deploying high sample-efficiency online reinforcement learning algorithms on real humanoid robots or complex hardware systems, focusing on addressing issues of safe exploration, environmental dynamic adaptation, and reward modeling under real physical feedback.
  • On-Robot Deployment and Closed-Loop Optimization: Lead high-performance deployment of policy models to edge-computing platforms such as NVIDIA Jetson, optimize them with TensorRT and ONNX, connect high-level vision-language planning with low-level WBC/MPC control, and systematically address latency, frequency mismatch, and high-frequency jitter.
  • Data-driven behavior alignment: Utilizing Human-in-the-loop to take over data and edge cases to build an online correction flywheel, continuously improving the model's generalization ability and physical common sense in open scenarios starting from failure modes.

Job Requirements

  • Background in computer science, robotics, or related fields, with a solid understanding of deep reinforcement learning algorithms such as PPO, SAC, TD3, as well as generative action architectures like Diffusion Policy and Transformer-based Control.
  • Familiar with StarVLA, π0.5, GR00T, Cosmos, or related World Model architectures; experience in VLM/LLM fine-tuning and alignment is preferred.
  • Must have practical experience in successfully deploying end-to-end policies on real complex hardware such as bipedal humanoid robots, quadruped robots, and high-degree-of-freedom dexterous hands, and be able to handle real physical issues such as motor characteristics, sensor noise, friction, and contact.
  • Proficient in Python/C, familiar with Linux/Tmux development environments, capable of developing high-performance operators or custom C deployment environments.
  • Possess strong automation awareness, skilled in using Agentic AI development tools to improve R&D efficiency and experiment iteration speed.

Preferred Qualifications

  • Able to run closed-loop Real-world RL in practice and achieve clear research or business results.
  • Deep understanding of the Sim-to-Real-to-Sim flywheel, with experience in system identification using real-world residual data to optimize simulation parameters in Isaac Sim, MuJoCo, etc.
  • Have deep expertise in high-dynamic motion control or precise dexterous manipulation.

Interested in This Role?

Submit your application below and our team will review your resume shortly.