Back to Open Positions
Experienced Roles

Senior AI Infrastructure Engineer (Foundation Model Training and Inference Optimization)

Beijing
Full-time
Senior-level experience
Relevant background

Job Description

  • You will own the end-to-end performance, scalability, and reliability of foundation models, including LLMs, VLMs, multimodal models, and embodied AI VLA models, from training through inference. You will design and optimize distributed training strategies for clusters ranging from hundreds to thousands of GPUs, continuously improving model FLOPs utilization (MFU), while developing high-performance inference engines that reduce latency, increase throughput, and reliably bring algorithmic advances into production.

Job Responsibilities

  • *Training Direction
  • Responsible for the design of distributed training architecture and performance optimization for large-scale models with hundreds of billions of parameters, improving training throughput and scalability on large GPU clusters.
  • Design and implement hybrid parallel strategies such as data parallelism, tensor parallelism, pipeline parallelism, sequence parallelism, and expert parallelism to address communication bottlenecks in multi-machine multi-GPU training.
  • Conduct full-stack performance optimization around computing, communication, storage, I/O, video memory, and scheduling, continuously improving training efficiency and MFU.
  • Optimize pre-training, SFT, and RLHF processes based on frameworks such as Megatron-LM, DeepSpeed, and PyTorch FSDP, and support evolution of multimodal, complex structures, and ultra-large parameter models.
  • Build fast or asynchronous checkpoints, fault tolerance, exception detection, and automatic recovery mechanisms to ensure the stable operation of long-cycle large-scale training.
  • Transform cutting-edge technologies such as long sequences, FP8/FP4, computation-communication overlap, efficient pipeline scheduling, and operator fusion into industrial-grade engineering capabilities.
  • *Reasoning Direction
  • Develop and optimize high-performance model inference engines to reduce service latency and improve throughput and resource utilization.
  • Deeply apply inference frameworks such as vLLM, TensorRT-LLM, SGLang, Triton, etc., and carry out kernel-level optimization for key paths like FlashAttention and PagedAttention.
  • Optimize core inference processes such as computation graph compilation, Continuous Batching, and KV Cache management.
  • Explore and implement model compression and deployment techniques such as PTQ, QAT, FP8, INT8, pruning, and knowledge distillation.
  • *Technology Evolution and Collaboration
  • Continuously follow infrastructure technologies such as efficient communication, low-precision computing, FlashAttention, and RL training frameworks, promoting the transformation of research results into productivity.
  • Collaborate with algorithms, models, and research teams to carry out prototype validation, performance evaluation, and production implementation, balancing model effectiveness with training and inference efficiency.

Job Requirements

  • Having research and development experience in large-scale distributed systems, GPU training or inference infrastructure, high-performance computing, ML systems, etc.; those who have been responsible for training or deploying models at the scale of thousands of cards are preferred.
  • Proficient in Python and C, with a solid foundation in data structures, algorithms, and concurrent programming, and skilled in PyTorch; familiarity with TensorFlow or JAX is preferred.
  • Familiar with Transformer internals and the architectures of leading foundation models such as Qwen, Llama, GPT, and DeepSeek.
  • Proficient in at least one distributed training framework such as Megatron-LM, DeepSpeed, or FSDP, with a deep understanding of parallel strategies like DP, TP, PP, EP, and SP, as well as the PyTorch, CUDA, and NCCL training stack.
  • Familiar with GPU architecture, memory hierarchy, and principles of distributed computing; candidates familiar with CUDA/Triton programming and with experience in developing custom high-performance operators are preferred.
  • Possess practical experience in training or inference performance optimization, capable of conducting system analysis and optimization around GPU memory, communication, throughput, and latency.
  • Proficient in using performance analysis and debugging tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, pdb, and gdb.

Preferred Qualifications

  • Familiar with compiler or runtime technologies such as XLA / OpenXLA, torch.compile, MLIR, TensorRT-LLM, etc.
  • Familiar with reinforcement learning algorithms and their large-scale parallel training frameworks.
  • Contributed to open-source projects such as Hugging Face, vLLM, DeepSpeed, and SGLang.
  • Published relevant papers in conferences such as OSDI, SOSP, NSDI, MLSys, NeurIPS, ICML, ICLR, and ACL.

Interested in This Role?

Submit your application below and our team will review your resume shortly.