Training & Inference Ownership. Own key components of the training-to-deployment pipeline — from distributed training execution through inference optimization, serving, and production handoff — ensuring models are delivered reliably, performantly, and cost-efficiently.
Large-Scale Distributed Training. Implement and operate distributed training strategies including PyTorch FSDP, Tensor Parallelism, and Pipeline Parallelism across multi-node GPU environments, ensuring correctness, stability, and scalability for large video and multimodal models.
Inference & Serving. Design and optimize inference and serving systems for large generative models, with a focus on latency, throughput, and cost across deployment targets.
Research-to-Production Bridge. Reduce the gap between trained model checkpoints and reliable production deployments — owning the practical work of hardening, validating, and operationalizing models at scale.
Performance & Cost-Aware Engineering. Identify and address inefficiencies across the training and inference stack — memory, communication, scheduling, and execution orchestration — with a clear focus on GPU efficiency and cost targets.
Collaboration with Research & Engineering Teams. Partner closely with applied researchers, ML engineers, and infrastructure teams to align training and inference systems with model architecture needs and product delivery timelines.