Find jobs
Pricing
FIND JOBSPRICINGSIGN INSIGN UP
← Back
HY
Hyphen Connect Limited

LLM Pre-training & Distributed Engineer (AI Infrastructure)

All locations for this role
United States, Hong Kong
United States, Singapore
United States, San Francisco Bay Area, USA
United States, Australia
United States, China
United States, Boston, USA
United States, Seattle, USA
United States, Oregon, USA
Location
United States, China
Type
full-time
Department
Engineering
Last seen
1h ago
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing  distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
•
Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
•
Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
•
Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
•
Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
•
Experience managing SLURM or Kubernetes-based GPU clusters.
•
Strong systems engineering background (C++, CUDA, Python).
Seekless
Product
Find jobsPricing
Legal
PrivacyTermsCookies
© 2026 SEEKLESS, INC. ALL RIGHTS RESERVED.BUILT WITH CARE