Find jobs
Pricing
FIND JOBSPRICINGSIGN INSIGN UP
← Back
EM
Embedding Vc

Member of Technical Staff - Efficient ML

Location
San Francisco Bay Area
Work mode
on-site
Type
full-time
Department
Moonlake
Last seen
23h ago
Introducing Moonlake, AI for creating world simulations.
Scope of Work
Training efficiency
•
Dataloaders, fusion, activation remat, gradient checkpointing.
•
FSDP/ZeRO/tensor+pipeline parallel; NCCL tuning.
GPU + kernel performance
•
Nsight profiling, Triton/CUDA kernels, fused ops.
•
Flash-attention–style speedups, sequence packing, KV-cache tricks.
Inference optimization
•
Low-latency serving, continuous batching, speculative decoding.
•
Quantization (GPTQ/AWQ), distillation, pruning.
Infra + reliability
•
SLURM/K8s multi-node jobs, checkpoint hygiene.
•
Determinism, env pinning, GPU failure handling.
We are committed to being an on-site, in-person team currently based in San Mateo
Seekless
Product
Find jobsPricing
Legal
PrivacyTermsCookies
© 2026 SEEKLESS, INC. ALL RIGHTS RESERVED.BUILT WITH CARE