•
4+ years in MLOps, ML platform engineering, or infra-focused ML roles
•
Deep familiarity with model lifecycle management tools: MLflow, Weights & Biases, DVC,
•
Experience with large model deployments (open-source LLMs preferred): LLaMA,
•
Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA)
•
Familiarity with inference serving: vLLM, TGI, Ray Serve, Triton Inference Server
•
Proficient with Terraform, Helm, K8s, and container orchestration
•
Experience with CI/CD for ML (e.g. GitHub Actions + model checkpoints)
•
Managed hybrid workloads across GPU cloud (Lambda, Modal, HuggingFace Inference,
•
Familiar with cost optimization (spot instance scaling, batch prioritization, model sharding)
**Agent + Data Pipeline Support:**●
Familiarity with LangChain, LangGraph, LlamaIndex or similar RAG/agent orchestration tools
Built embedding pipelines for multi-source documents (PDF, JSON, CSV, HTML)
Integrated with vector databases (Weaviate, Qdrant, FAISS, Chroma)
Implemented model-level RBAC, usage tracking, audit trails
Integrated with API rate limits, tenant billing, and SLA observability
Experience with policy-as-code systems (OPA, Rego) and access layers