As a Platform Engineering Architect you will focus on operationalizing, securing, and maturing artificial intelligence capabilities by building and maintaining the AI “path-to-prod” and a scalable AI run stack—the AI “plumbing.” Your work centers on providing a unified interface for AI model access, integrating foundational and reasoning models, AI agents, and generative AI. Key responsibilities include contributing to the K8s-based AI access platform, managing deployment of core AI services, integrating frontier models (e.g., Claude, GPT) and local inference engines (e.g., vLLM) along with designing MLOps pipelines. The role also requires providing expert recommendations for enterprise AI adoption, specifically in agentic orchestration and spec-driven development.
Core Requirements
· Professional Experience: 7+ years of combined experience in DevSecOps, Platform Engineering, or SRE.
· Kubernetes Mastery: Deep expertise in Administration and Development of Kubernetes clusters.
· Containerization: Advanced knowledge of Docker or equivalent container build tools.
· Cloud & IaC: Experience with Azure or AWS architecture and Infrastructure as Code (Terraform/Crossplane).
· Software Development: Proficiency in at least one backend or scripting language—ideally Python, Go, or Bash—to drive systems automation.
· Core Networking: Solid understanding of OSI Layer 4–7, including VPC/VNET configuration, DNS, Load Balancing, and SSL/TLS management.
· AI Ecosystem: Practical experience with modern AI/ML frameworks and tooling such as PyTorch, Hugging Face, LangChain, vLLM, Ray, MLflow, or equivalent open-source ecosystems.
· LLM & AI Infrastructure: Hands-on experience deploying, scaling, and securing AI/ML workloads on Kubernetes, including GPU-enabled clusters, model-serving platforms, and distributed inference/training systems.
· AI Platform Operations: Experience building internal AI platforms or developer enablement tooling that supports model lifecycle management, experimentation, inference endpoints, and reproducible AI workflows.
· MLOps & AI Delivery Pipelines: Familiarity with MLOps concepts and tooling, including automated model deployment, versioning, evaluation, observability, rollback strategies, and CI/CD integration for AI systems.Vector & Retrieval Systems: Working knowledge of vector databases, embedding pipelines, retrieval-augmented generation (RAG), and semantic search architectures.
· AI Security & Governance: Understanding of AI security concerns including model isolation, data handling controls, prompt injection risks, supply-chain security, and governance requirements for sensitive or regulated environments.