We are hiring a Platform Engineer to own the runtime our entire product sits on: infrastructure, orchestration, CI/CD, observability, and developer experience. You keep an AI-native, multi-tenant platform reliable, fast, and easy to build on.
WHY THIS ROLE
•
Own the foundation the whole product and team depend on.
•
Shape reliability, performance, and cost for a multi-tenant AI platform.
•
Multiply every engineer’s output through better tooling and infrastructure.
WHAT YOU’LL DO
•
Own Kubernetes (GKE), infrastructure-as-code (Terraform), and our cloud (GCP) footprint.
•
Build and operate CI/CD, observability (logs, metrics, traces), and orchestration (Temporal, gateway/Kong).
•
Improve reliability, performance, and cost across services and the data plane’s infrastructure.
•
Strengthen multi-tenant provisioning reliability and internal developer experience.
•
Lead on-call practices and incident response.
WHAT WE’RE LOOKING FOR
•
4+ years in platform/infrastructure/DevOps/SRE with production ownership.
•
Strong Kubernetes, Terraform, and a major cloud (GCP/AWS).
•
Solid CI/CD and observability experience, plus scripting in Python or Go.
•
A reliability mindset and calm under incidents.
BONUS IF YOU HAVE
•
Multi-tenant SaaS or data-platform infrastructure experience.
•
Temporal, Kong/API gateway, or ClickHouse operations experience.