Production-Scale Experience: Proven ability to operate and troubleshoot production workloads across multiple tenants or environments. Deep understanding of how distributed systems fail at scale and how to build in resilience.
Terraform & IaC Mastery: Strong hands-on experience with Terraform, including workspace strategies, state management, and automation patterns that scale. Comfortable solving state isolation issues and building reliable, reusable infrastructure code. Experience with Ansible and templating tools like Jsonnet is a plus.
Kubernetes in Production: Skilled at diagnosing deployment failures, interpreting pod logs, and debugging scheduling issues and rollback scenarios in live environments. Understands how pods, ReplicaSets, and controllers interact in production.
Programming & Code Analysis: Ability to read and debug code in Go and/or Ruby. Familiar with identifying performance issues, scalability concerns, and contributing to infrastructure tooling through thoughtful code analysis.
Large Scale Operations Background: Experience supporting infrastructure for many customers or environments simultaneously. Comfortable managing isolation, scaling, monitoring, and incident response across diverse workloads.
Architecture & Incident Response: Able to reason through complex systems and operational challenges. Brings on-call experience and can lead technical discussions and incident resolution efforts under pressure.
Customer-Focused Collaboration: Proven ability to work across teams and with internal or external customers to solve technical problems while maintaining service commitments and clear communication.
GitLab Platform Proficiency: Comfortable using GitLab as a daily tool for infrastructure automation, collaboration, and operational workflows.