Qualifications & Experience
• Bachelor’s degree in computer science, Computer Engineering, or equivalent practical experience
• 5+ years of engineering leadership experience managing teams of 20+ members
• 12+ years of experience in DevOps and Site Reliability Engineering within large-scale production environments
• 7+ years of hands-on experience deploying, operating, and optimizing production workloads using CI/CD pipelines in AWS and/or GCP environments
• 5+ years of experience managing distributed systems and streaming platforms, including Kafka, Cassandra, Elasticsearch, Spark, Flink, Storm, and cloud-native services such as AWS EMR, Dataproc, ElastiCache, Amazon RDS, or Cloud SQL
• 5+ years of automation and software development experience using Python, Go, and/or Rust, along with strong shell scripting expertise
• 5+ years of experience defining, implementing, and operationalizing metrics to monitor infrastructure and application health, reliability, and performance
• Working knowledge of Infrastructure as Code (IaC) principles and tools such as Terraform, CloudFormation, or equivalent frameworks