You should have:
● Strong hands-on Python skills—this is your primary language for automation and tooling.
● Experience in SRE, Production Engineering, Platform Engineering, or a related discipline with direct production ownership.
● Proven track record of building automation and diagnostic tooling that improved recovery times or reduced operational toil.
● Deep familiarity with cloud-native technologies—Kubernetes, containers, distributed systems—and how they fail in production.
● Experience with observability platforms such as Datadog, and a strong intuition for what “good” monitoring looks like.
● Exposure to Infrastructure as Code (Terraform) and GitOps-based deployment workflows (ArgoCD, GitHub Actions, or similar).
● Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, and Postgres.
● Strong analytical and problem-solving skills—you thrive on ambiguous, high-stakes production problems.
● A product mindset applied to operational tooling: you think about usability, adoption, and documentation when building internal solutions.
● Excellent communication skills and the ability to work fluidly across engineering, operations, and business stakeholders.
● Self-starter mentality—you identify opportunities, take initiative, and deliver with minimal supervision.
● Curiosity and a continuous learning mindset; fintech or financial industry background is a plus.