All about you:
5–10 years of experience in an SRE or SRE related operations role, including 3+ years supporting e commerce, financial services, or large scale SaaS platforms.
Excellent infrastructure troubleshooting and analytical problem solving skills.
Strong hands on experience with observability and monitoring tools such as Splunk, Dynatrace, or equivalent, with a proven ability to triage and investigate complex issues.
Familiarity with network telemetry tools such as SolarWinds and NetScout.
Proficiency in packet level debugging, including capturing traffic with tools like tcpdump and analyzing packets using Wireshark.
Broad understanding of end to end infrastructure supporting payment platforms—spanning platform services, networking, databases, and storage.
Experience with automation and Infrastructure as Code tools such as Chef, Ansible, and Terraform, as well as structured data formats (JSON/YAML).
Excellent communication skills with the ability to coordinate cross functional troubleshooting efforts and lead RCA processes to closure.
Demonstrated ability to troubleshoot complex production issues, perform root cause analysis, and drive long term corrective actions.
Experience partnering with development teams to shape architecture, define SLIs/SLOs, and embed reliability into services from design through operation.
Strong understanding of monitoring and observability ecosystems, including Prometheus, Grafana, ELK/EFK, Splunk, and OpenTelemetry.
Effective incident management skills with a structured, analytical approach to problem solving.