10-15 years of experience in SRE handling performance, architecture and design of applications
● Strong understanding of cloud computing, networking, Linux systems administration, containerization (e.g., Docker, Kubernetes), and infrastructure as code (e.g., Terraform, Ansible)
● Understanding of SRE principles, including SLOs, SLIs, SLAs, and error budgets.
● Experience in managing incident and retrospectives
● Experience in cloud cost management, cloud architecture
● In-depth knowledge of cloud computing platforms (e.g., AWS)
● Experience with infrastructure as code (IaC) tools and practices
● Experience with monitoring, logging & telemetry tools like New Relic, Splunk, ELK, Nagios, SolarWinds, Prometheus, AWS Cloudwatch, Datadog, Opentelemetry
● Expert in designing, creating and supporting Automation and Identify opportunities for self-healing systems, automated deployments, and other scalable solutions.
● Experience in performance engineering and identify opportunities for performance tuning and profiling
● Experience in prioritizing and managing technical roadmaps.
● Strong skills in stakeholder communication, requirements gathering, and documentation.
● Ability to work with cross-functional teams and build consensus around reliability goals
● Improve operational processes and team practices
● Problem-solving: Ability to analyze complex systems, troubleshoot issues, and devise effective solutions
● Excellent communication skills, with the ability to inspire and motivate cross-functional teams.
● Experience in dealing with the intricacies of large-scale distributed systems and ensuring their reliability and performance.