8+ years of hands-on technical experience in software engineering, infrastructure, or
operations roles, including a minimum of 5 years dedicated to Site Reliability Engineering.• Expert-level, well-rounded SRE skill set — proficient across monitoring/alerting, incident
response, capacity planning, performance optimization, CI/CD, and reliability engineering best
practices.
• Deep hands-on expertise with New Relic or a comparable observability platform; strong
preference for candidates who have led observability platform adoption or migration at scale.
• Demonstrated experience owning incident management programs: on-call processes,
escalation design, post-mortem culture, and measurable MTTR/MTTD improvement.
• Strong proficiency in Python, Bash, PowerShell, and other common SRE scripting and
automation technologies.
• Expert-level experience designing, building, and maintaining autonomous systems that handle
software build, deployment, testing, monitoring, and operations.
• Proficient hands-on experience with AWS (EC2, EKS/Kubernetes, CloudWatch, Lambda, S3,
IAM) and the broader cloud-native ecosystem.
• Strong communicator who proactively informs stakeholders, operates transparently, and can
bridge technical complexity for product and management audiences.
• Proven track record of mentoring engineers, leading initiatives to completion, and making
those around them measurably better.
• Bachelor’s degree in Computer Science, Information Systems, or a related field; equivalent
certifications (e.g., AWS certifications, Google Cloud Professional); or substantial comparable
direct work experience