System Reliability & Performance
• Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
• Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
• Lead incident response, conduct root cause analysis, and implement preventive measures
• Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
• Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
• Automate deployment pipelines, monitoring, and operational workflows
• Optimize cloud resource utilization and cost management
Engineering & Development
• Build and maintain internal tools and services to improve operational efficiency
• Collaborate with development teams to implement reliability best practices
• Conduct code reviews and provide technical guidance on system design
• Develop monitoring solutions, alerting systems, and observability frameworks
• Integrate security practices into CI/CD pipelines (SAST/DAST)
• Implement and maintain security controls across infrastructure and applications
• Ensure compliance with industry standards and regulatory requirements
• Conduct security assessments and vulnerability management
Leadership & Collaboration
• Mentor junior SRE team members and promote SRE culture across the organization
• Partner with software engineering teams to improve system reliability
• Drive technical initiatives and contribute to architectural decisions
• Document processes, runbooks, and technical specifications