• Lead a technical initiatives of site reliability support services before they go live through activities such as system design consulting, performance engineering, chaos testing, capacity planning and launch reviews.
• Increase, maintain and communicate service metrics once live by measuring and monitoring availability, latency, performance and overall system health.
• Scale systems sustainably through mechanisms, like automation, and evolve systems by pushing for changes that improve reliability, velocity and recommend performance tuning approaches.
• Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
• Review production incidents to identifying and driving solutions to prevent recurrence.
• Manage individual project priorities, deadlines, and deliverables.
• Create and maintain technology roadmaps.
• Look at all tasks with an eye for automation; then work to automate them.
• Build, manage and maintain robust dashboards reflecting system health.
• Applies expert technical capabilities across discipline(s) to coach and mentor technical talent.