Roles and Responsibilities: Observability and Monitoring:
• Design, implement, and maintain observability solutions for datacenter infrastructure.
• Develop, deploy, and maintain the operational and reliability components of a large-scale Observability and Telemetry collection platform, emphasizing performance at scale, real-time monitoring, logging, and alerting.
• Participate in and enhance the entire lifecycle of services, from inception and design to deployment, operation, and refinement.
• Develop and optimize monitoring systems to ensure high availability and performance.
• Create and manage dashboards, alerts, and reports to provide visibility into system health and performance.
Performance Optimization:
• Analyze and optimize the performance of datacenter systems and applications.
• Implement best practices for resource utilization and efficiency. Collaboration:
• Work closely with other engineering teams to understand and meet their observability and reliability requirements.
• Collaborate with hardware and software vendors to evaluate and integrate new technologies.
• Ensure that observability and reliability solutions comply with security policies and industry standards.
• Implement and maintain security measures to protect data and infrastructure.
Troubleshooting and Support:
• Provide support for observability and reliability-related issues, including debugging and resolving hardware and software problems.
• Develop and maintain documentation for troubleshooting procedures and best practices.
• Stay updated with the latest advancements in observability and SRE technologies and integrate them into the infrastructure.
• Continuously improve the reliability, scalability, and performance of datacenter services.