Responsibilities will include:
• Improve and maintain CI/CD, deployment workflows, and environment management across backend, web, and internal services
• Build, maintain and scale infrastructure across AWS and container based services
• Improve monitoring, alerting, logging, dashboards, tracing, and runbooks
• Work with engineers on safer deploys, rollback plans, and recovery from failures
• Automate repetitive operational work and improve internal tooling
• Maintain and improve infrastructure as code and deployment tooling
• Help improve failover planning, recovery procedures, and backup/restore testing for critical systems
• Support production systems and take part in on-call for critical services
• Manage and scale infrastructure across AWS, ECS, Docker, PostgreSQL, Redis, Celery, and Go/Python-based services
• Lead incident response and postmortems, and drive follow-up actions to reduce repeat issues
• Improve reliability, resilience, and operational readiness across critical systems