We are looking for a DevOps / Senior Cloud Infrastructure Engineer to join us on a freelance basis to build the observability and incident management backbone for autonomous kitchen operations.
4-5 month project | full time | remote
What we’re looking for
•
Senior-level experience as an AWS Cloud Engineer or Site Reliability Engineer.
•
Strong, demonstrable focus on metric-driven observability, monitoring, and alerting at scale.
•
Fluent, hands-on experience in Python for tooling and automation.
•
Hands-on experience with Terraform for Infrastructure as Code.
•
A proven track record of designing, architecting, and owning production systems end-to-end.
•
Ability to work completely independently without a detailed spec sheet or heavy direction.
•
Experience with Jira Service Management or similar ITSM/incident platforms is a plus.
•
Experience with Grafana dashboarding pipelines at scale is a plus.
•
Exposure to AI-assisted ops tooling (AI-Ops, runbook automation) is a plus.
•
Prior experience in a fast-moving hardware, robotics, or IoT fleet environment is a plus.
What you’ll do
•
Standardize and own the Tech Ops Incident Management Platform using Jira Service Management across our production fleets.
•
Automate incident resolution workflows with AI tooling, including runbook generation and assignment.
•
Design and implement proper documentation for all Tech Ops incident processes to ensure a clean handover.
•
Build out fleet management and task automation to support global 24/7 remote operations.
•
Own and customize the data-driven observability layer via Grafana across all internal tech teams.
•
Work closely with key leadership stakeholders (Head of Infrastructure & Cloud, VP Engineering) to independently drive architecture decisions.