We’re building a new category of interactive entertainment powered by AI characters, image, video, and real-time experiences. We’re one of the most-visited AI products globally, serving millions of users every day.
•
We work in-person in Sydney, Australia, and hire globally.
How we work
•
User-first: We build what people want. We invest time to understand our users and focus on adding value instead of extracting value.
•
High agency, high ownership: We own outcomes end-to-end. When something goes wrong, we take responsibility and fix it.
•
Urgency: We prioritize ruthlessly, increase leverage, and move at an exceptional pace.
What you’ll do
•
You’ll be our first dedicated site reliability engineer, owning reliability and core platform decisions as we scale to hundreds of millions of users.
Example projects
•
Improve uptime and reduce RTO across critical services.
•
Orchestrate and harden GPU clusters serving millions of AI generations per day.
•
Implement platform-wide observability (metrics, tracing, alerting) and enforce SLOs.
•
Optimize AWS infrastructure and reduce cloud spend without sacrificing performance.
What you’ll bring
•
5+ years operating production systems at scale.
•
Strong AWS experience (infra-as-code, high-scale compute, K8s/ECS or similar).
•
Deep observability and incident response experience.