The Cloud Engineering team is the core of our tech unit, the underlying infrastructure that every other team depends on to ship software and grow the business. Behind rewarded ad experience, the partner-facing dashboards, the ML pipelines, and the ad delivery systems serving millions of users daily, there is one team making sure it all stays online, performant, and cost-efficient.
400 million user events daily, traffic spikes of 150,000+ requests per second, and a data lake that feeds data scientists, BIs, and real-time ML inference across the entire ad stack. To handle this reliably and cost-efficiently, the team runs Kubernetes with event-driven autoscaling and Karpenter, manages Kafka clusters for real-time streaming, and provisions everything through Terraform. Rather than defaulting to managed AWS services, the team actively replaces them with open-source alternatives like Druid, Airflow, Grafana, Prometheus, Spark, Trino, MLFlow, Jupiter Hub, Elasticsearch, and Fluentbit as part of a deliberate move toward cloud-agnostic architecture. Cost efficiency is treated as an engineering constraint, not an afterthought, including AI-based anomaly detection on infrastructure costs.