Duties: Design, develop, and maintain large-scale data ingestion, storage, and querying systems
supporting AI/ML experiment tracking and analytics. Architect and optimize distributed systems
and backend services to improve performance, reliability, scalability, and cost-efficiency across
petabyte-scale datasets. Implement and enhance customer-facing APIs using modern
programming languages (e.g., Go, Python, TypeScript) and storage technologies such as
MySQL, Bigtable, ClickHouse, Postgres, and Kafka/PubSub. Build, deploy, and operate services
on Kubernetes-based infrastructure, using tools such as Terraform and major cloud platforms
(GCP, AWS, or Azure). Diagnose and resolve complex production issues related to distributed
systems, data pipelines, and storage/query performance. Collaborate with cross-functional teams
(product, revenue, and engineering) to define technical requirements, design system
architectures, and deliver new platform capabilities. Lead long-term architectural initiatives to
evolve metrics and storage infrastructure, ensuring secure, reliable, and cost-optimized operation.
Mentor junior engineers by providing technical guidance, code reviews, design feedback, and best
practices for high-quality software development.