• Genuinely hands-on coding: 8+ years of professional software engineering experience building and
operating large-scale distributed platforms; strong proficiency in Java and Python is required -
comfortable writing, reviewing, and debugging production code, not just directing others. Additional
experience with Go, C++, or Scala is a plus.
• Operational Excellence (OE) and Engineering Excellence (EE) mindset is a required skill, not a
nice-to-have: SLO/SLA ownership, incident management and blameless post-incident reviews,
proactive reliability and quality metrics, and continuous improvement built into how you ship, not
• Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or
equivalent practical experience.
• Deep experience with microservices architectures, distributed systems, event-driven patterns, API
design, service contracts, and enterprise system integration - backed by strong system design,
domain-oriented design (DDD), and low-level design (LLD) skills, with demonstrated ownership of
production code and the judgment to evaluate trade-offs, failure modes, scalability, and cost.
• Deep understanding of cloud-native engineering practices - CI/CD, containerization, Kubernetes or
equivalent orchestration, infrastructure automation, and production operations.
• Strong exposure to Big Data / data engineering technologies - Spark required at a working level, with
familiarity in Hadoop/HDFS, Hive, or Kafka - enough to flex into data engineering work confidently, not
just integrate with it from a distance.
• Full exposure to AI and Agentic technologies: hands-on experience integrating model APIs and
GenAI/LLM tooling, building or operating agentic workflows (tool-calling agents, multi-step
autonomous systems), working with embeddings/vector search, and applying MLOps practices
(feature stores, model lifecycle, retraining pipelines) - not just conceptual awareness.
• Strong knowledge of reliability engineering - observability, distributed tracing, logging, metrics,
alerting, incident response, capacity planning, and fault tolerance for large-scale systems.
• Strong grounding in security, privacy, and enterprise engineering practices, including authentication,
authorization, data protection, and secure API design.
• Proven ability to translate ambiguous technical/business problems into clear execution plans,
particularly within a newly formed team without fully established process.
• Strong ownership mindset and technical judgment; comfortable operating with high autonomy in a