Bachelor’s or master’s degree in computer science, Data Engineering, Data Science, or a related field, or equivalent practical experience.
3–6+ years of experience in Data Engineering, with a proven track record of designing and deploying data pipeline solutions at scale. At least 2+ years of experience building complex data science, AI/ML, or large-scale analytics solutions.
Proficiency in Python and SQL; familiarity with TypeScript/JavaScript or a systems programming language such as Go or Rust. Experience with Test-Driven Development (TDD), CI/CD pipelines, and modern software engineering best practices.
Experience with CI/CD platforms (e.g., GitHub Actions), containerization technologies (Docker), infrastructure-as-code principles, and cloud-native architecture patterns.
Hands-on experience building and operating data and AI platforms on AWS, including services such as S3, Glue, Lambda, IAM, Bedrock, and Databricks Unity Catalog. Experience designing secure, scalable, and production-grade cloud architectures is preferred.
Proficiency with Databricks (Spark), data orchestration platforms (Airflow), Tableau, Genie Spaces, and related business intelligence and analytics technologies.
Hands-on experience with AI coding tools (Claude Code, GitHub Copilot, Cursor, or equivalent) and Cortex AI or comparable LLM-serving platforms. Strong understanding of how large language models generate and reason about code, with practical experience in prompt engineering as an engineering discipline.
Strong understanding of data privacy, security, and regulatory requirements in healthcare environments. Experience with secrets management, audit controls, and compliance frameworks, including HIPAA, SOC 2, and 21 CFR Part 11.
Ability to design and operate systems that scale across both traditional ML infrastructure and emerging agentic AI architectures.