• Design, develop, and maintain production data pipelines on Databricks using Python, SQL, Apache Spark, and Delta Lake.
• Build Bronze-layer ingestion that reliably captures data from APIs, relational databases, flat files, cloud storage, and SaaS platforms — using dlt (dltHub) and Databricks-native ingestion where each fits — including incremental loading, pagination, watermarking, state management, and replay after failure.
• Develop Silver-layer transformations in dbt and Python over Delta Lake that cleanse, standardize, type, deduplicate, validate, conform, and enrich data so that it is reusable across domains. A meaningful share of this role is making messy source data trustworthy.
• Create Gold-layer data products: dimensional models, slowly changing dimensions, fact and bridge tables, aggregates, and serving tables aligned to how consumers actually query.
• Produce and maintain the curated datasets ML engineering trains and serves models from — feature and training tables that are versioned and reproducible, not one-off extracts.
• Author and maintain data contracts using the Open Data Contract Standard (ODCS) — schema with real semantics, named owner, known consumers, quality rules, and freshness expectations — and assess backward compatibility before every change.
• Implement data quality as code: uniqueness and not-null on keys at minimum, plus referential, accepted-value, freshness, and custom business-rule tests, surfaced to producers and consumers rather than buried in logs.
• Orchestrate ingestion and transformation as assets in Dagster, deployed to Dagster Cloud, and operate what you build across development, branch, and production deployments — schedules and sensors, asset dependencies, backfills, and run observability.
• Apply governance through Unity Catalog — catalogs, schemas, external locations, grants, row- and column-level security, and lineage — and handle credentials through Azure Key Vault rather than in code.
• Implement incremental and merge-based processing with Delta Lake (MERGE, schema evolution, time travel, OPTIMIZE) and tune Spark jobs, table layouts, and compute for performance and cost.
• Troubleshoot production failures, data-quality issues, source-system changes, and late-arriving or duplicate data — including backfills and recovery — and take part in the pod’s on-call rotation for the pipelines it owns, with root-cause analysis that closes the gap rather than reopening the ticket.
• Build and maintain CI/CD for data assets in Azure DevOps — automated tests and CI checks on dlt, dbt, and Dagster changes, promotion from development through branch deployments to production, and releases that are repeatable and auditable.
• Instrument what you own for observability: freshness, volume, quality, latency, and cost, with alerting tied to the SLAs and SLOs your contract commits to instead of depending on someone noticing.
• Work inside the platform’s control expectations — least-privilege access, secrets in Azure Key Vault, change management through pull request and pipeline, and audit evidence that falls out of the deployment path rather than being reconstructed later.
• Participate in code review and document architecture, runbooks, and data products so others can discover, trust, and reuse them.
• Work with data architects, analysts, product owners, and business stakeholders to translate requirements into maintainable data solutions.