● Build Spark ingestion jobs that land high-volume raw files into Iceberg tables with schema handling, bad-record quarantine, and idempotent batch replay.
● Develop and operate Trino SQL transform pipelines across data layers: validation and
typing, business-rule transforms, MERGE-based deduplication, and aggregate builds.
● Port existing warehouse SQL workloads to Trino and Spark SQL dialects, and validate results against source outputs.
● Automate Iceberg table maintenance: compaction, snapshot expiry, and orphan-file cleanup as scheduled workflows.
● Tune query and pipeline performance: partitioning strategy, file sizing, statistics, and resource-group behavior.
● Instrument pipelines with data-quality checks, reconciliation reports, and alerting, and participate in per-dataset validation during rollout phases.