• Production delivery of data pipelines inside banks.
• Customer, account, transaction, payments and AML data.
• Working within bank release, scheduling and change controls.
• Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning.
• Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R.
• Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports.
• ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation.
• Reconciliation, data quality and SLA monitoring.
• Performance at scale: tables of 1 billion+ rows and multi-year history.
•
Lead: 12+ years, including 6+ in banking.
•
Senior: 8+ years, including 4+ in banking ·
• Production delivery of data pipelines inside banks.
• Customer, account, transaction, payments and AML data.
• Working within bank release, scheduling and change controls.
• Teradata SQL, BTEQ, TPT and FastLoad/MultiLoad; query tuning.
• Spark (PySpark/Scala), Hive, Impala, HBase, Pig and Hue; Apache Iceberg and Trino; Python and R.
• Historical data processing: snapshots, SCD Type 2 and multi-grain backfills with control totals and validation reports.
• ETL/ELT development, CDC, metadata-driven automation frameworks and parameter-driven SQL generation.
• Reconciliation, data quality and SLA monitoring.
• Performance at scale: tables of 1 billion+ rows and multi-year history.
Integration and platforms
• Kafka, Spark Streaming, Informatica (PowerCenter, IDMC, IDL) and Talend; REST API development.
• Denodo, Snowflake and NoSQL databases.
• Airflow, Control-M or Autosys; Git, CI/CD, GitOps and Kubernetes.
Delivery and communication (Lead)
• Framework design, code standards, code reviews and estimation.
• Guiding a team of engineers and working with architects and analysts.