Building Scalable Data Pipelines
Design and develop high-quality, scalable ETL/ELT pipelines for processing big data using AWS analytical services. Leverage no-code tools and reusable Python libraries to ensure efficiency, maintainability, and reusability. Ensure data pipelines and platform components are scalable, performant, and cost-efficient.
Build and enhance reusable platform capabilities across ingestion, transformation, orchestration, data quality, cataloguing, and monitoring. Develop scalable and reliable AWS data pipelines and data products using standardised data architecture patterns such as medallion architecture, lakehouse, and other modern design paradigms.
Drive configuration-driven, declarative, and automated engineering approaches that reduce bespoke development and improve productivity. Champion automation across the data engineering lifecycle to minimize manual intervention and accelerate delivery.
Collaborating & Aligning with Project Goals
Work closely with cross-functional teams, including Tech Leads, Engineering Managers, and Business Analysts, to understand project objectives and deliver robust data solutions. Follow Agile/Scrum principles to drive consistent progress and iterative delivery.
Data Discovery & Root Cause Analysis
Perform data discovery and analysis to uncover data anomalies. Identify and resolve data quality issues through root cause analysis, and provide informed recommendations for data quality improvement and remediation.
Champion the integration of Claude Code and other LLM tools into our software development lifecycle (SDLC). Lead the team in using AI to accelerate coding, debugging, automated testing, and documentation.
CI/CD & Operational Excellence
Manage the automated deployment of code and ETL workflows within cloud infrastructure (AWS preferred) using tools such as GitHub Actions, AWS CodePipeline, or other modern CI/CD systems. Implement Infrastructure as Code (IaC), automated testing frameworks, observability solutions, security best practices, and operational reliability measures to ensure robust and resilient data platform operations.
Effective Time Management
Demonstrate strong organizational and time management skills. Prioritize tasks effectively and ensure the timely delivery of key project milestones.
Set the gold standard for the team by leading code reviews, defining CI/CD patterns, and enforcing data governance standards via AWS Lake Formation.
The AI Frontier (SageMaker & Bedrock)
Architect the infrastructure for our AI initiatives. Implement our initial AWS SageMaker footprint (Data Wrangler, Feature Store) and manage AWS Bedrock integrations (Knowledge Bases, RAG pipelines) to support downstream Data Science and GenAI initiatives.
Documentation & Data Mapping
Develop and maintain comprehensive data catalogs, including data mapping and documentation, to ensure data governance, transparency, and accessibility for all stakeholders.
Learning & Contributing to Best Practices
Continuously improve your skills by learning and implementing data engineering best practices. Stay updated on industry trends and contribute to team knowledge-sharing and codebase optimization.
What You Will Be Doing (Daily Execution & Process)
Build and evolve reusable data platform capabilities. Develop scalable ingestion, ETL/ELT and data processing solutions using AWS. Implement and support medallion data processing and data product patterns. Build automation around data quality, metadata, cataloguing, orchestration and monitoring.
Function as an integral part of an Agile engineering team, defining delivery phases, sub-activities, and milestones while ensuring development standards are rigorously enforced.
Collaborate with Engineering Managers and Business Analysts to produce accurate delivery estimates and oversee the transition from analysis through to final delivery.
Infrastructure as Code (IaC)
Create and maintain resilient infrastructure using Terraform or AWS CloudFormation and strong deployment practices enabled by automated tests to ensure environment parity and automationstability.
Implement comprehensive unit testing strategies and automated data quality checks to identify and resolve issues prior to deployment.
Observability & Monitoring
Establish and maintain observability practices including logging, monitoring, and alerting to ensure platform health, performance optimization, and proactive issue detection.
Cross-Functional Collaboration
Work closely with cross-functional teams to gather requirements and develop cloud-based applications, analytical services, and platforms.
Ensure all data solutions strictly adhere to company security guidelines and regulatory requirements. Implement security best practices throughout the data engineering lifecycle.
Document technical solutions and maintain a robust knowledge base for the team.