● Pipeline Development & Maintenance: Develop, deploy, and maintain scalable ETL/ELT
pipelines in GCP using Cloud Composer and PySpark on Dataproc following established
architectural patterns.
● Performance Optimization: Implement and utilize BigQuery features like Partitioning,
Clustering, and Materialized Views to improve query performance and contribute to GCP
cost management.
● Data Integration: Assist in the implementation of BigLake components to ensure
structured and unstructured data is accessible with unified governance.
● Code & Infrastructure Deployment: Utilize Terraform and CircleCI for deploying and
managing pipeline code and related infrastructure to ensure a modular and reproducible
environment.
● Data Operations & Monitoring: Implement robust logging, alerting, and data quality
checks within pipelines to actively monitor data flow and ensure data reliability.
● Collaboration: Work with senior engineers and cross-functional teams to understand
business requirements and translate them into efficient data solutions.
● Skill Development: Proactively seek out and learn new GCP services and tools to enhance
data processing capabilities.
Skills & Competencies
● Experience: Mid Level - 2 - 4 years of hands-on Data Engineering experience.
● GCP Proficiency: Solid knowledge of GCP services with a focus on BigQuery, DataForm,
Cloud Composer, Cloud Storage, Cloud functions and Dataproc.
● Transformation Tools: Proficiency in Dataform for managing version-controlled SQL
workflows and modeling within BigQuery.
● Strong Coding: Excellent Python and SQL skills. You write modular code that is easy to
test and maintain.
● Architectural Knowledge: Solid understanding of Data Modeling (Star Schema,
Snowflake) and Data Operations.
● Automation Mindset: Experience with CI/CD tools and Infrastructure-as-Code is essential.
● Domain Knowledge: Experience in the News or Media industry is a strong plus.