▸ Design and implement scalable batch and streaming data pipelines using Apache Spark, Kafka, and Flink
▸ Build and maintain the Bronze/Silver/Gold medallion architecture within the lakehouse (Delta Lake / Iceberg)
▸ Develop and optimize complex SQL and PySpark transformations for large-scale datasets
▸ Integrate structured, semi-structured, and unstructured data sources into the lakehouse
▸ Collaborate with data architects to evolve the physical and logical data models
▸ Implement data quality checks and monitoring using Great Expectations or dbt tests
▸ Write Infrastructure-as-Code for pipeline environments (Terraform, Helm)
▸ Participate in code reviews and enforce engineering standards and best practices
▸ Troubleshoot pipeline failures, performance bottlenecks, and data incidents
▸ Mentor junior and mid-level data engineers and contribute to internal knowledge sharing