• Master’s degree in Data Science, Computer Science or closely related field and 3 years of experience designing ETL pipelines. Will accept a Bachelor’s degree and 7 years of experience.
• Must have experience with:
• Designing and developing robust ETL pipelines using PySpark, Scala, SQL, and Shell scripting
• PostgreSQL, MySQL, or SQL Server
• Using multiple file formats in data pipelines, including Parquet, CSV, and ORC
• Big Data technologies such as Hadoop, Hive, Spark, or Kafka
• AWS tools such as Glue, IAM, EC2, EMR, RDS, S3, Athena, Step Functions, or Lambda
• Data mapping, validation, and modeling
• Building and maintaining CI/CD pipelines using GitHub or GitLab
• Reporting and BI tools such as Tableau or Power BI
• Working with Delta tables in Databricks
• RabbitMQ for queuing high-speed sensor data
• Natural Language Processing (NLP) techniques
• DevOps tools such as Terraform, Git, and Jira