• Strong hands-on experience with Big Data technologies including Hadoop, HDFS, Hive, Spark, and related distributed computing frameworks.
• Deep expertise in cloud data engineering on AWS, including storage, compute, networking, security, and scalable data architectures.
• Experience designing and developing modern data solutions using Databricks, Spark, Delta Lake, and cloud-native technologies.
• Proficiency in Python, Java, or Scala, building large-scale distributed data processing applications.
• Experience building robust batch and streaming pipelines supporting high-volume and high-velocity workloads.
• Understanding of modern Lakehouse architecture principles and open-table formats such as Apache Iceberg and Delta Lake.
• Skilled in data modelling, schema design, partitioning strategies, query optimization, and performance tuning.
• Familiarity with enterprise-scale data governance, security, privacy, and compliance, working with structured and unstructured datasets at scale.
• Experience with Infrastructure-as-Code (e.g., Terraform), modern CI/CD, and cloud-native monitoring and observability practices.
• Strong analytical and problem-solving skills, a passion for automation, and the ability to assess emerging technologies and recommend scalable, cost-effective solutions.
• Excellent communication and collaboration across technical and business teams; comfortable in Agile/Scrum environments managing multiple priorities.
• BS/MS degree in Computer Science, Software Engineering, Information Systems, or a related field.