· Design, develop, and optimize scalable data pipelines for ingesting, processing, and storing large volumes of structured and unstructured data.
· Implement distributed data processing solutions using technologies such as Spark, Kafka, and cloud-native services.
· Ensure data integrity, reliability, and high availability in mission-critical workloads.
· Model and design databases and data architectures to support advanced analytics, machine learning, and AI applications.
· Develop and maintain ETL processes for transforming and loading data from diverse sources.
· Deploy data infrastructure and distributed computing environments in the cloud using containerization (Docker, Kubernetes).
· Monitor, benchmark, and tune data processing applications for optimal performance and scalability.
· Collaborate with cross-functional stakeholders to understand business requirements and deliver tailored data solutions.
· Implement robust testing, observability, and monitoring solutions to maintain system health and performance.
· Stay current with technology trends, best practices, and industry standards in data engineering and platform development.