Basic 5+ years of experience in data engineering or related roles Experience with Python for data engineering and automation Experience with distributed computing frameworks like Apache Spark or Ray Experience integrating LLMs (e.g., OpenAl, Azure OpenAl, Anthropic, or similar) into applications via APIs, including prompt design, response parsing, and error handling Experience with distributed systems concepts including data partitioning and sharding strategies, fault tolerance and replication, consistency models and distributed consensus, load balancing and resource management Experience with distributed file systems (Azure Data Lake, S3, HDFS) Experience implementing REST APIs using Python frameworks such as FastAPI, Flask, or similar Experience with containerization and orchestration (Docker, Kubernetes) Experience with SQL and database optimization techniques Experience with data structures, algorithms, and software design patterns Experience with version control systems (Git) and CI/CD pipelines Bachelor’s or Master’s degree in Computer Science, Engineering, or related field, or equivalent practical experience Preferred Experience working with unstructured data (text, images, video, audio) and associated processing techniques Experience with multiple cloud platforms (e.g., AWS, Azure, GCP) Knowledge of data governance and security best practices Experience with machine learning pipelines and MLOps Experience with LLM orchestration frameworks (e.g., LangChain, LlamaIndex) Familiarity with responsible Al practices, including bias mitigation, content filtering, and token cost optimization for LLM-based applications Contributions to open-source projects Excellent problem-solving skills and ability to work with complex, ambiguous requirements Strong communication skills and ability to collaborate across teams