Job Duties: Design, develop, and maintain scalable backend data pipelines for processing large-scale content catalog and metadata using distributed data processing frameworks such as Apache Spark. Develop and optimize high-throughput data processing workflows capable of ingesting, transforming, and aggregating large-scale datasets with optimal performance and reliability. Design and implement secure, scalable RESTful APIs to expose data services to internal platforms and downstream consumers. Architect and build data integration solutions that connect heterogeneous data sources and destinations across the platform, ensuring data consistency and availability. Collaborate with data scientists and machine learning engineers to integrate models and algorithms into production environments, including supporting feature engineering and model-serving pipelines. Implement monitoring, alerting, and observability frameworks for data pipelines to ensure reliability, performance, and high availability at scale. Optimize data storage and retrieval mechanisms, including data modeling, indexing strategies, and query performance tuning across relational databases, NoSQL stores, and cloud-based data warehouses. Contribute to architectural decisions for backend and data infrastructure, partner with cross-functional stakeholders to translate business requirements into technical solutions, and apply data engineering best practices including data quality validation, error handling, and recovery mechanisms.