Find jobs
Pricing
Free Resume Checker
Blog
Sign in
Sign up
FIND JOBSPRICINGFREE RESUME CHECKERBLOGSIGN INSIGN UP
TA
Takeda

Data Engineering Professional II

LocationIND - Bengaluru
Typefull-time
SeniorityMid
Experience4+ yrs
DepartmentDigital, Data & Technology (DD&T) - R&D MLOPs
Company size10,000+ people
First seenOct 6, 2026 · 5d ago
Verified live1d ago
At a glanceSummarised by Seekless from the posting.
Must have9
Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience)
4+ years of hands-on experience in MLOps, ML engineering, data engineering, or DevOps for data/ML systems
Strong Databricks experience: Delta Lake, MLflow, Unity Catalog, Jobs/Workflows, and Spark (PySpark)
Strong AWS experience across compute, storage, and IAM, plus at least one ML service (SageMaker and/or Bedrock)
Proficiency in Python for production code (packaging, testing, typing), plus solid SQL
Experience building CI/CD pipelines and using Git-based workflows
Working knowledge of containerization (Docker) and orchestration (Kubernetes/EKS or ECS)
Experience with Infrastructure as Code (Terraform, CloudFormation, or CDK)
Understanding of ML lifecycle concepts: experiment tracking, model registry, feature stores, and model monitoring/drift
Nice to have6
Experience delivering software or ML in a GxP / regulated (FDA, EMA) life-sciences environment; familiarity with CSV/CSA, GAMP 5, 21 CFR Part 11, ALCOA+
Exposure to pharma/biotech data domains: clinical trial data (CDISC/SDTM/ADaM), regulatory submissions, pharmacovigilance/safety, real-world data (RWD/RWE), or omics/biomarker datasets
Experience operationalizing LLM/GenAI workloads (e.g., AWS Bedrock), including RAG, prompt/version management, evaluation, and guardrails
Familiarity with handling PHI/PII under HIPAA/GDPR and with data-governance tooling
Streaming/event-driven data (Kafka/Kinesis), data-observability tooling, and feature-store frameworks
Relevant certifications: AWS (ML Specialty, Solutions Architect, or DevOps Engineer) and/or Databricks (Data Engineer, ML Engineer)
Eligibility1
By clicking the “Apply” button, I understand that my employment application process with Takeda will commence and that the information I provide in my application will be processed in line with Takeda’s Privacy Notice and Terms of Use. I further attest that all information I submit in my employment application is true to the best of my knowledge.
Skills
Databricks
Delta Lake
MLflow
Unity Catalog
Feature Store
Databricks Workflows/Jobs
Spark (PySpark)
AWS
SageMaker
Bedrock
S3
Lambda
ECS/EKS
Step Functions
ECR
IAM
CloudWatch
Infrastructure as Code (Terraform, CloudFormation/CDK)
Python
SQL
CI/CD (GitHub Actions, GitLab CI, Jenkins)
Git
Docker
Kubernetes/EKS
ECS
GitHub Actions
GitLab CI
Jenkins
Kafka
Kinesis
AWS ML Specialty
AWS Solutions Architect
AWS DevOps Engineer
Databricks Data Engineer
Databricks ML Engineer
Legal
By clicking the “Apply” button, I understand that my employment application process with Takeda will commence and that the information I provide in my application will be processed in line with Takeda’s Privacy Notice and Terms of Use. I further attest that all information I submit in my employment application is true to the best of my knowledge.
Job Description
Data Engineering Professional II
Digital, Data & Technology (DD&T) - R&D MLOPs
About the Role
We are seeking an MLOps Engineer to operationalize machine learning and generative AI across our R&D and enterprise data ecosystem. You will build and maintain the platforms, pipelines, and controls that move models from notebook experiments into validated, production-grade, GxP-compliant services: supporting use cases that span clinical development, regulatory operations, pharmacovigilance, real-world evidence, and translational/biomarker research.
This is a hands-on engineering role at the intersection of data engineering, ML lifecycle automation, and regulated-systems discipline.
You will work primarily in Databricks and AWS, partnering with data scientists, platform/cloud engineering, quality, and regulatory teams to ship models that are reproducible, monitored, auditable, and trustworthy.
Key Responsibilities
ML Lifecycle & Pipeline Automation
•
Design, build, and operate end-to-end ML pipelines (data ingestion → feature engineering → training → validation → deployment → monitoring) using Databricks (Delta Lake, MLflow, Unity Catalog, Feature Store, Workflows/Jobs) and AWS services.
•
Implement CI/CD for ML and data assets (e.g., GitHub Actions, GitLab CI, or Jenkins), including automated testing, environment promotion (dev → test → prod), and reproducible builds.
•
Stand up and maintain model registries, model versioning, and artifact lineage so every deployed model is traceable to its data, code, and configuration.
Cloud & Platform Engineering (AWS)
•
Build and manage ML infrastructure on AWS — e.g., SageMaker, Bedrock, S3, Lambda, ECS/EKS, Step Functions, ECR, IAM, CloudWatch — using Infrastructure as Code (Terraform or CloudFormation/CDK).
•
Integrate Databricks with AWS securely (Unity Catalog governance, cross-account access, VPC/networking, KMS encryption, secrets management).
•
Optimize compute and cost (cluster policies, autoscaling, spot strategy, job orchestration) without compromising performance or compliance.
Production Monitoring & Reliability
•
Implement model and data monitoring: drift detection, data-quality checks, performance/SLA tracking, and automated alerting/retraining triggers.
•
Establish observability and incident-response practices for ML services; participate in on-call/runbook ownership as needed.
•
Maintain feature stores and data contracts to ensure consistency between training and serving.
Regulated-Environment & Compliance Engineering
•
Build ML systems that meet GxP expectations and support Computer System Validation (CSV) / Computer Software Assurance (CSA), GAMP 5, 21 CFR Part 11, and data-integrity (ALCOA+) requirements.
•
Implement audit trails, electronic records/signatures controls, access controls, and change-management workflows suitable for validated environments.
•
Handle PII/PHI and sensitive R&D data in line with HIPAA, GDPR, and internal privacy/data-governance policies (de-identification, anonymization, role-based access).
•
Author and maintain technical documentation, validation deliverables, and SOP-aligned procedures; partner with Quality/QA and Regulatory on audits and inspections.
Collaboration & Enablement
•
Work under the guidance of Director, Solution Engineering/Solution Architect to produce artifacts and deliverables that adhere to best practices at Takeda.
•
Partner with data scientists to productionize models (including LLM/GenAI and RAG applications) and to translate research code into robust, maintainable services.
•
Contribute reusable templates, accelerators, and self-service tooling that raise the engineering bar across teams.
•
Promote MLOps best practices, mentor peers, and document standards.
Required Qualifications
•
Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience).
•
4+ years of hands-on experience in MLOps, ML engineering, data engineering, or DevOps for data/ML systems.
•
Strong Databricks experience: Delta Lake, MLflow, Unity Catalog, Jobs/Workflows, and Spark (PySpark).
•
Strong AWS experience across compute, storage, and IAM, plus at least one ML service (SageMaker and/or Bedrock).
•
Proficiency in Python for production code (packaging, testing, typing), plus solid SQL.
•
Experience building CI/CD pipelines and using Git-based workflows.
•
Working knowledge of containerization (Docker) and orchestration (Kubernetes/EKS or ECS).
•
Experience with Infrastructure as Code (Terraform, CloudFormation, or CDK).
•
Understanding of ML lifecycle concepts: experiment tracking, model registry, feature stores, and model monitoring/drift.
Preferred / Pharma-Specific Qualifications
•
Experience delivering software or ML in a GxP / regulated (FDA, EMA) life-sciences environment; familiarity with CSV/CSA, GAMP 5, 21 CFR Part 11, ALCOA+.
•
Exposure to pharma/biotech data domains: clinical trial data (CDISC/SDTM/ADaM), regulatory submissions, pharmacovigilance/safety, real-world data (RWD/RWE), or omics/biomarker datasets.
•
Experience operationalizing LLM/GenAI workloads (e.g., AWS Bedrock), including RAG, prompt/version management, evaluation, and guardrails.
•
Familiarity with handling PHI/PII under HIPAA/GDPR and with data-governance tooling.
•
Streaming/event-driven data (Kafka/Kinesis), data-observability tooling, and feature-store frameworks.
•
Relevant certifications: AWS (ML Specialty, Solutions Architect, or DevOps Engineer) and/or Databricks (Data Engineer, ML Engineer).
What Success Looks Like (First 12 Months)
•
Production ML/GenAI pipelines run reproducibly with full lineage, monitoring, and automated promotion across validated environments.
•
Deployment lead time and manual handoffs are measurably reduced through reusable templates and CI/CD.
•
Models in production are monitored for drift and quality, with documented retraining and rollback procedures.
•
Engineering artifacts meet inspection-readiness standards and pass internal QA review.
Locations
IND - Bengaluru
Worker Type
Employee
Worker Sub-Type
Regular
Time Type
Full time
More roles at Takeda
Plasma Center Licensed Practical Nurse (LPN)Plasma Center EMT-IAssociate Director, Scientific Communications Lead, Solid Tumors, Global Medical Affairs OncologyPhlebotomist (On the Job Training!)Production Engineer I
Seekless
Seek less: one search across companies' own career pages, instead of a dozen job boards.
Product
Find jobsCompaniesPricingFree Resume Checker
Company
AboutBlog
Legal
PrivacyTermsCookies
© 2026 SEEKLESS. ALL RIGHTS RESERVED.BUILT WITH CARE