As a Lead AI Engineer, you will:
• Design, develop, and maintain MLOps capabilities and advanced AI and machine learning systems that address specific business challenges.
• Implement models into production, building scalable training pipelines and deployment frameworks that handle large data volumes and high request rates.
• Administer and maintain Databricks workspaces, including provisioning and configuration, cluster and compute policies, job orchestration, runtime and library upgrades, catalog and access management, secrets, monitoring, and cost optimization.
• Deploy and maintain AI and machine learning infrastructure through infrastructure as code and automated release pipelines, keeping environments repeatable, secure, auditable, and consistent.
• Build and optimize data ingestion, preprocessing, and feature engineering workflows that support model training and inference.
• Automate model training, testing, deployment, and update workflows following CI/CD best practices.
• Build and maintain domain-specific feature and model monitoring, tracking performance metrics and drift and updating models to sustain high-quality outputs.
• Implement onboarding and operating standards for platform users, including naming and packaging conventions, validation rules, and exception handling.
• Ensure the operational stability and scalability of AI systems, adhering to ethical guidelines and contributing to the organization’s AI infrastructure.
• Influence stakeholders and partner with data science, platform, and product teams to translate requirements into technical solutions.
• Guide and mentor junior engineers through on-the-job experiences and code and design reviews, fostering continuous improvement across the discipline.
Required Qualifications
• Master’s degree with 3+ years of relevant experience, or Bachelor’s degree with 5+ years, in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Engineering, or a related field; equivalent practical experience considered.
• Hands-on MLOps experience across model monitoring, feature catalogs, experiment tracking, model registry, and lifecycle CI/CD pipelines.
• Hands-on experience administering Databricks workspaces, including cluster and compute policies, job orchestration, runtime and library upgrades, permissions, and secrets.
• Experience deploying and maintaining infrastructure through infrastructure as code and automated pipelines, including environment provisioning, configuration management, and controlled release and rollback.
• Strong experience with Spark and distributed data processing.
• Proficiency in Python, PySpark, and SQL.
• Hands-on experience with CI/CD and build tooling such as Git, Jenkins, Maven, and Artifactory.
• Experience building and optimizing feature engineering and large-scale data processing workflows.
• Strong understanding of machine learning and deep learning techniques, model lifecycle management, and production AI systems.
• Experience with model deployment, evaluation, observability, optimization, and operational support.
• Experience with cloud operations across public and private cloud environments.
• Ability to communicate technical concepts clearly, work independently, and mentor other engineers.