Own model deployment and serving infrastructure — manage containerised model endpoints, versioning, traffic management, and rollback mechanisms across environments
Implement monitoring and observability — establish model performance monitoring, data drift detection, and alerting frameworks that give the team early warning of degradation in production
Govern LLM usage across CUBE’s platform — take ownership of LLM provider relationships (OpenAI, Azure OpenAI, Anthropic), manage API access and versioning, track token consumption, and optimise cost across workloads
Manage the LLM gateway and prompt versioning — maintain tooling such as LangSmith or Helicone for tracing, evaluation, and prompt lifecycle management in production environments
Support experiment tracking and model registry — ensure that experiments are reproducible, models are catalogued, and the path from experiment to production is governed and auditable
Collaborate with data scientists and engineers — work closely with the Lead Data Scientist and Data Engineering team to understand modelling requirements and translate them into reliable production systems
Champion engineering best practices — bring CI/CD, infrastructure as code, and automated testing discipline to ML workflows; reduce toil and manual intervention wherever possible
Drive platform reliability and cost efficiency — monitor infrastructure spend, identify optimisation opportunities, and ensure platform SLAs are met