· 6+ years in data science, applied ML, or AI engineering, including 2+ years building LLM-powered products. Healthcare experience is a plus.
· Deep NLP and GenAI experience. Statistical and classical machine learning is good to have on top of that.
· Strong hands-on Python — building highly scalable, performant enterprise applications, plus optimization technique.
· Hands-on experience with deep learning frameworks: PyTorch and/or HuggingFace transformers.
· At least one shipped GenAI product with a genuinely complex architecture — multiple agents, memory, retrieval, and agent OTEL/tracing in production.
· Working command of modern fine-tuning: PEFT methods, with LoRA and QLoRA preferred.
· Hands-on experience with at least one ML platform — Databricks, Azure ML, or SageMaker.
· Experience leading engineers, formally or as a technical lead — you have owned other people’s output, not only your own.
· Strong written and spoken communication, with a customer-focused instinct in both conversation and documentation.
· Preferably a Master’s in Computer Science, Computer Engineering, or a related field.
The engineering baseline we assume
Everything above sits on top of independent delivery. This role assumes you can already do the following without supervision:
· Build production-grade RAG and LLM/SLM-powered features end to end with limited supervision.
· Work fluently in at least one orchestration framework — LangGraph, LlamaIndex, CrewAI, or equivalent — to compose multi-step, tool-using flows.
· Design retrieval pipelines and tune them for relevance: chunking, embeddings, vector stores, re-ranking.
· Implement prompt engineering, function and tool calling, and reliable structured-output parsing.
· Write and run evals — golden sets, LLM-as-judge — to measure quality and catch regressions before customers do.
· Containerize and deploy services (Docker, REST/gRPC) with an eye on latency, token cost, and basic guardrails.
· Document well and participate actively in code review.