Simile’s Research organization builds the systems that every step of the model lifecycle runs on: data ingestion and schema design, distributed training, evaluation, serving, and monitoring. We are the reason a researcher’s hypothesis can become a production simulation in days rather than quarters.
Two things make this problem unusual. First, our research-to-product pipeline is unusually tight - the experimental methods we validate on Monday are integrated into systems customers use to make high-stakes decisions. Second, simulating a society means running inference over populations of agents, not single requests. A single customer study can mean millions of model calls with interdependent state. Cost per simulation and latency per agent are not back-office metrics for us; they determine what research is even possible to run.
As a Member of Technical Staff in ML Systems, you will build the platform our researchers train, evaluate, and deploy on - and own it through the last mile, where a trained checkpoint becomes a production service serving millions of interdependent agent calls at a cost per simulation we can afford.
This is a role for someone who is energized by both halves of that. You will spend some weeks designing the data schemas and training pipelines a research team depends on, others profiling a serving path to find where the FLOPs and GPU memory are going, and others still bringing up cluster nodes or deleting the third redundant copy of a code path. The common thread is leverage: every improvement you make compounds across every researcher and every simulation we run.
We are looking for engineers who find it gratifying to see their work pushed to its absolute limits, and who own problems end-to-end - including the last mile of deployment that most people would rather hand off.