Type: Full-time permanent contract or an Internship with a potential follow-up offer
Location: San Francisco or remote with future relocation to San Francisco (sponsored)
Start: ASAP
About us
•
We’re building state-of-the-art context compression. Our mission is to become the “Cloudflare for LLMs” — a compression layer embedded into most LLM pipelines by default.
•
We’re a team of ex-EPFL MSc/PhDs from dlab. We started by publishing papers, then got into YC and started making money helping companies cut their LLM costs.
•
We run the business like a research lab: form hypotheses, kill the ones that don’t work, double down on the ones that do.
About you:
•
A cracked full-stack engineer who enjoys a high-paced startup environment, takes pride in what they build and owns it end to end.
•
Solid understanding of cloud infrastructure, deployment, and production systems on AWS.
•
Python/basic ML Ops skills. Experience in scaling AI infra products is a plus.
•
Proactive, strong communicator with fast response time, team player
Tech Requirements
•
Strong backend engineering fundamentals
•
Experience with concurrency and distributed systems
•
Experience deploying and scaling production backend services on AWS
•
Ability to work across systems (Python + light frontend)