Rollout and deployment. Orchestrating how our agents and services roll out across many customer cloud environments: deployment strategies, per-customer configuration, automated health checks, and the monitoring that catches problems before customers do.
Inference and performance. Keeping model inference fast, reliable, and cost-effective at scale — serving infrastructure, GPU workloads, and the performance work that keeps agents responsive as volume grows.
Core infrastructure. Kubernetes, multi-account AWS, CI/CD, observability (traces, metrics, logs, alerting, SLOs), disaster recovery, and cost management.
Security posture. Access controls, secrets management, network security, image scanning, dependency auditing, and compliance work (SOC 2, enterprise security) as customer requirements demand.
Infrastructure as code. Defining, provisioning, and evolving all infrastructure through code — designing modules, managing state, and thinking hard about blast radius.