The AI Insights team builds customer-facing AI capabilities across CoreWeave’s Mission Control portfolio. We combine machine learning, observability, and production software engineering to help engineers and customers understand workload health, diagnose infrastructure issues, and identify opportunities to improve efficiency, capacity, and performance.
Our work spans telemetry from metrics, logs, traces, alerts, and operational events. We are building the intelligence layer that turns this data into trustworthy, actionable insights—grounded in evidence and integrated into the tools where people operate CoreWeave infrastructure.
This is not a role focused on building a generic chatbot. You will help build the underlying ML systems, services, evaluation frameworks, and product capabilities that make AI-powered troubleshooting and optimization reliable in real-world environments.
As a Staff Machine Learning Engineer, you will be a technical leader on the AI Insights team. You will define and implement the machine learning systems that power anomaly detection, signal correlation, incident understanding, recommendations, and other intelligence capabilities across CoreWeave’s observability and cloud platforms.
You will work across the full lifecycle: framing problems, developing models and algorithms, designing evaluation methodology, building data and inference services, integrating with production systems, and operating the result at scale. You will partner closely with software engineers, product managers, researchers, and domain experts to turn promising ideas into reliable customer experiences.