Domain Architecture & Technical Direction
● Own and evolve the technical direction for a defined HPC systems domain, such as Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments or service automation.
● Make architectural decisions that balance software quality, operational realities, customer needs, and long-term maintainability.
● Define how proven Slurm implementations should be packaged, automated and exposed as a service.
● Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operating models across teams.
● Act as the technical escalation point for the most complex issues within the domain.
Cross-Team Engineering Leverage
● Establish shared patterns for automation, service lifecycle management, observability, reliability and supportability across the HPC platform.
● Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems and platform tooling.
● Create reusable modules, automation, deployment patterns, and reference implementations that increase engineering leverage.
● Identity and correct avoidable technical divergence, duplicated effort and fragile operating models
● Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation and production operations.
Delivery, Reliability & Influence
● Lead technically critical initiatives spanning 2-4 teams or a defined HPC platform area.
● Unblocked delivery by clarifying technical direction and reducing ambiguity in complex system design problems.
● Contribute hands-on where needed to de-risk or accelerate critical work.
● Influence engineering teams without formal authority through strong judgement, design clarity and practical solutions.
● Partner with adjacent cloud-native software engineers so HPC implementations build on shared platform patterns rather than separate ones.
KPIs
● Technical direction across a defined HPC domain
● Delivery of critical initiatives across 2-4 teams
● Reduction in technical divergence and duplicated effort
● Reliability and supportability of Slurm-based HPC services