· Design, develop, and optimize HPC software running on large-scale Linux clusters, including distributed and parallel workloads (MPI, multithreading, GPU-accelerated pipelines, containerized workloads).
· Optimize application performance and power utilization across CPU, memory, storage, and network subsystem, with attention to throughput, latency, and scaling behavior.
· Develop and maintain system-level tooling for cluster bring-up, diagnostics, monitoring including component power usages, and health checks.
· Work closely with algorithms, systems and application teams to understand and translate workload characteristics into power-efficient HPC software solutions.
HPC Systems & Hardware Awareness
· Collaborate with hardware and systems teams to define HPC node, storage, and interconnect requirements based on software and algorithm needs.
· Understand and influence CPU/GPU selection, memory sizing, PCIe layout, NUMA behavior, and network topology to ensure optimal software performance.
· Participate in HW/SW co-debug activities, including performance bottlenecks, stability issues, and failure analysis.
Rack & Infrastructure Engineering
· Understand rack-level integration of HPC systems, focusing on power, cooling, cabling, networking, and physical layout considerations.
· Understand data-center and lab constraints such as power budgets, thermal limits, network drops, and serviceability.
· Contribute to best practices, and design reviews for new platforms and refresh cycles.
Cross-Functional Collaboration
· Act as a technical bridge between software, hardware, systems teams.
· Provide clear technical documentation covering software and system architecture, deployment flows, performance assumptions.