Operations: Expertise and leadership in verifying the quality of operations of Linux supercomputers, infrastructure systems, and other research computing systems. This includes, but is not limited to, performing daily system checks, analyzing system logs, troubleshooting hardware and software problems, monitoring and analyzing storage/infrastructure/job performance, helping users recognize job performance problems, writing scripts to enhance monitoring, and responding to unplanned system events such as power outages. This may require travel to various HPC sites to maintain physical installation of systems located off-site.
Proactively perform hardware maintenance on the clusters, cluster infrastructure, and other systems as needed. This includes diagnosing and fixing problems which includes, but is not limited to, running diagnostics, re-seating dimms, replacing hard disks, calling vendors for RMA support, replacing mother boards, and return shipping replacement parts.
Plan and perform software maintenance on both the clusters and the cluster infrastructure as needed. This includes, but is not limited to, installing operating systems, installing security patches, installing or upgrading drivers, upgrading firmware, installing or upgrading software licenses, installing or upgrading software specific to HPC cluster management. (50%)
Research:Investigates, architects and implements new technology as appropriate to add new features to both the user environment and to our deployment environment. This requires the ability to work without training to take a new technology through installation to production. This also includes the ability to develop and document procedures related to that technology and to train other members of the group. (25%)
Customer Support: Respond to tickets which include complaints, requests, troubleshooting, assessing storage options, etc. Provide training to groups or individuals as needed. (15%)
Other duties as assigned. (10%)