California law requires most employers to post a pay range.
At a glanceSummarised by Seekless from the posting.
Must have6
Significant experience in enterprise storage, distributed systems or cloud infrastructure, with technical leadership at Senior or Staff level
Deep understanding of file systems and storage technologies, including S3, POSIX, NFS and storage performance
Strong Linux systems knowledge, including kernel-level troubleshooting and debugging
Strong coding ability in Python or C++
Proven ability to diagnose complex issues using tools such as strace, tcpdump and perf
Genuine interest in working directly with customers and taking ownership of complex problems through to resolution
Nice to have5
Experience with DDN, VAST, Weka or similar scale-out storage/file systems
Familiarity with observability platforms such as Prometheus, Grafana, ELK or OpenTelemetry
Knowledge of replication, consistency models and data integrity mechanisms
Experience supporting AI/ML, LLM training or other high-performance computing environments
Experience using AI tools for log analysis, troubleshooting, automated RCA or reducing MTTR
Eligibility1
This position requires participation in an on-call rotation to provide after-hours support as needed
Skills
Python
C++
Linux
S3
POSIX
NFS
strace
tcpdump
perf
Prometheus
Grafana
ELK
OpenTelemetry
DDN
VAST
Weka
About the role
DDN is seeking a Sustaining Engineer to join our Infinia Core team. This is a senior, technical engineering role focused on owning complex customer escalations through to resolution.
You’ll own complex technical escalations end-to-end, from root-cause analysis and incident response through to mitigation, customer communication and product improvements. You’ll also help shape engineering best practice, mentor other engineers and drive the use of AI and automation to improve reliability and diagnostics.
If you enjoy solving difficult technical problems, working across the stack and seeing the direct impact of your engineering work on real customers, this is an opportunity to make a significant contribution.
About Infinia
Infinia is DDN’s next-generation, software-defined storage platform, built from the ground up for AI and accelerated computing. It combines separate control and data planes, all-flash performance, sub-millisecond latency and multi-tenancy for demanding enterprise and hyperscale AI and GPU workloads.
What You’ll Do
•
Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.
•
Own complex customer escalations from diagnosis through to resolution, mitigation and RCA.
•
Lead live incident response, war rooms and cross-functional investigations with Engineering, QA and Field teams.
•
Debug complex distributed-systems, storage and performance issues across the system, protocol and application layers.
•
Reproduce customer issues and feed findings into product and reliability improvements.
•
Develop runbooks, troubleshooting guidance and performance-tuning practices.
•
Act as a technical authority on Infinia internals, mentoring engineers and influencing architectural best practice.
•
Partner with Field CTOs, Solutions Architects and Sales Engineers on strategic customer issues.
•
Use AI, automation and observability to improve diagnostics, reliability and MTTR.
•
Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.
•
This position requires participation in an on-call rotation to provide after-hours support as needed.
What You’ll Bring
Must-Haves
•
Significant experience in enterprise storage, distributed systems or cloud infrastructure, with technical leadership at Senior or Staff level.
•
Deep understanding of file systems and storage technologies, including S3, POSIX, NFS and storage performance.
•
Strong Linux systems knowledge, including kernel-level troubleshooting and debugging.
•
Strong coding ability in Python or C++.
•
Proven ability to diagnose complex issues using tools such as strace, tcpdump and perf.
•
Genuine interest in working directly with customers and taking ownership of complex problems through to resolution.
Nice-to-Haves
•
Experience with DDN, VAST, Weka or similar scale-out storage/file systems.
•
Familiarity with observability platforms such as Prometheus, Grafana, ELK or OpenTelemetry.
•
Knowledge of replication, consistency models and data integrity mechanisms.
•
Experience supporting AI/ML, LLM training or other high-performance computing environments.
•
Experience using AI tools for log analysis, troubleshooting, automated RCA or reducing MTTR.