NVIDIA · 9 hours ago
Senior Site Reliability Engineer - Storage
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. They are seeking a Senior Site Reliability Engineer focused on HPC storage to design, implement, and optimize on-prem storage solutions while collaborating with engineering teams.
AI InfrastructureArtificial Intelligence (AI)Consumer ElectronicsFoundational AIGPUHardwareSoftwareVirtual Reality
Responsibilities
Design, implement an on-prem HPC infrastructure supplemented with cloud computing to support the growing IT needs of NVIDIA
Design and implement advanced storage solutions, such as high-performance NFS, S3-compatible object storage, and distributed storage systems
Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources
Document the general procedures and practices, perform technology evaluations, related to distributed file systems
Collaborate across teams to better understand developers' workflows and gather their infrastructure requirements
Influence and guide methodologies for building, testing, and deploying applications to ensure optimal performance and resource utilization
Qualification
Required
BS in Computer Science (or equivalent experience) with 8+ years of relevant experience, MS with 5+ years of experience or Ph.D. with 3 years of experience
Deep experience with storage protocols such as nfs, NVMe/TCP, S3 and Lustre (LNet)
Experience with containerization technologies like Kubernetes and their integration with storage solutions
Proficiency in one or more programming languages (Python, GO) is a must
Experience working with monitoring and configuration management tools such as Chef, Ansible, Puppet, Saltstack, etc
Background with cloud infrastructure - AWS, Azure or Google Cloud
Experience with multiple monitoring stacks such as Prometheus+Grafana, Elasticsearch+Kibana
Excellent communication and collaboration skills
Preferred
Knowledge of HPC and AI solution technologies from CPU's and GPU's to high speed interconnects and supporting software
Experience with RDMA (InfiniBand or RoCE) fabrics
Background with HPC cluster management tools such as Slurm, PBS, LSF, etc
Passionate and experienced in AI methodologies
Benefits
Equity
Benefits
Company
NVIDIA
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.
H1B Sponsorship
NVIDIA has a track record of offering H1B sponsorships. Please note that this does not
guarantee sponsorship for this specific role. Below presents additional info for your
reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
Represents job field similar to this job
Trends of Total Sponsorships
2025 (1877)
2024 (1355)
2023 (976)
2022 (835)
2021 (601)
2020 (529)
Funding
Current Stage
Public CompanyTotal Funding
$4.09BKey Investors
ARPA-EARK Investment ManagementSoftBank Vision Fund
2023-05-09Grant· $5M
2022-08-09Post Ipo Equity· $65M
2021-02-18Post Ipo Equity
Recent News
GlobeNewswire
2026-01-13
2026-01-13
Company data provided by crunchbase