Senior Solutions Architect, Cloud Infrastructure and DevOps - NVIS

NVIDIA

2.7

(9)

Multiple Locations (Remote)

#JR1988918

Position summary

ement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!

What you'll be doing:

  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting Manage Linux job/workload schedulers and orchestration tools.

  • Develop and maintain continuous integration and delivery pipelines .

  • Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources.

  • Deploy monitoring solutions for the servers, network and storage.

  • Perform troubleshooting bottom up from bare metal, operating system, software stack and application level.

  • Being a technical resource, develop, re-define and document standard methodologies to share with internal teams Support Research & Development activities and engage in POCs/POVs for future improvements .

What we need to see:

  • BS/MS/PhD or equivalent experience in Computer Science, Data Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering fields with at least 8 years work or research experience in networking fundamentals, TCP/IP stack, and data center architecture.

  • Knowledge of HPC and AI solution technologies from CPU's and GPU's to high speed interconnects and supporting software.

  • Direct design, implementation and management experience with cloud computing platforms (e.g. AWS, Azure, Google Cloud).

  • Experience with job scheduling workloads and orchestration technologies such as Slurm, Kubernetes and Singularity.

  • Excellent knowledge of Windows and Linux (Redhat/CentOS and Ubuntu) networking (sockets, firewalld, iptables, wireshark, etc.) and internals, ACLs and OS level security protection and common protocols e.g. TCP, DHCP, DNS, etc.

  • Experience with multiple storage solutions such as Lustre, GPFS, zfs and xfs. Familiarity with newer and emerging storage technologies.

  • Python programming and bash scripting experience.

  • Comfortable with automation and configuration management tools including Jenkins, Ansible, Puppet/Chef, etc.

  • Deep knowledge of Networking Protocols like InfiniBand, Ethernet Deep understanding and experience with virtual systems (for example VMware, Hyper-V, KVM, or Citrix).

  • Strong written, verbal, and listening skills in English are critical.

Ways to stand out from the crowd:

  • Knowledge of CPU and/or GPU architecture .

  • Knowledge of Kubernetes, container related microservice technologies.

  • Experience with GPU-focused hardware/software (DGX, CUDA.)

  • Background with RDMA (InfiniBand or RoCE) fabrics.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking individuals in the world working for us. If you're creative and autonomous, we want to hear from you.