NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can tackle, and that matter to the world. This is our life’s work, to amplify human imagination and intelligence. Make the choice, join our diverse team today!
We are looking for an outstanding architect for a Senior System Engineer role for system bringup and datacenter applications. Be a key player to the most exciting computing hardware and software to contribute to the latest breakthroughs in artificial intelligence, Omniverse, and GPU computing. Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs. You will work with the latest Accelerated computing and Deep Learning software, or Omniverse software, and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.
What you’ll be doing:
Provide engineering solutions to enable large scale scheduling and resource management for performance for GPU Computing products and software stacks, ensure technical relationships with internal and external engineering teams, and assisting systems, machine learning/deep learning engineers in building creative solutions based on NVIDIA technology.
Be an internal reference for scheduling, IO and other datacenter and large-scale GPU-accelerated system solutions among the NVIDIA technical community.
What we need to see:
6+ years of experience using in accelerated computing for datacenter/HPC solutions.
Proven years of OS and server level automation, CI/CD process and DevOps experience using Python, SHELL, Ansible, Jenkins.
Strong server and Linux(Ubuntu, RedHat, CentOS, SuSE, Fedora and etc…) troubleshooting and debugging experience in a bare-metal/KVM/K8S environment.
Experience working data center infrastructure for bare metal provisioning, testing and bringup.
Knowledge of SRE principles (observability, SLOs, logging, etc.).
Strong experience in FW, BMC/OpenBMC, Network protocol, internal/external enterprise storage devices, PCIe buses and devices, IO sub-devices, CPU and memory, ACPI, UEFI, redfish.
Strong verbal and written communication skills.
Ability to multitask effectively in a dynamic environment and action driven with strong analytical and analytical skills.
BS (or equivalent experience) in Engineering, Mathematics, Physics, or Computer Science. MS or PhD desirable.
Ways to stand out from the crowd:
Background with Host management systems (DHCP, Redfish, UEFI) and host security services such as TPM, TXT, and SecureBoot.
Exposure to telemetry catalog and observability stack.
Exposure to container technology and software defined network.
Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/ We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.