Die ganze Ausschreibung von Nebius B.V.
Das ist der Job
We offer competitive salaries ranging from $170k-$300k + equity based on your experience.
Darum lohnt es sich
We’re looking for a Senior Software Systems Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack.
You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments.
The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.
In this position, you will be responsible for: • Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. • Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. • Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. • Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. • Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation.
We expect you to have: • 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). • 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. • Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python).
It would be a plus if you have: • Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. • Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). • Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. • Background in Software-Defined Networking (SDN) and experience with HPC cluster networking. • Understanding of QEMU/KVM virtualization and managing virtualized environments. • Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems. • Familiarity with collective communication libraries like MPI and NCCL for distributed computing.
We conduct coding interviews as part of the process.
Bereit?
Bewerbung wird direkt an Nebius B.V. übergeben — kein Konto nötig.