TechNuggets Academy

NVIDIA InfiniBand Networking

Free NVIDIA-Certified Professional: AI Networking practice — 6 questions on NVIDIA InfiniBand Networking, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · NVIDIA InfiniBand Networking
A 2,048-node GPU training cluster uses NVIDIA Quantum-2 InfiniBand with UFM as the subnet manager. The primary UFM instance running on a dedicated server fails during a training job. Fabric connectivity should be preserved automatically. Which mechanism BEST explains how this is achieved without operator intervention?
InfiniBand supports master/standby SM redundancy; standby SMs poll the master at a configured interval, and on failure the standby with the best (numerically highest, per IB spec priority semantics as configured) priority takes over, performs subnet discovery/sweep, and reprograms LIDs and forwarding tables — all without host application intervention.
Question 2 of 6 · NVIDIA InfiniBand Networking
An architect is designing a non-blocking fat-tree (CBB — constant bisectional bandwidth) InfiniBand fabric for 1,024 GPU-attached HCAs using 40-port leaf switches with equal uplink and downlink port counts. What downlink-to-uplink port ratio on each leaf switch is required to guarantee full non-blocking bisectional bandwidth?
A true non-blocking (CBB) fat-tree requires the aggregate downlink bandwidth to hosts to equal the aggregate uplink bandwidth to the spine on every leaf switch — a 1:1 port ratio. Any tapering ratio (2:1, 3:1, etc.) introduces oversubscription, which the exam distinguishes from CBB design.
Question 3 of 6 · NVIDIA InfiniBand Networking
A cluster is running distributed training with NCCL all-reduce operations across 512 GPUs on Quantum-2 switches with SHARPv3 enabled. Which statement correctly describes what SHARP offloads to accomplish for this workload?
SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) offloads collective reduction operations (e.g., all-reduce) to the switch network by building an aggregation tree; switches perform partial reductions in-flight, cutting the data volume that must traverse the fabric and reducing endpoint compute/communication overhead — this is its defining exam-tested capability.
Question 4 of 6 · NVIDIA InfiniBand Networking
A fabric administrator observes that despite having multiple equal-cost paths between leaf switches (via LMC > 0), certain uplinks remain congested while others are underutilized, even though Adaptive Routing (AR) is enabled at the switch level. What is the MOST likely root cause?
Adaptive Routing requires that the SM's routing algorithm (e.g., fat-tree routing algorithm with LMC>0) actually populate multiple valid forwarding entries (AR groups) per destination at each switch. If the SM does not generate genuinely redundant paths in its LFT/AR group computation, AR has nothing to load-balance across even if enabled on switch ports.
Question 5 of 6 · NVIDIA InfiniBand Networking
On an NVIDIA Quantum InfiniBand fabric, an administrator wants to guarantee that management traffic (subnet management packets) is never starved by high-priority GPU training RDMA traffic, even under heavy congestion. Which InfiniBand QoS mechanism inherently reserves a dedicated lane for this purpose?
InfiniBand architecture reserves Virtual Lane 15 (VL15) exclusively for subnet management packets (SMPs). VL15 is never used for data traffic and is excluded from SL2VL mapping and standard flow control, guaranteeing management traffic always has a dedicated path regardless of data-plane congestion.
Question 6 of 6 · NVIDIA InfiniBand Networking
A hyperscale AI training deployment plans to scale beyond 4,000 GPU nodes and is evaluating Dragonfly+ versus a traditional 3-tier fat-tree topology for the InfiniBand fabric. Which statement correctly captures the primary trade-off the exam expects you to know?
Dragonfly+ topologies reduce cabling, switch tier count, and cost at very large scale by using fewer global (inter-group) links compared to a fully non-blocking multi-tier fat-tree, but this comes at the cost of needing robust adaptive routing to avoid congestion on shared global links, since the topology is not inherently uniformly non-blocking end-to-end like a well-provisioned fat-tree.
Ready for the real thing?

The full course: two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed explanations.

undefined $34.99 with code FREETEST34 — valid through Oct 11.

Get my $34.99 deal →