TechNuggets Academy

Spectrum-X Ethernet and RoCE for AI

Free NVIDIA-Certified Professional: AI Networking practice — 6 questions on Spectrum-X Ethernet and RoCE for AI, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Spectrum-X Ethernet and RoCE for AI
A customer deploys a 2,000-GPU cluster on Spectrum-X (Spectrum-4 switches + BlueField-3 SuperNICs) running large all-reduce collectives. They observe high tail latency during collective operations despite low average link utilization across the fabric. Which action BEST addresses this specific symptom?
High tail latency with low average utilization is the classic signature of microburst/incast congestion during collectives. Spectrum-X's value proposition is closed-loop, telemetry-based congestion control combined with fine-grained adaptive routing (packet spraying) that reacts to transient bursts far faster than average-utilization metrics suggest, smoothing tail latency.
Question 2 of 6 · Spectrum-X Ethernet and RoCE for AI
Which NVIDIA component is primarily responsible for performing the real-time, per-flow RTT telemetry measurement and closed-loop rate adjustment used by NVIDIA's programmable congestion control on a Spectrum-X fabric?
NVIDIA's telemetry-based programmable congestion control is executed at the NIC: the BlueField-3 SuperNIC measures per-flow round-trip telemetry and adjusts injection rates in a closed loop, working alongside switch-based INT data.
Question 3 of 6 · Spectrum-X Ethernet and RoCE for AI
An architect is tuning a Spectrum-X RoCEv2 fabric where RoCE traffic is mapped to priority 3. Which configuration pairing correctly prevents both packet loss and PFC deadlock/storm conditions?
Best practice is to set ECN thresholds below the PFC trigger point so ECN marks and throttles senders before pause frames are ever needed, while a bounded PFC watchdog timeout detects and clears stalled/stuck queues, protecting against deadlock and PFC storms.
Question 4 of 6 · Spectrum-X Ethernet and RoCE for AI
A team is deploying a 64-GPU training pod with predictable, static traffic patterns. Their top priority is the lowest possible per-hop switch latency, and they do not need multi-tenant Ethernet integration or convergence with existing datacenter Ethernet services. Which fabric choice BEST fits these requirements?
For a small, deterministic pod where absolute lowest latency matters and Ethernet multi-tenancy/convergence isn't required, NVIDIA Quantum InfiniBand's native RDMA transport and lower per-hop switch latency make it the better fit than an Ethernet-based fabric.
Question 5 of 6 · Spectrum-X Ethernet and RoCE for AI
In a RoCEv2 lossless Ethernet fabric, what is the fundamental functional difference between PFC and ECN?
PFC (IEEE 802.1Qbb) is a link-level, per-priority pause mechanism operating hop-by-hop to prevent buffer overflow and packet loss, while ECN (RFC 3168, used via DCQCN/PCC in RoCEv2) is an end-to-end congestion signal that marks packets in-flight so the sender can gracefully reduce its rate, avoiding traffic being halted outright.
Question 6 of 6 · Spectrum-X Ethernet and RoCE for AI
Which statement BEST describes how Spectrum-X's adaptive routing differs from traditional Ethernet ECMP hashing in AI training fabrics?
Unlike static per-flow ECMP hashing, which can create hotspots on large elephant flows, Spectrum-X spreads traffic at finer granularity across equal-cost paths and depends on the BlueField-3 NIC's RoCE transport to reassemble out-of-order packets, achieving much better load distribution for AI collective traffic.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →