TechNuggets Academy

High-Performance AI Networking

Free Cisco Certified Specialist - Data Center AI Infrastructure practice — 6 questions on High-Performance AI Networking, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · High-Performance AI Networking
A four-tier leaf-spine AI fabric using RoCEv2 experiences intermittent complete traffic stalls on a subset of no-drop queues, with switches showing sustained PFC pause frames on the same priority in a circular pattern across four leaf switches. No single link is oversubscribed and ECN is enabled end-to-end. What is the MOST likely root cause?
Circular pause dependencies across multiple switches on the same lossless priority, with no single oversubscribed link, is the classic signature of a PFC deadlock, where each switch is waiting on buffer space held up by pause frames from the next hop in a loop.
Question 2 of 6 · High-Performance AI Networking
An engineer is tuning WRED-based ECN marking on Nexus switches carrying RoCEv2 traffic in the no-drop queue. Current thresholds are min-threshold 150KB and max-threshold 1500KB with a 100% mark-probability. Congestion Notification Packets (CNPs) are arriving so late that GPUs still experience PFC pauses before senders throttle. What change BEST addresses this?
Lowering the ECN min/max thresholds causes marking to begin earlier in the queue buildup, so CNPs are generated and returned to senders before the queue depth reaches the level that triggers PFC pause, restoring the intended layered congestion response (ECN first, PFC as last resort).
Question 3 of 6 · High-Performance AI Networking
A data center architect is calculating required PFC headroom buffer on a leaf switch to prevent no-drop queue drops before a pause takes effect on 100G links with 100m of cabling between leaf and GPU NIC. Which factor is NOT directly part of this headroom calculation?
PFC headroom buffer sizing depends on link speed, cable length/propagation delay, and switch reaction time -- collectively determining how much data can arrive before the pause takes effect. ECMP path count affects load balancing and traffic distribution but is not a factor in the per-link headroom buffer formula.
Question 4 of 6 · High-Performance AI Networking
Multiple GPU servers behind different leaf switches are performing an AllReduce operation, causing many-to-one incast traffic toward a single leaf-facing port. Despite ECMP being configured across four equal-cost spine uplinks, one uplink consistently carries far more RoCEv2 flow traffic than the others, worsening congestion. What is the BEST explanation and fix?
RoCEv2 encapsulates RDMA traffic over UDP with a fixed destination port (4791) but a varying, often entropy-rich source UDP port per queue pair; if ECMP hashing ignores the UDP source port and hashes only on IP fields, many flows collapse onto the same path (polarization). Enabling hash inputs that include the UDP source port restores even distribution across ECMP links.
Question 5 of 6 · High-Performance AI Networking
An architect must ensure that a burst of malfunctioning traffic from a single misbehaving host does not indefinitely pause an entire no-drop queue class fabric-wide, which would stall unrelated healthy flows sharing that priority. Which mechanism is specifically designed to detect and recover from this condition?
PFC watchdog specifically monitors for a priority queue that remains paused beyond a configured threshold (indicating a stuck or malicious sender) and can take corrective action such as temporarily disabling that priority's pause behavior on the port, preventing a fabric-wide deadlock/storm from a single bad actor.
Question 6 of 6 · High-Performance AI Networking
During RoCEv2 fabric design review, an engineer states that RoCEv2 packets can be routed across Layer 3 boundaries in a spine-leaf fabric because RoCEv2 encapsulates the InfiniBand transport header directly inside an IP/UDP packet. Which detail of this encapsulation is MOST important for correctly configuring QoS classification on the switches?
RoCEv2 uses UDP encapsulation with a well-known destination port of 4791, which switches use in QoS classification ACLs/policies to identify RoCEv2 traffic and map it into the correct no-drop CoS/priority for PFC and ECN treatment.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →