TechNuggets Academy

Troubleshooting Tools

Free NVIDIA-Certified Professional: AI Networking practice — 6 questions on Troubleshooting Tools, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Troubleshooting Tools
A 400G InfiniBand port shows no link-down or link-error-recovery events, yet application throughput on that path has quietly degraded over several days. Running mlxlink on the port shows FEC-corrected block counts climbing steadily and effective/raw BER trending upward, while uncorrected block count is still zero. What should the engineer conclude and do?
Rising raw/effective BER with increasing FEC-corrected blocks is the classic mlxlink early-warning signature of a marginal optic, connector, or cable. FEC is masking the problem today, but a trending BER means uncorrectable errors and link flaps are imminent — this is exactly the proactive maintenance signal mlxlink is designed to expose before an outage occurs.
Question 2 of 6 · Troubleshooting Tools
A network team needs a real-time, fabric-wide congestion and hotspot heatmap spanning several hundred Quantum-2 switches, correlating port utilization with buffer occupancy across the entire topology, without manually querying each device. Which approach best satisfies this requirement?
UFM (Unified Fabric Manager) is purpose-built for centralized, real-time telemetry collection across the entire IB/Ethernet fabric, including congestion and buffer-occupancy dashboards, and is the exam's expected answer for any fabric-scale, holistic monitoring requirement.
Question 3 of 6 · Troubleshooting Tools
A RoCEv2 leaf-spine fabric experiences sudden severe throughput collapse with symptoms of head-of-line blocking spreading across multiple unrelated flows. The team needs to confirm whether this is being driven by a PFC pause storm rather than ECN-based congestion signaling. Which counter is the correct one to examine to confirm a PFC storm?
A PFC storm is directly confirmed by per-priority pause frame counters rising rapidly and remaining elevated on multiple ports simultaneously — this is the specific, unambiguous signature of PAUSE-driven head-of-line blocking spreading through the fabric, distinct from ECN/DCQCN-based congestion response.
Question 4 of 6 · Troubleshooting Tools
On an InfiniBand switch port, the symbol_error_counter is steadily incrementing while the link_error_recovery_counter remains at zero and the port has never transitioned to Down. What does this combination of counter behavior indicate?
symbol_error_counter increments on physical-layer bit/symbol errors detected on the wire. link_error_recovery_counter only increments when those errors accumulate enough to force a retraining event. A rising symbol error count with zero link error recoveries means low-level physical errors are occurring and being tolerated/corrected without yet requiring the link to retrain — an early warning worth investigating (cable, connector, or transceiver quality) before it escalates.
Question 5 of 6 · Troubleshooting Tools
During a large all-to-all collective operation on a Spectrum-X RoCEv2 fabric, GPU jobs report intermittent long-tail latency spikes. Telemetry shows CNP (Congestion Notification Packet) counters rising sharply on several spine-facing ports, while PFC pause frame counters on those same ports remain near zero. What is the correct interpretation of this telemetry pattern?
Rising CNP counts with near-zero PFC pause activity is the textbook signature of DCQCN working correctly: ECN marks packets under incipient congestion, receivers generate CNPs, and senders reduce rate — successfully preventing the need for PFC pause. During legitimate incast from an all-to-all collective, this pattern reflects expected, healthy congestion control response rather than a fault to be remediated.
Question 6 of 6 · Troubleshooting Tools
NCCL all-reduce throughput across a multi-rail Quantum-2 InfiniBand fabric is well below expected bandwidth. UFM telemetry shows one spine switch's uplinks running near 100% utilization while all other spine switches average around 40%. What should the engineer verify next to isolate the root cause?
A single hotspot spine switch while others run at ~40% utilization is the classic signature of static or ECMP-only routing failing to load-balance traffic, or adaptive routing being disabled/misconfigured. The correct next diagnostic step is to verify AR is enabled and actively rebalancing flows via the subnet manager before considering hardware replacement.
Ready for the real thing?

The full course: two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed explanations.

undefined $34.99 with code FREETEST34 — valid through Oct 11.

Get my $34.99 deal →