Free NVIDIA-Certified Professional: AI Infrastructure practice — 6 questions on System and Server Bring-up, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · System and Server Bring-up
A company just racked a new DGX H100 node. After OS and driver installation, engineers must verify all 8 GPUs have full NVLink connectivity through the 4 onboard NVSwitches before moving the node into production. Which command should they run to confirm the NVLink topology matches the expected DGX H100 all-to-all connectivity?
nvidia-smi topo -m prints the GPU-to-GPU interconnect matrix, showing NVx entries for NVLink-connected pairs, which lets engineers confirm full NVSwitch fabric connectivity across all 8 GPUs.
Question 2 of 6 · System and Server Bring-up
After initial bring-up of an NVIDIA-Certified server with 8x H100 GPUs, the deployment team wants to run an extended hardware diagnostic that stresses the GPUs and validates NVLink bandwidth before handing the system to production workloads. Which command correctly performs this validation?
dcgmi diag -r 3 runs DCGM's 'long' diagnostic level, which includes memory, PCIe, and NVLink bandwidth stress tests appropriate for thorough bring-up validation before production use.
Question 3 of 6 · System and Server Bring-up
During BIOS validation on a new NVIDIA-Certified server intended for multi-GPU training with GPUDirect RDMA over InfiniBand, engineers observe that GPU peer-to-peer communication and RDMA throughput are far below expected values. Which BIOS/PCIe setting should be checked and corrected first?
ACS enforces isolation between PCIe devices, forcing peer-to-peer GPU traffic and GPUDirect RDMA traffic through the root complex/CPU instead of a direct path. NVIDIA reference bring-up procedures require ACS to be disabled on the root ports serving GPUs and RDMA NICs to restore expected throughput.
Question 4 of 6 · System and Server Bring-up
While bringing up a new 4-node DGX H100 SuperPOD segment, NCCL all-reduce tests intermittently fail with topology errors on one specific node while the other three nodes complete successfully. All four nodes were imaged from the same OS and driver template. What should be investigated FIRST as the most likely root cause?
Firmware inconsistencies such as mismatched GPU VBIOS or NVSwitch firmware between otherwise identically imaged nodes are a well-known bring-up defect that causes NVLink/NVSwitch topology mismatches and intermittent NCCL failures. Firmware versions should be audited and aligned across the cluster during bring-up before deeper network debugging.
Question 5 of 6 · System and Server Bring-up
During bring-up documentation review for a DGX H100 system, a new engineer asks how NVLink and NVSwitch relate to each other. Which statement is correct?
NVLink is the high-bandwidth GPU-to-GPU interconnect link/protocol, while NVSwitch is the on-board switching ASIC (four per DGX H100 baseboard) that interconnects multiple NVLinks so every GPU can communicate with every other GPU at full bandwidth in an all-to-all fabric.
Question 6 of 6 · System and Server Bring-up
A newly racked DGX H100 fails to complete a normal power-on sequence: the chassis fans spin up, but the system never POSTs and no output appears on console redirection. Following standard bring-up troubleshooting procedure, what should be checked FIRST via the BMC before escalating to hardware replacement?
Because the host never completes POST, host OS-based tools cannot run and reflashing firmware is premature. The correct first step is out-of-band troubleshooting through the BMC: review the System Event Log and sensor data for power-supply, voltage rail, or CPU/GPU fault flags to identify the hardware condition preventing POST.
Ready for the real thing?
The full course: two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed explanations.
$199.99$34.99 with code FREETEST33 — valid through September 9.