✅ Free practice — no sign-up📝 Real exam-style questions💡 Detailed explanations💸 30-day money-back via Udemy
Question 1 of 12 · System and Server Bring-up
A newly racked DGX H100 system passes POST and boots into its OS, but before joining the cluster the bring-up engineer must confirm that all NVLink connections between the eight GPUs are active. Which command should be used to validate NVLink link status?
nvidia-smi nvlink -s reports the status (active/inactive) and speed of every NVLink connection on each GPU, which is the standard bring-up check for NVLink health.
Question 2 of 12 · Server and Network Installation and Configuration
A customer is deploying a DGX H100 SuperPOD using a rail-optimized InfiniBand network design. In this topology, how should the ConnectX-7 adapter connected to GPU 0 on every node in the pod be cabled?
Rail-optimized design dedicates one leaf switch per GPU index (rail) so that all-reduce/collective traffic for a given GPU rank across nodes stays on a single switch, minimizing hops and maximizing bisection bandwidth.
Question 3 of 12 · Physical Layer Management
A technician is cabling a rail-optimized DGX H100 SuperPOD fabric using NDR InfiniBand (400Gb/s per port). The distance between the compute rack and the NDR leaf switch rack is 50 meters. Which cable/transceiver type should be used for this link?
At NDR (400Gb/s) signal rates, passive copper DAC is only reliable for very short runs (roughly 1.5-2m) due to signal attenuation. Runs of 50m require active optical cables or pluggable transceivers over fiber to maintain signal integrity.
Question 4 of 12 · Cluster Management and Orchestration
A cloud provider runs a shared DGX-based Kubernetes cluster serving multiple untrusted tenants. Each tenant must receive a hardware-isolated GPU partition with dedicated memory and fault isolation, and no tenant may observe another's workload activity. Which GPU Operator feature BEST satisfies these requirements?
MIG creates hardware-partitioned instances with separate memory, cache, and compute slices, and GPU Operator exposes each MIG instance as a distinct nvidia.com/gpu resource, giving tenants real isolation.
Question 5 of 12 · Troubleshooting and Optimization
A DGX A100 node's NCCL-based multi-GPU training job shows a 40% throughput drop compared to baseline. GPU SM utilization looks normal, and nvidia-smi shows no thermal or power throttling. Which command should the administrator run FIRST to check for a degraded NVLink connection between GPU pairs?
nvidia-smi nvlink -s (and nvlink -e for error counts) directly reports the state and error counters of each NVLink connection per GPU, which is the fastest way to isolate a degraded intra-node link when compute utilization is normal but throughput is down.
Question 6 of 12 · System and Server Bring-up
An administrator must update the BIOS, BMC, and GPU VBIOS firmware on a DGX H100 as part of initial bring-up, before OS re-installation, using NVIDIA's officially supported firmware update mechanism for DGX and NVIDIA-Certified Systems. Which tool should be used?
nvfwupd is NVIDIA's supported firmware update utility for DGX and NVIDIA-Certified Systems, used to update BIOS, BMC, GPU VBIOS, and other system firmware components as part of bring-up.
Question 7 of 12 · Server and Network Installation and Configuration
A customer's data center standardized on Ethernet switches and wants to build a lossless RDMA fabric for GPU-to-GPU communication without deploying a separate InfiniBand network. Which technology should be configured to meet this requirement?
RoCEv2 (RDMA over Converged Ethernet v2) enables RDMA over standard Ethernet infrastructure but requires a lossless fabric, achieved by configuring Priority Flow Control (802.1Qbb) and ECN/DCQCN congestion control end-to-end on switches and NICs.
Question 8 of 12 · Physical Layer Management
During bring-up validation of a new NDR InfiniBand leaf-spine fabric, a port that should negotiate at 400Gb/s instead links up at 200Gb/s (HDR speed), and `mlxlink` output shows elevated raw bit error rate (BER) on that port. What is the MOST likely physical layer cause?
Elevated raw BER combined with a link auto-downtraining to a lower speed is a classic physical layer symptom of a dirty/damaged optical connector, degraded fiber, or a cable exceeding its rated length/signal budget — the link firmware falls back to a lower, more reliable speed.
Question 9 of 12 · Cluster Management and Orchestration
Which NVIDIA GPU Operator component continuously exports per-GPU utilization, memory, temperature, and ECC error metrics for consumption by Prometheus in a BCM-managed Kubernetes cluster?
DCGM Exporter, deployed as part of the GPU Operator, wraps NVIDIA Data Center GPU Manager to expose Prometheus-format metrics including utilization, memory, temperature, power, and ECC/Xid errors per GPU.
Question 10 of 12 · Troubleshooting and Optimization
An administrator managing a Base Command Manager-controlled GPU cluster wants a single tool that can run multi-level GPU health diagnostics — including memory bandwidth tests, NVLink validation, and extended stress tests — and can be scheduled cluster-wide for proactive health monitoring. Which tool should be used?
DCGM's dcgmi diag command offers multiple diagnostic run levels (r1 through r4/r3) covering deployment checks, integration tests, medium stress tests, and extended stress/NVLink validation, and it integrates natively with Base Command Manager for cluster-wide scheduled health checks.
Question 11 of 12 · System and Server Bring-up
On a DGX H100 system with NVSwitch, GPU-to-GPU communication over NVLink will not initialize correctly if a particular system service fails to start. Which service must be running to properly bring up the NVSwitch fabric before workloads are launched?
nv-fabricmanager configures the NVSwitch interconnect fabric and must be running and healthy for GPUs to communicate over NVLink on multi-GPU NVSwitch-based systems like the DGX H100.
Question 12 of 12 · Server and Network Installation and Configuration
An enterprise wants strict separation of duties so that host server administrators cannot alter the network, security, or storage virtualization functions running on a BlueField-3 DPU. Which BlueField deployment mode should be configured?
Zero Trust mode isolates the DPU's Arm subsystem and control-plane functions from the host, preventing host administrators from managing or tampering with the DPU's security, network, and storage services; management is restricted to a separate trusted entity.
Ready for the real thing?
The full course: two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed explanations.
$199.99$34.99 with code FREETEST33 — valid through September 9.
The exam fee is approximately $400 and varies by region — confirm current pricing with the certification vendor before you book.
What topics are on the exam?
It covers 5 domains: System and Server Bring-up (~22%), Server and Network Installation and Configuration (~22%), Physical Layer Management (~18%), Cluster Management and Orchestration (~20%), Troubleshooting and Optimization (~18%). The full course has a dedicated chapter, lab and practice-test coverage for each.
Is this practice test really free?
Yes — all questions on this page are free with explanations and no sign-up. The paid Udemy course adds two full-length timed exams, video lessons and hands-on labs.
How do I get the discount?
Use code FREETEST33 at checkout for $34.99 (list $199.99) through September 9 — the enroll button applies it automatically.
Will this prepare me for the real exam?
The questions mirror the real exam's style and are mapped to the official domains. This is exam-focused preparation — combine the free test with the full course's timed simulations to gauge your readiness.