TechNuggets Academy

Compute and Acceleration

Free Cisco Certified Specialist - Data Center AI Infrastructure practice — 6 questions on Compute and Acceleration, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Compute and Acceleration
An architect is designing an 8-GPU AI training server that will use GPUDirect RDMA over RoCEv2 for east-west traffic. To achieve the lowest latency and avoid saturating the CPU's inter-socket interconnect (e.g., UPI/Infinity Fabric), which design requirement must be enforced?
GPUDirect RDMA performs best when the NIC and GPU sit on the same PCIe switch/root complex, allowing peer-to-peer DMA transfers that never cross the CPU socket interconnect. This is the basis of rail-optimized NIC-per-GPU design.
Question 2 of 6 · Compute and Acceleration
A design calls for a single Cisco UCS server supporting up to 8 SXM-based H100/H200 GPUs interconnected via a full NVSwitch mesh for large-scale LLM training. Which Cisco platform meets this requirement?
The UCS C885A M8 is Cisco's dense AI server purpose-built to house 8x SXM GPUs with NVSwitch/NVLink full-mesh interconnect, targeting large-scale training workloads.
Question 3 of 6 · Compute and Acceleration
An architect is validating the PCIe fabric for a GPU accelerator installed in a PCIe Gen5 x16 slot. What is the approximate total bidirectional bandwidth available to that single GPU over the PCIe link (excluding NVLink)?
PCIe Gen5 delivers approximately 3.94 GB/s per lane per direction; a x16 link yields roughly 63-64 GB/s per direction, or about 128 GB/s aggregate bidirectional bandwidth.
Question 4 of 6 · Compute and Acceleration
A cloud provider needs to run multiple independent inference tenants on a single physical GPU while guaranteeing hardware-level fault and memory isolation between tenants. Which capability should be configured on the GPU?
MIG partitions a supported GPU (A100/H100 class) into isolated hardware instances, each with dedicated compute cores, cache, and memory, providing true fault and security isolation between tenants.
Question 5 of 6 · Compute and Acceleration
Which statement correctly differentiates an NVSwitch-based GPU interconnect from a PCIe switch fabric within an AI training server?
NVSwitch fabrics (e.g., in H100 SXM systems) deliver several hundred GB/s of non-blocking, any-to-any GPU bandwidth without involving the CPU, whereas PCIe switch fabrics offer lower per-link bandwidth and may route traffic through the CPU root complex.
Question 6 of 6 · Compute and Acceleration
A cluster of AI training nodes each has 8 GPUs with dedicated RDMA NICs. To achieve non-blocking east-west all-reduce traffic across the fabric, each GPU's NIC on every node is connected to a leaf switch determined by GPU index (GPU0 to leaf1, GPU1 to leaf2, etc.) across all nodes. What is this design pattern called, and what problem does it solve?
Rail-optimized topology maps each GPU's NIC (by rank/index) to the same leaf switch across all nodes, creating dedicated non-blocking paths ('rails') for collective communication patterns like all-reduce, avoiding oversubscription and unpredictable ECMP hashing behavior.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →