TechNuggets Academy

AI/ML Workloads and Cluster Patterns

Free Cisco Certified Specialist - Data Center AI Infrastructure practice — 6 questions on AI/ML Workloads and Cluster Patterns, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · AI/ML Workloads and Cluster Patterns
A 512-GPU training cluster performs synchronous all-reduce for gradient synchronization at the end of every mini-batch. All GPUs finish computation within microseconds of each other and simultaneously push large gradient tensors onto the fabric. Which fabric characteristic is MOST critical to prevent this behavior from stalling the training job?
Synchronized all-reduce produces simultaneous many-to-many incast; a non-blocking lossless fabric with PFC/ECN absorbs this burst and prevents drops that would ripple through the collective and stall every GPU (straggler effect).
Question 2 of 6 · AI/ML Workloads and Cluster Patterns
Which network design is specifically recommended for connecting GPU server NICs in a large-scale AI training cluster to guarantee equal, predictable bandwidth and minimal hop count between any two GPUs regardless of their physical location in the fabric?
Rail-optimized leaf-spine designs align each GPU NIC 'rail' to dedicated spine switches, delivering consistent low-hop, non-blocking any-to-any bandwidth — the standard recommended pattern for AI training clusters.
Question 3 of 6 · AI/ML Workloads and Cluster Patterns
Which statement BEST captures the fundamental difference between training and inference workload traffic patterns on the data-center fabric?
This captures the recognized distinguishing pattern: training's heavy synchronized east-west collectives (all-reduce/all-to-all) versus inference's largely north-south request/response pattern, with the noted exception of multi-GPU model-parallel inference.
Question 4 of 6 · AI/ML Workloads and Cluster Patterns
A cluster uses GPUDirect RDMA over RoCEv2 for gradient synchronization across the leaf-spine fabric. Engineers observe collective operations occasionally stalling due to buffer overflow drops during synchronized bursts. Which configuration change directly addresses this at the network layer?
RoCEv2 requires a lossless network path. A PFC no-drop queue prevents buffer-overflow drops for that traffic class, while ECN provides end-to-end congestion signaling to throttle senders before drops occur — the standard lossless Ethernet configuration for RDMA-based AI fabrics.
Question 5 of 6 · AI/ML Workloads and Cluster Patterns
A large language model is split using tensor parallelism across 8 GPUs within a single node, while data parallelism is used across nodes for the remaining scale. Which of these two parallelism strategies places the strictest requirement on network latency and typically must be confined within a high-bandwidth, low-latency domain (e.g., NVLink or a single rail) rather than spanning the general fabric?
Tensor parallelism splits computation within a layer, requiring synchronization of partial results (all-reduce/all-gather) at every forward and backward step, making it extremely latency sensitive and generally confined to NVLink-connected GPUs within a node or rail rather than the broader Ethernet fabric.
Question 6 of 6 · AI/ML Workloads and Cluster Patterns
A training cluster checkpoints model state to a shared storage tier every 30 minutes. All 1,024 GPUs write checkpoint data simultaneously, causing a massive synchronized burst of traffic toward the storage fabric. What should the network design provision to prevent this periodic burst from degrading storage-tier performance and causing timeouts?
Periodic simultaneous checkpoint writes create a significant, predictable incast burst; the fabric toward storage must be sized and buffered for peak burst bandwidth, not just average throughput, to avoid drops and timeouts during checkpointing.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →