TechNuggets Academy

AI Connectivity and Transport Models

Free Cisco Certified Specialist - Data Center AI Infrastructure practice — 6 questions on AI Connectivity and Transport Models, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · AI Connectivity and Transport Models
A design team builds an AI training pod using a rail-optimized topology, where each GPU's NIC in every server connects to a leaf switch dedicated to that GPU's index (e.g., all GPU0 ports go to Leaf-1, all GPU1 ports go to Leaf-2). During a maintenance event, Leaf-3 (the GPU2 rail) goes offline. What is the MOST accurate description of the impact on the cluster?
In a rail-optimized topology, each rail (GPU index) is mapped to its own leaf switch across the entire pod, so a single rail switch failure isolates the impact to that specific GPU index on every server, not the whole fabric or a physical subset of servers.
Question 2 of 6 · AI Connectivity and Transport Models
An AI cluster running large all-reduce collective operations experiences persistent flow collisions and hotspots on specific ECMP links, even though overall fabric utilization is well below capacity. Which Cisco Nexus 9000 feature should be enabled to resolve this without redesigning the physical topology?
Nexus 9000 Dynamic Load Balancing (DLB) uses flowlet-based switching to dynamically spread elephant flows across ECMP paths based on real-time link utilization, directly addressing hash-based collisions common with large collective operations.
Question 3 of 6 · AI Connectivity and Transport Models
A Cisco Nexus 9000-based leaf-spine fabric for GPU training is initially provisioned with a 3:1 oversubscription ratio between leaf uplinks and downlinks, matching a typical enterprise data center design. Why is this configuration considered a design error for the AI cluster?
AI training workloads generate synchronized, bursty RDMA traffic during collective operations (all-reduce, all-gather); a 1:1 non-blocking fabric is standard guidance to avoid congestion and tail latency that directly slow down job completion time, unlike typical mixed enterprise traffic that tolerates oversubscription.
Question 4 of 6 · AI Connectivity and Transport Models
Within a single GPU server, engineers need the highest possible bandwidth and lowest latency for GPU-to-GPU communication during tensor-parallel operations, while the fabric connecting that server to other servers in the pod uses RoCEv2 over Ethernet. What best describes this architectural distinction?
NVLink/NVSwitch provides an ultra-high-bandwidth, low-latency scale-up interconnect for GPUs within a node or tightly coupled chassis, while RoCEv2 over Ethernet (or InfiniBand) is used for scale-out connectivity between nodes across the fabric — these are distinct transport domains with different bandwidth and topology characteristics.
Question 5 of 6 · AI Connectivity and Transport Models
A storage architecture team wants GPUs to read training data directly from NVMe-based storage arrays with minimal CPU involvement, using the same lossless Ethernet fabric already deployed for GPU-to-GPU RDMA traffic. Which transport combination BEST meets this requirement?
NVMe-oF over RoCEv2 combined with GPUDirect Storage allows data to move directly from NVMe storage to GPU memory over the same RDMA-capable lossless Ethernet fabric, bypassing the CPU and minimizing latency, which satisfies the stated requirement to reuse the existing RDMA fabric.
Question 6 of 6 · AI Connectivity and Transport Models
A shared AI training cluster must support multiple tenants with overlapping IP address space and strict Layer 2 segmentation between tenant workloads, while still preserving ECMP-based RDMA transport for RoCEv2 traffic within each tenant's segment. Which fabric design BEST satisfies both requirements?
BGP EVPN with a VXLAN overlay on top of a Layer 3 ECMP underlay provides tenant-level Layer 2 segmentation and overlapping IP support via VNIs, while the underlying ECMP fabric continues to support multipath RoCEv2 RDMA transport, satisfying both multi-tenancy and performance requirements simultaneously.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →