TechNuggets Academy
300-640 DCAI

Free Cisco Certified Specialist - Data Center AI Infrastructure Practice Test

12 exam-style questions with full explanations — no sign-up. Score yourself, then close your gaps with the full course.

Exam fee ~$3006 exam domainsLevel Professional2 timed practice tests in the course
✅ Free practice — no sign-up📝 Real exam-style questions💡 Detailed explanations💸 30-day money-back via Udemy
Question 1 of 12 · AI/ML Workloads and Cluster Patterns
A distributed training job splits a large model across 512 GPUs and uses all-reduce operations to synchronize gradients after every mini-batch. Which traffic pattern will dominate the fabric during this job?
All-reduce is a collective operation where every GPU exchanges gradient data with every other participating GPU, generating sustained, synchronized many-to-many east-west traffic across the fabric on every training step.
Question 2 of 12 · High-Performance AI Networking
An AI training fabric built on Cisco Nexus 9000 switches begins losing GPU-to-GPU RDMA throughput during large collective operations. Investigation shows that a single congested port is repeatedly sending PFC PAUSE frames upstream, and after several seconds an entire pod of switches stops forwarding RoCEv2 traffic on that priority, even though the original congestion point has since cleared. Which mechanism should be enabled to detect and automatically recover from this condition?
PFC Watchdog monitors a no-drop queue for prolonged PAUSE storms (PFC deadlock/HOL blocking) and automatically drops frames or shuts the interface to recover forwarding once a configured detect interval is exceeded, then restores traffic after the recovery interval.
Question 3 of 12 · AI Connectivity and Transport Models
A design team is building a distributed GPU training cluster and must select a fabric topology that guarantees full bisectional bandwidth between any two leaf switches, since GPU collective operations (all-reduce, all-to-all) generate heavy east-west traffic across the entire fabric. Which topology BEST satisfies this requirement?
AI training clusters require predictable, high-bandwidth any-to-any communication for collective operations. A non-blocking spine-leaf CLOS fabric with 1:1 oversubscription ensures every leaf-to-leaf path has equal, full bandwidth, which is the standard Cisco reference architecture for GPU fabrics.
Question 4 of 12 · Compute and Acceleration
A data center architect is designing an AI training cluster requiring 8 GPUs per node with a full non-blocking NVLink switch fabric for large language model training. Which Cisco compute platform should be selected?
The HGX 8-GPU baseboard integrated into the C880 M8 uses NVSwitch to provide a full non-blocking NVLink mesh among all 8 GPUs, enabling the all-to-all GPU communication bandwidth required for LLM training.
Question 5 of 12 · Storage and Data Pipelines
A GPU cluster training a large language model experiences significant GPU idle time whenever checkpoints are written every 30 minutes. Each 2TB checkpoint must be written in under 5 minutes to minimize idle GPU time. Which storage solution BEST meets this requirement?
Parallel file systems distribute I/O across many nodes and NVMe drives in parallel, and running over an RDMA fabric provides the low-latency, high-throughput write path needed to flush a 2TB checkpoint within minutes without stalling the training loop.
Question 6 of 12 · Operations and AIOps
A GPU training cluster using RoCEv2 over a Cisco Nexus 9000 fabric experiences intermittent job slowdowns. The network team suspects PFC pause storms are causing congestion spreading (head-of-line blocking) across the fabric. Which Nexus Dashboard tool should be used FIRST to confirm and pinpoint the congestion source?
NDI ingests hardware-level flow telemetry and PFC/ECN counters from the ASIC to detect microbursts, buffer congestion, and pause-frame storms, and can pinpoint the exact switch/port/queue causing head-of-line blocking.
Question 7 of 12 · AI/ML Workloads and Cluster Patterns
A company is designing a GPU cluster for large-scale distributed training and wants to guarantee any-to-any GPU communication with minimal oversubscription. Which fabric design pattern BEST meets this requirement?
AI training clusters require a non-blocking, rail-optimized spine-leaf design so that any GPU can communicate with any other GPU at full bandwidth without contention, matching the demands of collective operations like all-reduce and all-to-all.
Question 8 of 12 · High-Performance AI Networking
An engineer is tuning RoCEv2 congestion behavior on a Cisco Nexus 9000 leaf switch. The requirement is to have the switch begin marking packets with CE (Congestion Experienced) in the IP header once the no-drop queue depth exceeds a configured threshold, before PFC pause is triggered. Which configuration accomplishes this?
ECN marking with WRED-based min/max thresholds is configured within a network-qos policy-map; when average queue depth crosses the min threshold, ECN-capable packets are marked CE instead of being dropped, enabling DCQCN-based rate reduction before PFC engages.
Question 9 of 12 · AI Connectivity and Transport Models
An architect is choosing a transport protocol for GPU-to-GPU RDMA traffic across an Ethernet-based AI fabric, and must ensure lossless delivery without adopting a separate InfiniBand fabric. Which transport model should be selected?
RoCEv2 (RDMA over Converged Ethernet v2) is the standard transport for RDMA over Ethernet fabrics, and it requires a lossless underlay achieved through Priority Flow Control (PFC) for congestion isolation and ECN with DCQCN for end-to-end congestion management. This is the Cisco-recommended model for Ethernet-based AI clusters that avoid a separate InfiniBand fabric.
Question 10 of 12 · Compute and Acceleration
An AI cluster requires GPUDirect RDMA between GPUs and the network fabric to minimize latency during distributed training. Which requirement must be satisfied at the server design level?
For efficient GPUDirect RDMA peer-to-peer transfers, the GPU and its associated SmartNIC should reside on the same PCIe switch/root complex, avoiding cross-NUMA or cross-root-complex hops and keeping the data path off the CPU.
Question 11 of 12 · Storage and Data Pipelines
An architect needs to minimize CPU involvement and reduce latency when GPUs read training data directly from NVMe storage in an AI cluster. Which technology should be implemented?
GPUDirect Storage creates a direct data path between GPU memory and local or remote NVMe storage, bypassing the CPU and system memory bounce buffers, which reduces latency and CPU overhead for data-intensive AI workloads.
Question 12 of 12 · Operations and AIOps
Which single capability of Nexus Dashboard Insights specifically detects sub-microsecond buffer spikes that standard 30-second polling intervals would miss, which is critical for diagnosing transient congestion on lossless AI fabrics?
NDI's microburst detection samples ASIC-level buffer/queue statistics at very fine granularity (well below standard SNMP polling intervals), surfacing transient buffer spikes that cause packet drops or PFC pauses even when average utilization looks normal.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →

300-640 DCAI exam — quick answers

How much does the 300-640 DCAI exam cost?

The exam fee is approximately $300 and varies by region — confirm current pricing with the certification vendor before you book.

What topics are on the exam?

It covers 6 domains: AI/ML Workloads and Cluster Patterns (Content area (confirm official 300-640 blueprint weights on Cisco's site)), High-Performance AI Networking (Content area (confirm official 300-640 blueprint weights on Cisco's site)), AI Connectivity and Transport Models (Content area (confirm official 300-640 blueprint weights on Cisco's site)), Compute and Acceleration (Content area (confirm official 300-640 blueprint weights on Cisco's site)), Storage and Data Pipelines (Content area (confirm official 300-640 blueprint weights on Cisco's site)), Operations and AIOps (Content area (confirm official 300-640 blueprint weights on Cisco's site)). The full course has a dedicated chapter, lab and practice-test coverage for each.

Is this practice test really free?

Yes — all questions on this page are free with explanations and no sign-up. The paid Udemy course adds two full-length timed exams, video lessons and hands-on labs.

Will this prepare me for the real exam?

The questions mirror the real exam's style and are mapped to the official domains. This is exam-focused preparation — combine the free test with the full course's timed simulations to gauge your readiness.

More free practice by exam domain:
AI/ML Workloads and Cluster Patterns →High-Performance AI Networking →AI Connectivity and Transport Models →Compute and Acceleration →Storage and Data Pipelines →Operations and AIOps →