TechNuggets Academy

Storage and Data Pipelines

Free Cisco Certified Specialist - Data Center AI Infrastructure practice — 6 questions on Storage and Data Pipelines, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Storage and Data Pipelines
An 8-node GPU training cluster must checkpoint model state every 30 minutes. Each node writes an 80GB checkpoint shard, and training must resume within 5 minutes of checkpoint completion to avoid GPU idle time, meaning all shards must land on shared storage within that window. Which storage design BEST satisfies this requirement?
640GB of aggregate checkpoint data in 300 seconds requires ~2.1GB/s sustained, and a parallel file system with multiple I/O nodes over RDMA-capable 100GbE links provides ample aggregate bandwidth with low latency and no single bottleneck.
Question 2 of 6 · Storage and Data Pipelines
An architect wants GPU memory to perform direct DMA transfers to and from storage using GPUDirect Storage (GDS), bypassing host CPU memory staging entirely. Which storage access protocol MUST be deployed to support this capability?
GPUDirect Storage requires an RDMA-capable transport so the GPU can DMA directly to/from the storage target; NVMe-oF over RoCEv2 provides the RDMA fabric needed for this direct path.
Question 3 of 6 · Storage and Data Pipelines
A fabric carries both RoCEv2 GPU collective traffic and NVMe-oF storage traffic on the same switches. The design must prevent either lossless flow from causing packet loss or blocking the other. Which PFC/QoS configuration should be applied?
Separating the two lossless flows into distinct no-drop classes with independently tuned PFC/ECN thresholds prevents head-of-line blocking between GPU and storage traffic while preserving losslessness for both.
Question 4 of 6 · Storage and Data Pipelines
A training pipeline for 64 GPU nodes performs shuffled batch loading of millions of small image files averaging 150KB each, requiring very high random-access IOPS across the cluster. Which storage architecture BEST meets this requirement?
A parallel file system with distributed metadata and RDMA client access is purpose-built to sustain high random small-file IOPS across many concurrent clients without a single metadata or throughput bottleneck.
Question 5 of 6 · Storage and Data Pipelines
In a Cisco AI data center reference design, datasets are tiered across hot, warm, and cold storage layers. Which statement correctly describes the intended role of the hot tier?
The hot tier is designed for performance-critical, actively used data, using low-latency, high-throughput media to feed live training and inference workloads.
Question 6 of 6 · Storage and Data Pipelines
A cluster requires shared storage accessed both by GPU compute nodes needing NVMe-oF/RDMA access for training data and by traditional CPU-based ETL nodes needing standard NFS access to the same dataset. Which architecture best supports both access methods without duplicating data?
A parallel/distributed file system that natively exposes both RDMA client access and an NFS gateway over one namespace lets both node types read the same data concurrently without duplication or staleness.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →