TechNuggets Academy

Operations and AIOps

Free Cisco Certified Specialist - Data Center AI Infrastructure practice — 6 questions on Operations and AIOps, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Operations and AIOps
A GPU training cluster using RoCEv2 experiences intermittent complete stalls where multiple switches show PFC pause frames looping between ports, but no single congested endpoint can be identified as the root cause. Which Nexus Dashboard Insights capability BEST identifies and helps recover from this condition?
PFC deadlock is a circular buffer-dependency condition detected through pause-frame loop analysis; NDI surfaces this as an anomaly and correlates it with PFC watchdog/storm recovery mechanisms to break the loop.
Question 2 of 6 · Operations and AIOps
An architect wants to confirm that a proposed ACL and QoS policy change to enable lossless RoCEv2 transport will not disrupt existing PFC/ECN configurations before pushing it to production switches in the AI fabric. Which capability should be used?
Pre-Change Analysis simulates the impact of a proposed configuration change against the current assurance/compliance state before it is committed, flagging conflicts such as broken PFC/ECN policy interactions.
Question 3 of 6 · Operations and AIOps
Which telemetry mechanism must be enabled on Cisco Cloud Scale ASIC-based Nexus switches to give Nexus Dashboard Insights per-flow buffer occupancy and drop visibility needed for troubleshooting RoCEv2 congestion?
Flow Telemetry leverages ASIC-level hardware export of per-flow buffer occupancy, latency, and drop counters at microsecond granularity, which is required to troubleshoot transient RoCEv2 congestion events.
Question 4 of 6 · Operations and AIOps
In Nexus Dashboard Insights, what does an increasing Anomaly Score trend for a fabric object across multiple analysis epochs indicate, as opposed to a single one-time Advisory notification?
A rising anomaly score across successive epochs signals a recurring or worsening issue on that object, which NDI uses to help prioritize root-cause investigation over transient one-off events.
Question 5 of 6 · Operations and AIOps
During distributed GPU training, nodes report periodic latency spikes lasting under 100 microseconds that standard SNMP-based monitoring never detects. Which Nexus Dashboard Insights capability is specifically designed to capture and report these sub-millisecond buffer events?
Hardware-exported Flow Telemetry captures microsecond-level buffer occupancy spikes (microbursts) that are invisible to polling-based tools with multi-second or multi-minute intervals.
Question 6 of 6 · Operations and AIOps
A new leaf switch being added to an AI fabric via NDFC POAP obtains an IP address via DHCP and retrieves the POAP bootstrap script, but the subsequent configuration download step times out. What is the MOST likely cause?
After the DHCP and script retrieval stages succeed, the config download step relies on NDFC correctly mapping the switch's serial number to a defined switch/role in the fabric template; a missing or incorrect mapping causes the download to stall or time out.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →