TechNuggets Academy

Identify Data Needs

Free PMI Certified Professional in Managing AI practice — 6 questions on Identify Data Needs, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Identify Data Needs
A CPMAI project team is building a fraud-detection model. During Identify Data Needs, exploratory data analysis on two years of historical transaction data reveals 0.3% fraud cases (severe class imbalance) and 40% missing values in the three features SMEs identified as most predictive of fraud. The team must decide whether to proceed to Data Preparation on schedule. Per the CPMAI go/no-go discipline, what is the BEST course of action?
CPMAI explicitly builds in a go/no-go decision point after data evaluation. When core predictive features are severely incomplete and the target class is critically underrepresented, the responsible action is to formally flag this to stakeholders and pursue additional data sourcing or scope renegotiation rather than silently absorbing the risk downstream — this is the defining discipline that separates AI projects from traditional software projects.
Question 2 of 6 · Identify Data Needs
During Identify Data Needs for a customer churn model, the required internal data lives in a finance-owned data warehouse that requires governance council approval to access, and a needed third-party behavioral dataset requires a licensing and compliance review before use. The project sponsor wants a firm 3-week delivery commitment for the training dataset. What should the CPMAI-trained project lead do?
Identify Data Needs explicitly includes mapping data sources along with their ownership, access permissions, licensing, and compliance requirements. The correct practice is to surface and resolve these gating dependencies early and set realistic commitments based on confirmed access, rather than assuming approvals or skipping governance.
Question 3 of 6 · Identify Data Needs
Which role is PRIMARILY responsible for defining acceptable data quality thresholds and enforcing organizational data governance policies during the Identify Data Needs sub-phase of a CPMAI project?
The Data Steward owns data quality standards, governance policy enforcement, and stewardship of data assets, and is the role CPMAI identifies as a required stakeholder to engage when identifying data needs and evaluating whether data meets governance and quality standards.
Question 4 of 6 · Identify Data Needs
A predictive maintenance project ingests continuous IoT sensor data streaming at several gigabytes per hour from manufacturing equipment. During Identify Data Needs, what infrastructure characteristic must be specified to prevent bottlenecks once Data Preparation begins?
Coordinating the AI workspace and infrastructure is an explicit Identify Data Needs activity. High-velocity, high-volume streaming data requires elastic ingestion and storage provisioned for sustained throughput and burst capacity, identified up front so Data Preparation and later phases are not blocked by infrastructure limits.
Question 5 of 6 · Identify Data Needs
A team is defining data requirements for a model predicting rare equipment failures, where the historical failure rate is under 0.1% of all records. To ensure the eventual training and evaluation sets adequately represent failure events, what sampling approach should be specified in the Identify Data Needs data requirements?
Defining sampling requirements is an explicit Identify Data Needs activity. For extremely rare-event targets, requirements must call for stratified sampling that deliberately oversamples the minority class so downstream training and evaluation sets contain enough positive examples to build and validate a usable model.
Question 6 of 6 · Identify Data Needs
Why does CPMAI treat the data go/no-go decision as fundamentally distinct from a requirements sign-off in a traditional software development project?
CPMAI's core axiom is that AI projects are data projects — feasibility hinges on whether sufficient, quality, representative data actually exists to support the intended outcome. This is unlike traditional software, where once requirements are defined, functionality can generally be engineered regardless of data availability. The go/no-go decision exists specifically to test this data-dependent feasibility before committing further investment.
Ready for the real thing?

The full course: two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed explanations.

$109.99 $34.99 with code FREETEST33 — valid through September 14.

Get my $34.99 deal →