Free CompTIA DataX practice — 6 questions on Modeling, Analysis, and Outcomes, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Modeling, Analysis, and Outcomes
A data scientist is building a churn prediction model to predict whether a customer will churn within the next 30 days. The dataset includes a feature called 'support_tickets_after_churn_decision', which captures the number of support tickets the customer opened in the 30 days following the churn event. The model achieves 98% accuracy on the holdout test set, but performance drops to 62% in production. What is the MOST likely cause of this discrepancy?
The feature is derived from events occurring after the churn outcome has already occurred, meaning it cannot exist at the time a real prediction must be made. This creates artificially inflated test performance that collapses once the feature is unavailable (or meaningless) in production.
Question 2 of 6 · Modeling, Analysis, and Outcomes
A data science team is building a fraud detection model where fraudulent transactions represent only 0.3% of all transactions. The business requires the model to rank transactions by fraud risk so that only the top 1% highest-risk transactions are sent for manual review. Which evaluation metric is MOST appropriate for comparing candidate models in this scenario?
With extreme class imbalance and a ranking-based business use case, PR-AUC focuses on performance for the minority (fraud) class and is not inflated by the large number of true negatives, making it the most informative metric for comparing model quality on rare-event ranking tasks.
Question 3 of 6 · Modeling, Analysis, and Outcomes
An analyst discovers that income data is missing for high-income survey respondents at a substantially higher rate than for low-income respondents, and no observed variable in the dataset explains this pattern. Which missing data mechanism best describes this situation, and what is an appropriate strategy?
Because the probability of missingness is driven directly by the (unobserved) income value itself rather than by any observed variable, this is MNAR. Standard imputation methods assume MAR and will be biased; approaches like selection models or pattern-mixture models are needed to explicitly account for the missingness mechanism.
Question 4 of 6 · Modeling, Analysis, and Outcomes
A credit default prediction model outputs predicted probabilities. The business assigns a cost of $500 for a false negative (a missed default) and $50 for a false positive (unnecessarily denying credit to a good customer). Which approach should the data scientist use to select the classification threshold instead of the default 0.5 cutoff?
Because the costs of false negatives and false positives are explicitly asymmetric and quantified, the threshold should be chosen to minimize total expected business cost using those weights directly, rather than a metric that implicitly assumes equal error costs.
Question 5 of 6 · Modeling, Analysis, and Outcomes
Which statement correctly differentiates wrapper methods from filter methods in feature selection?
Wrapper methods (e.g., recursive feature elimination) select features by repeatedly training and evaluating a specific model, which is expensive but captures interactions between features; filter methods (e.g., correlation, chi-square, mutual information) score features independently of any model, making them fast but potentially blind to feature interactions.
Question 6 of 6 · Modeling, Analysis, and Outcomes
A loan officer asks the data science team to explain why one specific applicant's loan application was denied by the model, so the officer can communicate the specific reasons to that applicant as required by regulation. Which explainability technique is MOST appropriate for this request?
SHAP values provide a local, instance-level explanation, attributing the contribution of each feature to one specific prediction, which is exactly what is needed to explain an individual applicant's adverse decision as regulations often require.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.