TechNuggets Academy

Engineering Features for Machine Learning

Free Certified Artificial Intelligence Practitioner practice — 6 questions on Engineering Features for Machine Learning, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Engineering Features for Machine Learning
A data scientist builds a churn-prediction model using historical customer transaction data. One engineered feature, "customer_lifetime_value," is computed using each customer's ENTIRE transaction history, including transactions that occurred after the prediction date used for that training row. Another feature, "average_purchase_last_90_days," is computed using only data available up to the prediction date. During evaluation, the model shows unusually high accuracy that collapses in production. Which BEST describes the problem and the fix?
Using post-prediction-date data to construct a training feature is a textbook case of target/temporal leakage: the model learns from information it would never have access to at real prediction time, inflating offline metrics and collapsing in production. The fix is to recompute the feature using a point-in-time cutoff consistent with when the prediction would actually be made.
Question 2 of 6 · Engineering Features for Machine Learning
A dataset contains a categorical feature "zip_code" with 15,000 unique values. The team wants to encode it for a gradient-boosted tree model while minimizing both dimensionality and the risk of target leakage. Which approach is MOST appropriate?
High-cardinality categorical features are best handled with target or frequency encoding computed strictly within cross-validation folds (or using out-of-fold schemes) with smoothing for rare categories. This keeps dimensionality low while preventing the encoding from leaking test-set target information into training.
Question 3 of 6 · Engineering Features for Machine Learning
During feature selection for a linear regression model, a data scientist computes the Variance Inflation Factor (VIF) for each numeric predictor. Which VIF value is generally used as a threshold indicating problematic multicollinearity that warrants removing or combining features?
A VIF greater than 5 (and especially above 10) is the commonly cited rule-of-thumb threshold indicating that a predictor is highly correlated with other predictors, inflating coefficient variance and warranting removal, combination, or dimensionality reduction.
Question 4 of 6 · Engineering Features for Machine Learning
A dataset has missing values in the "income" column. Investigation shows the missingness resulted purely from an intermittent database logging error unrelated to the income value itself or to any other observed variable in the dataset. Which missing-data mechanism does this describe?
MCAR (Missing Completely At Random) describes a scenario where the probability of a value being missing is unrelated to both the missing value itself and any other observed variables — exactly the case of a random logging error, as opposed to missingness driven by data content.
Question 5 of 6 · Engineering Features for Machine Learning
A bank builds a credit-risk model that explicitly excludes "race" as an input feature. However, an audit finds that two engineered features, "zip_code" and "neighborhood_median_income," are strongly correlated with race due to historical housing segregation, and the model exhibits disparate impact against certain racial groups. What is the MOST appropriate feature-engineering response?
This is a classic 'fairness through unawareness' failure: excluding a protected attribute does not prevent bias if correlated proxy variables remain. The correct engineering practice is to explicitly identify proxy features, assess and mitigate their disparate impact, and document the fairness analysis — addressing the ethical/business risk directly rather than assuming exclusion of the literal attribute is sufficient.
Question 6 of 6 · Engineering Features for Machine Learning
A data scientist is preparing a fraud-detection dataset where only 2% of transactions are labeled fraudulent. To address the class imbalance, they apply SMOTE to generate synthetic minority-class samples across the ENTIRE dataset, and only AFTER that split the data into training and test sets. What is the problem with this workflow?
SMOTE creates synthetic samples by interpolating between nearest neighbors in the minority class. If SMOTE is applied before the train/test split, synthetic training samples can be generated using information derived from points that end up in the test set (and vice versa), leaking test-set information into training and producing an overly optimistic evaluation. SMOTE must be applied only to the training set after splitting.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →