TechNuggets Academy

Training and Tuning ML Systems and Models

Free Certified Artificial Intelligence Practitioner practice — 6 questions on Training and Tuning ML Systems and Models, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Training and Tuning ML Systems and Models
A team tunes hyperparameters for a gradient boosting model using 5-fold cross-validation on the full training set, selecting the hyperparameter combination with the best average CV score. They then report that same CV score as the expected generalization performance of the final model. What is wrong with this approach, and what is the BEST fix?
Because the same folds were used to both select hyperparameters and estimate final performance, the reported score is optimistically biased. Nested cross-validation separates these two purposes — the inner loop tunes hyperparameters, the outer loop provides an unbiased generalization estimate.
Question 2 of 6 · Training and Tuning ML Systems and Models
A team plots learning curves for a classifier trained on 2 million examples. Training accuracy plateaus at 99%, validation accuracy plateaus at 72%, and the gap does not narrow even after adding substantially more training data. Which action will MOST directly address this problem?
A persistent large gap between training and validation performance that does not close with more data is the signature of high variance (overfitting). Reducing model complexity or adding stronger regularization directly targets this.
Question 3 of 6 · Training and Tuning ML Systems and Models
A fraud detection dataset has a class ratio of roughly 1:1000 (fraud:non-fraud). The team wants the loss function to automatically emphasize hard-to-classify examples during training, rather than simply reweighting classes uniformly or generating new minority-class samples. Which technique is BEST suited to this requirement?
Focal loss adds a modulating factor that down-weights well-classified (easy) examples and focuses training on hard, misclassified examples, dynamically adapting per example rather than applying a static class-level adjustment.
Question 4 of 6 · Training and Tuning ML Systems and Models
A team must tune 15 hyperparameters for a deep learning model, where each training run takes 8 hours on expensive GPU clusters. Exhaustively evaluating a grid of combinations is infeasible. Which hyperparameter search strategy is MOST appropriate to minimize the number of expensive evaluations while still converging on strong hyperparameters?
Bayesian optimization builds a surrogate model (e.g., a Gaussian process) of the objective function and uses an acquisition function to select the most promising hyperparameters to evaluate next, making it far more sample-efficient for expensive, high-dimensional searches.
Question 5 of 6 · Training and Tuning ML Systems and Models
A 50-layer feedforward neural network using sigmoid activations throughout is being trained. The team observes that weights in the first 10 layers barely change across many epochs, while later layers update normally. What is the MOST likely cause and the BEST fix?
Sigmoid activations saturate for large-magnitude inputs, producing near-zero derivatives. As gradients backpropagate through many layers, these small derivatives multiply together and shrink toward zero, so early layers receive almost no gradient signal. Using ReLU-family activations and/or residual (skip) connections mitigates this by preserving gradient flow.
Question 6 of 6 · Training and Tuning ML Systems and Models
A linear regression model is regularized on a dataset containing groups of highly correlated features. The team wants coefficients on irrelevant features driven to exactly zero (automatic feature selection) while also having correlated features within a group tend to be selected or dropped together, rather than arbitrarily keeping one and zeroing the rest. Which regularization approach BEST meets both requirements?
Elastic Net combines the L1 penalty's sparsity-inducing property (driving irrelevant coefficients to zero) with the L2 penalty's grouping effect, which encourages correlated features to receive similar coefficients rather than arbitrarily selecting one from the group.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →