Free CompTIA DataX practice — 6 questions on Machine Learning, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Machine Learning
A data scientist is training an SVM classifier on a dataset with 50,000 engineered features but only 300 training samples (severe p >> n scenario). Which kernel choice minimizes overfitting risk while keeping training time reasonable?
When the number of features vastly exceeds the number of samples, the data is already likely to be linearly separable in the high-dimensional space, so a linear kernel provides sufficient capacity without adding unnecessary nonlinear complexity, reducing overfitting and avoiding costly kernel matrix computations.
Question 2 of 6 · Machine Learning
As the value of k increases in a k-nearest neighbors classifier, what happens to the model's bias and variance?
Larger k averages predictions over more neighbors, smoothing the decision boundary. This smoothing increases bias (the model becomes less flexible and can miss local patterns) while reducing variance (predictions become less sensitive to noise in any single training point).
Question 3 of 6 · Machine Learning
A single decision tree model exhibits very low training error but significantly higher validation error, indicating high variance and severe overfitting. Which ensemble technique is BEST suited to directly address this specific problem?
Random Forest uses bootstrap aggregation (bagging) combined with random feature subsetting, which averages out the high variance of individual overfit trees, directly targeting the variance problem described without increasing bias much.
Question 4 of 6 · Machine Learning
A data scientist needs a nonlinear dimensionality reduction technique that preserves both local and global structure, scales efficiently to datasets with over 1 million rows, and supports transforming new, previously unseen data points after the model is fit. Which technique should be chosen?
UMAP is built on a learned fuzzy topological representation that supports a proper transform() method for new data, scales significantly better than t-SNE on large datasets due to its optimized nearest-neighbor and embedding graph construction, and is designed to better preserve global structure alongside local neighborhoods compared to t-SNE.
Question 5 of 6 · Machine Learning
The standard self-attention mechanism in a Transformer architecture computes pairwise attention scores between all tokens in a sequence. What is the computational complexity of this operation with respect to sequence length n?
Standard self-attention computes an n x n attention score matrix (every token attending to every other token), giving quadratic O(n^2) time and memory complexity in sequence length, which is a well-known limitation motivating techniques like sparse or linear attention for long sequences.
Question 6 of 6 · Machine Learning
A dataset contains 200 features, many of which are highly correlated with one another. The data scientist wants automatic feature selection (driving irrelevant coefficients toward zero) while maintaining stable coefficient estimates despite the multicollinearity. Which regularization approach BEST meets these requirements?
Elastic Net combines the L1 penalty (which drives coefficients to exactly zero for feature selection) with the L2 penalty (which stabilizes coefficient estimates and handles grouped, correlated predictors by shrinking them together rather than arbitrarily picking one), directly satisfying both stated requirements.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.