Free Microsoft Certified: Azure Data Scientist Associate practice — 6 questions on Explore data and train models, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Explore data and train models
You configure a sweep job to tune hyperparameters for a gradient boosting model. Training is expensive, so you want to terminate poorly performing runs early, but only after each run has completed at least 5 intervals, and only if a run's primary metric is worse than 10% below the best run seen so far at that interval. Which early termination policy configuration meets this requirement?
BanditPolicy terminates a run if its primary metric falls outside the slack_factor (allowed slack) relative to the best performing run, evaluated at each interval, and delay_evaluation=5 defers policy application until after the 5th interval — exactly matching the 10% slack and 5-interval delay requirement.
Question 2 of 6 · Explore data and train models
A data scientist needs to define a reusable data asset with a fixed schema, apply column type conversions and drop unneeded columns at read time, and have that same transformed, tabular view of the data reused consistently across multiple training jobs and AutoML runs. Which Azure ML data asset type should they register?
MLTable is the only data asset type that stores a schema and a set of transformation steps (column type conversion, column drop, etc.) alongside the data reference, so the same tabular abstraction is reproduced identically wherever the asset is consumed, including by AutoML jobs.
Question 3 of 6 · Explore data and train models
You submit a command job for distributed PyTorch training across 4 compute nodes, using 4 GPU processes per node with native PyTorch DistributedDataParallel. Which SDK v2 job configuration correctly expresses this?
PyTorchDistribution is the SDK v2 distribution class purpose-built for native PyTorch DDP training; process_count_per_instance sets processes (GPUs) per node and instance_count sets the number of nodes, matching the 4x4 scenario exactly.
Question 4 of 6 · Explore data and train models
You are configuring an AutoML forecasting job to predict daily product sales, where sales from exactly 7 and 14 days prior are known to be strong predictors due to weekly and biweekly ordering cycles. Which forecasting configuration setting should you use to have AutoML generate these specific lagged features?
target_lags explicitly instructs AutoML forecasting to create feature columns using the target value from the specified number of periods in the past (7 and 14 days), which is exactly what's needed to capture the known weekly/biweekly relationship.
Question 5 of 6 · Explore data and train models
In an Azure ML pipeline job composed of multiple components, what determines whether a given component step's output is reused from a previous run (cached) instead of being re-executed?
Azure ML pipeline step reuse (caching) is based on hashing the component's inputs, source code, and environment; if this signature matches a previous successful run and the component is configured to allow reuse, the cached output is returned instead of re-running the step.
Question 6 of 6 · Explore data and train models
Your team registered a data asset named 'sales-data' (version '1') pointing to a folder of CSV files. A month later, new files are added to the source location, and you need to register this as a new, traceable version of the same logical dataset while keeping the original version intact for reproducibility of past experiments. Which approach is correct?
Azure ML data assets are versioned under a shared name; calling create_or_update with the same name and an incremented version number registers a new, independent, immutable version while preserving version '1' unchanged, giving full lineage and reproducibility across both experiment runs.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.