Free Databricks Certified Machine Learning Professional practice — 6 questions on Model Development, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Model Development
A data scientist needs to train 2,000 independent store-level demand forecasting models on a Databricks cluster with 16 worker nodes. Each individual model's training data fits comfortably in the memory of a single executor, and each model trains in under 5 seconds with scikit-learn. The goal is to train all 2,000 models as quickly as possible by using the cluster's parallelism. Which approach is correct?
applyInPandas (the pandas Function API) is designed exactly for this pattern: many independent, single-node-sized model-training tasks that need to run in parallel across a Spark cluster. Grouping by store ID sends each group's pandas DataFrame to a worker task, where a normal single-node scikit-learn model is trained and results are returned as rows in a Spark DataFrame.
Question 2 of 6 · Model Development
A team is building a training set with the Feature Engineering client for a churn model. The label table has an event_timestamp column marking when each label was observed, and the feature table is updated continuously in near real time. To avoid label leakage from features that were written after the label event occurred, which configuration correctly enforces a point-in-time correct join when calling create_training_set?
FeatureLookup's timestamp_lookup_key parameter tells create_training_set which column in the label DataFrame to use for an as-of join, ensuring only feature values with a timestamp at or before the label's event_timestamp are retrieved, which is the built-in mechanism for point-in-time correctness in the Feature Engineering client.
Question 3 of 6 · Model Development
A team wants to run Optuna hyperparameter tuning trials in parallel across the worker nodes of a multi-node Databricks cluster, rather than running all trials sequentially on the driver. Which approach correctly achieves distributed trial execution with Optuna on Databricks?
Optuna integrates with joblib, and by registering the joblib-spark backend, calls to study.optimize(..., n_jobs=-1) dispatch trials as Spark tasks that run in parallel across the cluster's executors, which is the supported way to distribute Optuna trials on Databricks.
Question 4 of 6 · Model Development
During an Optuna hyperparameter sweep, a data scientist wants every trial's parameters and metrics automatically logged to MLflow as a separate nested child run underneath one parent run representing the entire sweep. Which implementation correctly achieves this nested run structure?
Nested runs in MLflow require an active parent run context; calling mlflow.start_run(nested=True) inside the objective function while the parent run is still open creates a proper child run for each trial, linking it under the parent via parentRunId automatically.
Question 5 of 6 · Model Development
A team wraps a scikit-learn classifier together with a custom preprocessing step inside an mlflow.pyfunc.PythonModel subclass and logs it with mlflow.pyfunc.log_model() for deployment to a Model Serving endpoint. Why is explicitly defining a model signature at logging time particularly important in this scenario?
For custom PyFunc models, the logged signature is the mechanism MLflow and Model Serving use to validate that incoming JSON/pandas payloads match the expected column names and types, catching malformed requests before they reach potentially fragile custom preprocessing code.
Question 6 of 6 · Model Development
A real-time fraud detection model needs a feature representing the number of seconds between the current transaction's request time and the customer's most recent prior transaction. This value depends on the exact moment the scoring request arrives and cannot be known in advance. Which Feature Engineering approach is correct for this requirement?
On-demand features are computed at inference/scoring time from raw inputs supplied in the request, which is required here because the feature depends on the exact request timestamp and cannot be materialized in advance in any feature table.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.