TechNuggets Academy

Design and prepare a machine learning solution

Free Microsoft Certified: Azure Data Scientist Associate practice — 6 questions on Design and prepare a machine learning solution, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Design and prepare a machine learning solution
A company requires that all resources associated with its Azure ML workspace (storage account, key vault, container registry) be reachable only over private endpoints, and that compute nodes in the workspace have no direct path to the public internet except for a small, explicitly approved list of destinations. Which configuration meets this requirement?
'Allow only approved outbound' on a managed virtual network restricts outbound traffic from compute to an explicit allow-list, and private endpoints on storage, key vault, and container registry ensure those dependencies are never reached over the public internet — satisfying both requirements simultaneously.
Question 2 of 6 · Design and prepare a machine learning solution
A data scientist runs ad-hoc training jobs a few times per week. She does not want to create, size, or manage a compute cluster, and wants nodes to spin up only for the duration of each job and scale to zero automatically afterward with no ongoing management. Which compute option should she use?
Serverless compute in Azure ML lets you submit a job by simply specifying resource requirements (VM size, instance count) without creating or managing a compute target at all — Azure ML provisions, scales, and tears down the underlying nodes automatically per job.
Question 3 of 6 · Design and prepare a machine learning solution
A training job requires CUDA 12.2 and a proprietary library that is not available in any curated Azure ML environment. The environment must produce identical, reproducible results every time the job is rerun, including across different compute targets. What should the data scientist do?
A custom environment built from a Dockerfile lets you pin the exact CUDA base image and install the proprietary library at build time. Azure ML versions the resulting environment image, so referencing it by name:version guarantees the exact same environment is used on every rerun, regardless of compute target.
Question 4 of 6 · Design and prepare a machine learning solution
Multiple data scientists share an Azure ML workspace and each has different, individually-assigned RBAC permissions on folders within an ADLS Gen2 account. The requirement is that when each user accesses the data from a notebook, only their own Azure AD permissions determine what they can read — no shared credential should grant broader access. Which datastore configuration satisfies this?
An identity-based datastore does not store a shared secret; instead it passes through the calling user's own Azure AD identity to the storage account, so ADLS Gen2 RBAC/ACL permissions already assigned per-user are enforced exactly as configured.
Question 5 of 6 · Design and prepare a machine learning solution
A platform team wants training jobs, environments, and online endpoints defined as version-controlled specification files that can be deployed identically across dev, test, and prod through Azure DevOps pipeline stages, with no interactive or manual step. Which approach best fits this requirement?
CLI v2 is designed around declarative YAML specification files for jobs, environments, components, and endpoints. These YAML files are naturally version-controlled and can be applied non-interactively via 'az ml' commands in CI/CD pipeline stages, giving identical, repeatable deployments across environments.
Question 6 of 6 · Design and prepare a machine learning solution
A data scientist is preparing tabular training data with mixed column types (dates, categoricals, numerics) for AutoML, and later needs the exact same schema and type inference applied consistently when scoring new data in a batch endpoint. Which data asset type should be created for the training data?
An MLTable data asset encodes the schema, column types, and transformation/read logic in its MLTable definition file, so the same data-loading and type-inference behavior is reproduced consistently whenever the asset is referenced, including at batch scoring time — which is exactly what AutoML expects for tabular data.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →