✅ Free practice — no sign-up📝 Real exam-style questions💡 Detailed explanations💸 30-day money-back via Udemy
Question 1 of 12 · Databricks Intelligence Platform
A data engineer schedules a nightly ETL notebook that runs for 45 minutes once per day. There is no need for interactive access to the compute outside of this scheduled run. Which cluster configuration minimizes cost while meeting the requirement?
Job clusters are created automatically when a Lakeflow Job runs and are terminated immediately when the job finishes, incurring lower job-compute DBU rates and no idle cost between runs.
Question 2 of 12 · Data Ingestion and Loading
A data engineer is using Auto Loader to ingest JSON files that occasionally include new columns added by the upstream source. The pipeline must automatically add these new columns to the target Delta table's schema without failing the stream. Which cloudFiles.schemaEvolutionMode value should be configured?
addNewColumns is the default and correct mode when new columns should be automatically merged into the target table schema, restarting the stream to pick up the change without data loss.
Question 3 of 12 · Data Transformation and Modeling
A pipeline receives daily CDC files for a silver 'customers' table. Each file contains inserts, updates, and deletes, identified by a boolean 'is_deleted' flag on the source records. Which single statement correctly applies all three change types to the target Delta table in one operation?
MERGE INTO is purpose-built for CDC/upsert patterns and lets you combine conditional DELETE, UPDATE, and INSERT clauses against the same target in a single atomic transaction, keyed on the matched condition.
Question 4 of 12 · Working with Lakeflow Jobs
A Lakeflow job has the following task graph: Task A runs first. Task B and Task C both depend on Task A and run in parallel. Task D depends on both B and C, with its dependency condition set to 'All done'. During a job run, Task B fails but Task C succeeds. What happens to Task D?
The 'All done' dependency condition means Task D runs once all upstream tasks have completed, whether they succeeded, failed, or were skipped.
Question 5 of 12 · Implementing CI/CD
A data engineering team wants to deploy the same Lakeflow Job definition to dev, staging, and prod workspaces, each with different cluster sizes and job schedules, from a single source-controlled project. Which approach BEST meets this requirement?
Databricks Asset Bundles let you define a project once in databricks.yml and use the 'targets' section to override cluster size, schedule, and workspace host per environment (dev/staging/prod), enabling consistent, repeatable deployment from a single source-controlled codebase.
Question 6 of 12 · Troubleshooting, Monitoring, and Optimization
A streaming job continuously appends small batches of data to a Delta table every few seconds. Over time, downstream queries against this table have become significantly slower due to a large number of small underlying files. What should the data engineer do to resolve this?
OPTIMIZE compacts the many small files produced by frequent streaming writes into fewer, larger files, which directly improves read performance by reducing file-listing and open/close overhead.
Question 7 of 12 · Governance and Security
A company stores employee salary data in the `hr.employees.compensation` table. Only members of the `hr_analysts` group should see actual salary values; all other users should see NULL for that column, while still being able to query the rest of the table. Which approach BEST implements this in Unity Catalog?
Dynamic views using CASE WHEN combined with is_member() (or current_user()) let you conditionally return real values or NULL per user/group at query time, enabling column-level masking while keeping the rest of the table accessible to everyone.
Question 8 of 12 · Databricks Intelligence Platform
In Unity Catalog, a data engineer wants to fully qualify a table reference in a SQL query so that it unambiguously identifies the exact table across the metastore. Which three-level namespace format does Unity Catalog use?
Unity Catalog organizes data using a three-level namespace: catalog (top-level container), schema (equivalent to a database), and table, e.g. main.sales.orders.
Question 9 of 12 · Data Ingestion and Loading
A nightly batch job uses COPY INTO to load new CSV files from cloud storage into a Delta table. Due to an orchestration bug, the job runs twice on the same day against the same source directory. What is the result?
COPY INTO maintains metadata about which source files have already been loaded into the target table, making repeated executions idempotent — already-processed files are automatically skipped.
Question 10 of 12 · Data Transformation and Modeling
A gold-layer daily sales summary needs to be automatically and incrementally recomputed on a schedule by a Lakeflow Declarative Pipeline, with results persisted so downstream BI queries don't pay recomputation cost on every read. Which object type should be defined for this?
Materialized views in Lakeflow Declarative Pipelines are declaratively defined, automatically and incrementally refreshed by the pipeline engine on each update, and their results are physically persisted for fast downstream reads.
Question 11 of 12 · Working with Lakeflow Jobs
A data engineer configures a task in a Lakeflow job to retry up to 2 times with a minimum interval of 5 minutes between attempts, and sets an email alert to fire 'on failure'. The task fails on the first attempt but succeeds on the retry. What happens regarding the alert?
Failure alerts are based on the final status of the task run. Since the retry succeeded, the task's overall result is success, so no failure notification is triggered.
Question 12 of 12 · Implementing CI/CD
A team wants their CI/CD pipeline to run 'databricks bundle validate' on every pull request and 'databricks bundle deploy --target prod' only after a merge to the main branch. Which setup accomplishes this?
GitHub Actions (or any CI runner) can invoke the Databricks CLI to run 'bundle validate' on pull_request events and 'bundle deploy --target prod' on push to main, which is the standard pattern for automating Asset Bundle deployment through CI/CD.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.
Databricks Data Engineer Associate exam — quick answers
How much does the Databricks Data Engineer Associate exam cost?
The exam fee is approximately $200 and varies by region — confirm current pricing with the certification vendor before you book.
What topics are on the exam?
It covers 7 domains: Databricks Intelligence Platform (~15%), Data Ingestion and Loading (~18%), Data Transformation and Modeling (~22%), Working with Lakeflow Jobs (~15%), Implementing CI/CD (~12%), Troubleshooting, Monitoring, and Optimization (~10%), Governance and Security (~8%). The full course has a dedicated chapter, lab and practice-test coverage for each.
Is this practice test really free?
Yes — all questions on this page are free with explanations and no sign-up. The paid Udemy course adds two full-length timed exams, video lessons and hands-on labs.
Will this prepare me for the real exam?
The questions mirror the real exam's style and are mapped to the official domains. This is exam-focused preparation — combine the free test with the full course's timed simulations to gauge your readiness.