Free Databricks Certified Data Engineer Associate practice — 6 questions on Databricks Intelligence Platform, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Databricks Intelligence Platform
A team runs a nightly Lakeflow Job that ingests 2TB of data and finishes in 20 minutes. The job runs once per day, and the team wants to minimize both compute cost and cluster startup latency without managing infrastructure sizing manually. Which compute configuration BEST meets these requirements?
Serverless compute for jobs eliminates manual cluster sizing, starts rapidly (no need to keep it running), and bills only for the compute used during the 20-minute run, matching a job that runs once daily and needs minimal management overhead.
Question 2 of 6 · Databricks Intelligence Platform
An engineer runs the following command on a production Delta table: VACUUM sales_transactions RETAIN 0 HOURS. Which statement correctly describes the outcome?
Delta Lake enforces a safety check (spark.databricks.delta.retentionDurationCheck.enabled) that blocks VACUUM calls with a retention period below 168 hours (7 days) to prevent breaking time travel and concurrent readers/writers. Running with 0 hours retention fails unless this check is disabled via SQL configuration, and doing so is explicitly a documented but dangerous exam trap.
Question 3 of 6 · Databricks Intelligence Platform
A data platform team wants to enforce that only members of the 'finance_analysts' group can query the table finance.reporting.quarterly_revenue, regardless of which workspace or cluster is used to access it. Which Databricks capability should they configure to achieve this?
Unity Catalog governs data access centrally at the catalog.schema.table level using GRANT/REVOKE statements, and these permissions are enforced consistently across all workspaces and compute resources attached to the metastore, independent of which cluster or notebook is used.
Question 4 of 6 · Databricks Intelligence Platform
In a medallion architecture pipeline, an engineer has raw JSON files landing in the bronze layer with duplicate records, inconsistent field names, and occasional malformed rows. Business users need clean, deduplicated, and conformed data joined with a reference dimension table before it can be aggregated for dashboards. In which layer should this deduplication, schema conformance, and joining logic be implemented?
The silver layer is designed to hold cleansed, deduplicated, and conformed data — including joins with reference/dimension tables — providing a validated foundation before gold-layer aggregation for business consumption.
Question 5 of 6 · Databricks Intelligence Platform
Which mechanism does Delta Lake use to provide ACID transaction guarantees and enable time travel queries against a table?
Delta Lake records every write as an ordered, versioned commit in JSON (and periodic checkpoint) files inside the _delta_log directory alongside the data files. This transaction log provides atomicity, consistency, and the version history needed for time travel queries like VERSION AS OF or TIMESTAMP AS OF.
Question 6 of 6 · Databricks Intelligence Platform
A gold-layer table is queried almost exclusively with filters on the customer_region column, and the table has grown to hundreds of small files after months of incremental writes. Query performance has degraded significantly. Which command should be run to BEST address both the file fragmentation and the filtering performance for this workload?
OPTIMIZE compacts small files into larger ones to reduce fragmentation, and ZORDER BY colocates related data for the specified column, enabling effective data skipping when queries filter on customer_region — directly addressing both stated problems.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.