Free Databricks Certified Data Engineer Associate practice — 6 questions on Working with Lakeflow Jobs, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Working with Lakeflow Jobs
A Lakeflow Job uses a file arrival trigger on a Volume path where upstream systems continuously drop small files every few seconds. The engineer notices the job is starting far more often than needed, consuming cluster startup overhead. Which trigger setting should be adjusted to reduce run frequency without disabling the file arrival trigger?
File arrival triggers expose a 'Minimum time between triggers' setting (in seconds) that throttles how often the job can fire even if new files keep arriving, directly solving the over-triggering problem.
Question 2 of 6 · Working with Lakeflow Jobs
A task is configured with Max retries = 2 and Minimum retry interval = 2 minutes. The task fails on its first attempt after running for 10 minutes. Which statement correctly describes the retry behavior?
Lakeflow Jobs retry policy enforces a fixed minimum interval between the failed attempt and the next retry attempt, and stops after the configured max retries is reached — there is no automatic exponential backoff.
Question 3 of 6 · Working with Lakeflow Jobs
An engineer configures a For Each task to iterate over 500 partition values, invoking the same notebook task once per value. To avoid overwhelming the cluster, they want to cap how many iterations run in parallel. What is the maximum concurrency value that can be set for a For Each task?
The For Each task in Lakeflow Jobs supports a configurable concurrency setting, but the maximum allowed value is capped at 100 concurrent iterations regardless of iteration count.
Question 4 of 6 · Working with Lakeflow Jobs
A team has an existing ingestion job used by three separate downstream Lakeflow Jobs. Instead of duplicating the ingestion tasks into each of the three jobs, the team wants a single source of truth for the ingestion logic while still triggering it as part of each downstream workflow. Which approach BEST achieves this?
The Run Job task type lets one job invoke another existing job as a task, enabling modular, reusable workflows without duplicating task definitions — updates to the ingestion job automatically apply everywhere it's referenced.
Question 5 of 6 · Working with Lakeflow Jobs
Task A runs a notebook that computes a validation record count. This value must control whether an If/else condition task allows Task C to run downstream. Which implementation correctly wires this together?
dbutils.jobs.taskValues.set() in the upstream task publishes a named value scoped to that task run, and downstream If/else or notebook tasks can reference it directly via the {{tasks.<task_key>.values.<key>}} syntax in the task's condition or parameters.
Question 6 of 6 · Working with Lakeflow Jobs
A workflow has Task X depending on upstream Tasks A, B, and C, which run in parallel. The engineer wants Task X to execute a cleanup/alerting step whenever at least one of A, B, or C fails, but Task X should be skipped if all three succeed. Which 'Run if' dependency setting on Task X achieves this?
The 'At least one failed' Run if condition triggers the dependent task only when one or more of its upstream dependencies fail, matching the requirement to run cleanup on failure and skip on full success.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.