TechNuggets Academy

Data Ingestion and Loading

Free Databricks Certified Data Engineer Associate practice — 6 questions on Data Ingestion and Loading, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Data Ingestion and Loading
A streaming Auto Loader job ingests JSON files with cloudFiles.schemaEvolutionMode set to 'none', and cloudFiles.rescuedDataColumn is NOT configured. New source files begin arriving with a column that is not present in the target table's current inferred schema. What happens to the data in that new column?
With schemaEvolutionMode set to 'none', schema evolution is disabled and new columns are simply ignored. Unlike 'rescue' mode, this mode does not automatically capture unmatched data — the rescued data column only captures data if cloudFiles.rescuedDataColumn is explicitly configured.
Question 2 of 6 · Data Ingestion and Loading
A data engineer runs the exact same COPY INTO command twice against a Delta table, pointing at the same source directory whose contents have not changed since the first run. What happens on the second execution?
COPY INTO is idempotent by design: it tracks which source files have already been loaded into the target table using metadata stored in the Delta transaction log, and automatically skips files it has already processed on subsequent runs.
Question 3 of 6 · Data Ingestion and Loading
A company needs to continuously ingest data from Salesforce into Delta tables with built-in change data capture handling and minimal custom engineering effort. Which approach BEST meets this requirement?
Lakeflow Connect provides fully managed, prebuilt connectors for SaaS applications and databases such as Salesforce, handling authentication, incremental extraction, and CDC automatically with minimal engineering effort.
Question 4 of 6 · Data Ingestion and Loading
By default, when Auto Loader infers the schema of a CSV or JSON data source, what data type does it assign to all inferred columns unless cloudFiles.inferColumnTypes is explicitly set to true?
By default, cloudFiles.inferColumnTypes is false, so Auto Loader infers all columns as StringType during schema inference for CSV and JSON sources. Setting cloudFiles.inferColumnTypes to true is required to get more specific inferred types.
Question 5 of 6 · Data Ingestion and Loading
A data engineer previously loaded a batch of files using COPY INTO, but later discovers the source files contained corrupted values that have since been corrected in place at the same file paths. Which clause should be added to the COPY INTO statement to force these already-loaded files to be reprocessed?
COPY_OPTIONS ('force' = 'true') tells COPY INTO to ignore its normal idempotency tracking and reprocess files it has already loaded, which is necessary when the same file paths contain corrected data.
Question 6 of 6 · Data Ingestion and Loading
A company ingests files into a cloud storage bucket that receives millions of new files per day. An Auto Loader job using the default directory listing mode is experiencing increasing latency in discovering new files as the number of files grows. Which configuration change would BEST address this?
File notification mode (cloudFiles.useNotifications = true) uses cloud provider event/queue services (e.g., AWS SQS/SNS, Azure Event Grid, GCP Pub/Sub) to discover new files incrementally and efficiently, avoiding the scalability limits of full directory listings at high file volumes.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →