TechNuggets Academy

Working with Data for Gen AI

Free SnowPro Specialty: Gen AI practice — 6 questions on Working with Data for Gen AI, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Working with Data for Gen AI
A team stores document embeddings generated by different embedding models, resulting in vectors with varying magnitudes but consistent semantic direction. Which Snowflake vector function should they use in a similarity search query to rank results primarily by semantic direction rather than raw vector magnitude?
VECTOR_COSINE_SIMILARITY normalizes for vector length by measuring the angle between vectors, so it ranks results based on semantic direction rather than magnitude, which is critical when embeddings come from sources with inconsistent scaling.
Question 2 of 6 · Working with Data for Gen AI
A support team needs a search experience over a growing knowledge base that combines exact keyword matches (e.g., product SKU numbers) with semantic similarity, refreshes automatically as new articles are added, and requires no manual index management or infrastructure provisioning. Which Snowflake capability should they use?
Cortex Search Service provides managed hybrid search (keyword plus semantic/vector) with automatic embedding generation and incremental refresh based on a configured target lag, eliminating manual index or infrastructure management.
Question 3 of 6 · Working with Data for Gen AI
Which configuration constraint applies when calling SNOWFLAKE.CORTEX.SPLIT_TEXT_RECURSIVE_CHARACTER to chunk documents for a RAG pipeline?
SPLIT_TEXT_RECURSIVE_CHARACTER requires overlap to be smaller than chunk_size (overlapping more than the chunk itself is invalid), and the format argument selects whether markdown-aware splitting (headers, code blocks) is applied or plain-text splitting is used.
Question 4 of 6 · Working with Data for Gen AI
What is the maximum number of dimensions supported by Snowflake's VECTOR data type?
Snowflake's VECTOR data type supports up to 4096 dimensions, which comfortably accommodates common embedding sizes (e.g., 768 or 1024) as well as larger custom embedding models.
Question 5 of 6 · Working with Data for Gen AI
A pipeline extracts text from scanned PDF invoices containing multi-column tables of line items and totals. Preserving row and column relationships in the extracted text is critical for downstream RAG accuracy. Which configuration of SNOWFLAKE.CORTEX.PARSE_DOCUMENT should be used?
PARSE_DOCUMENT's LAYOUT mode is designed to preserve document structure — including tables, headers, and reading order — by outputting structured markdown, making it the correct choice when tabular relationships must be retained for RAG.
Question 6 of 6 · Working with Data for Gen AI
A team is building a multilingual RAG pipeline using SNOWFLAKE.CORTEX.EMBED_TEXT_1024 with the multilingual-e5-large model, and plans to feed those embeddings into a Cortex Search Service for retrieval. Which statement about aligning the two is correct?
Cortex Search Service is configured with a source text column and an EMBEDDING_MODEL setting; it generates and maintains its own semantic index from raw text, so pre-computed EMBED_TEXT_1024 vectors are not what the service consumes for its internal search index — a common source of confusion.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →