TechNuggets Academy

Snowflake Gen AI and LLM Concepts

Free SnowPro Specialty: Gen AI practice — 6 questions on Snowflake Gen AI and LLM Concepts, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Snowflake Gen AI and LLM Concepts
A legal team is building a Streamlit app that uses SNOWFLAKE.CORTEX.COMPLETE to answer questions about 100-page contracts, and they want to pass the entire contract as context in a single call without chunking or RAG. Given typical context window sizes for Cortex LLM functions models, which model choice BEST supports this requirement?
mistral-large2 offers a roughly 128K token context window in Cortex, large enough to hold a ~100 page contract in a single call without chunking, unlike the smaller-context models.
Question 2 of 6 · Snowflake Gen AI and LLM Concepts
A table has a column defined as EMBEDDING_VECTOR VECTOR(FLOAT, 1024) storing document embeddings used for similarity search. Which Cortex function must be used to generate new embeddings that are dimensionally compatible with this column?
EMBED_TEXT_1024 returns a 1024-dimension vector matching the column definition, allowing VECTOR_COSINE_SIMILARITY or VECTOR_L2_DISTANCE to run without a dimension mismatch error.
Question 3 of 6 · Snowflake Gen AI and LLM Concepts
A developer calls SNOWFLAKE.CORTEX.SUMMARIZE() on a long support ticket and needs the output capped at approximately 50 words for a dashboard tile. SUMMARIZE() does not accept a length or max_tokens parameter. What is the BEST approach to enforce this length constraint?
COMPLETE() accepts a free-form prompt where length and style instructions can be embedded directly, giving control over output length that SUMMARIZE() intentionally does not expose.
Question 4 of 6 · Snowflake Gen AI and LLM Concepts
For SNOWFLAKE.CORTEX.COMPLETE(), which statement about tokens and context windows is CORRECT?
The model's context window is a shared token budget consumed jointly by the input prompt and the generated output; exceeding it causes truncation or an error, regardless of how the tokens are split between input and output.
Question 5 of 6 · Snowflake Gen AI and LLM Concepts
A team combines Cortex Search with COMPLETE() to build a RAG chatbot over internal documentation, believing this will eliminate hallucinations entirely. During UAT, the model still occasionally invents details not present in the retrieved documents. What is the MOST accurate explanation of this behavior for the exam?
RAG grounds generation in retrieved context, which reduces hallucination risk, but the LLM's generation step remains probabilistic and can still fabricate details not actually present in the retrieved documents — RAG mitigates, it does not eliminate, hallucination.
Question 6 of 6 · Snowflake Gen AI and LLM Concepts
A company needs to score the sentiment of 20 million customer reviews per day for a cost-sensitive batch analytics job, with no need for custom categories or explanations. Which Cortex function is the MOST cost- and performance-efficient choice for this specific task?
SENTIMENT() is a purpose-built, lightweight Cortex function optimized specifically for sentiment scoring, making it more cost- and latency-efficient at scale than invoking a general-purpose LLM for the same narrow, well-defined task.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →