TechNuggets Academy

Building and Deploying Gen AI Applications

Free SnowPro Specialty: Gen AI practice — 6 questions on Building and Deploying Gen AI Applications, with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Building and Deploying Gen AI Applications
A team must deploy a custom fine-tuned LLM for real-time inference using Snowpark Container Services. Profiling shows the model requires more GPU memory than a single NVIDIA A10G GPU (24 GB) can provide, and the team wants to avoid quantizing the model. Which compute pool instance family should they select to run the model across multiple GPUs on one node?
GPU_NV_M provisions multiple GPUs per node (unlike GPU_NV_S, which has a single A10G with 24 GB), letting an inference framework shard the model's weights across GPUs to exceed the memory ceiling of one GPU without quantization.
Question 2 of 6 · Building and Deploying Gen AI Applications
A container running in Snowpark Container Services must call an external third-party embedding API over HTTPS from inside the running service. Which combination of steps correctly enables this outbound call?
Outbound network access from an SPCS service to an external endpoint requires a network rule defining the allowed host/port, wrapped in an External Access Integration (EAI), which is then referenced in the service specification so the container is permitted to make the call.
Question 3 of 6 · Building and Deploying Gen AI Applications
In a Snowpark Container Services service specification YAML, which field tells Snowflake to check that a container has finished starting up and is able to accept traffic before routing requests to it?
readinessProbe defines a health check (e.g., an HTTP GET on a given path/port) that Snowflake uses to determine when a container instance is ready to receive traffic, preventing requests from being routed to a container that is still initializing.
Question 4 of 6 · Building and Deploying Gen AI Applications
A team needs to generate embeddings for 5 million documents once, using a GPU-accelerated container image, and wants the compute pool to release the node automatically as soon as processing completes rather than remaining active for future requests. Which SPCS execution approach fits this requirement?
EXECUTE JOB SERVICE runs a container as a job-type service designed for run-to-completion batch workloads; once the container process exits, the job service terminates and its compute pool node can be reclaimed automatically, matching the one-time batch requirement.
Question 5 of 6 · Building and Deploying Gen AI Applications
A Streamlit in Snowflake app needs to display real-time predictions produced by a custom PyTorch model deployed as a long-running SPCS service. What is the correct, supported way for the Streamlit app's Python code to obtain predictions from that service?
Snowflake supports creating a service function (a UDF defined with SERVICE and ENDPOINT) that routes SQL calls to a specific endpoint of a running SPCS service. A Streamlit in Snowflake app can call this UDF via a SQL query (e.g., through the Snowpark session), which is the officially supported in-account integration pattern.
Question 6 of 6 · Building and Deploying Gen AI Applications
A pipeline must run a nightly GPU batch-inference job in SPCS after new data lands in a table, and then update a results table once the job finishes. Which orchestration design correctly implements this within Snowflake?
Snowflake Tasks provide native scheduling (CRON or dependency-based) for stored procedures. A stored procedure can call EXECUTE JOB SERVICE to launch the batch container job and, once it completes, run follow-on SQL to update the results table — this is the standard orchestration pattern for scheduled SPCS batch jobs.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →