Free IBM Certified watsonx Generative AI Engineer - Associate practice — 6 questions on Deployment, with explanations. No sign-up.
Full 12-question mixed test →
Question 1 of 6 · Deployment
An engineering team builds a custom RAG orchestration in Python that performs vector search, formats a prompt, and calls a foundation model, all in one workflow. They need to expose this entire workflow as a single deployable scoring endpoint in a watsonx.ai deployment space. Which asset type should they create and deploy?
An AI service lets engineers package arbitrary Python code (retrieval logic, prompt assembly, and model invocation) and deploy it as a single REST-callable scoring endpoint within a deployment space.
Question 2 of 6 · Deployment
A company needs to score 5 million records overnight against a tuned foundation model with no real-time requirement, and wants to minimize cost and operational overhead. Which watsonx.ai deployment approach is BEST?
Batch deployment jobs process large datasets asynchronously against a deployed asset, which is more cost-efficient and operationally simpler than making millions of individual synchronous calls.
Question 3 of 6 · Deployment
You are deploying a custom fine-tuned foundation model (not one of the shared base models) to an online deployment in a watsonx.ai deployment space. Which configuration parameter must you explicitly specify to ensure adequate compute/GPU resources are allocated for serving?
Custom/tuned foundation models are not served on shared multi-tenant infrastructure like the built-in base models, so a hardware_spec must be defined in the deployment configuration to reserve the compute/GPU tier needed to host and serve the model.
Question 4 of 6 · Deployment
After deploying a foundation-model-based classification AI service, the team notices performance degrading over several months as customer language patterns evolve, and continuously labeled ground truth is not readily available. Which monitor should be configured in watsonx.governance/OpenScale to detect this issue?
A drift monitor compares production input/prediction distributions against the training baseline without requiring continuously labeled ground truth, making it the appropriate way to detect this kind of gradual data/concept shift.
Question 5 of 6 · Deployment
What is the key architectural difference between an 'online' deployment and a 'batch' deployment in watsonx.ai Deployment Spaces?
Online deployments provide a synchronous REST scoring endpoint that returns a result per request immediately, while batch deployments submit asynchronous jobs that process an entire dataset and write results once the job completes.
Question 6 of 6 · Deployment
A production AI service deployment (v1) is serving live consumers. The team wants to test updated prompt orchestration logic without disrupting current consumers, then gradually shift traffic to the new logic. What is the recommended watsonx.ai lifecycle approach?
watsonx.ai supports asset revisions/versions, allowing engineers to create and validate a new deployment version alongside the existing one and control a gradual rollout without breaking active consumers.
Ready for the real thing?
The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.