TechNuggets Academy

Retrieval-Augmented Generation (RAG)

Free IBM Certified watsonx Generative AI Engineer - Associate practice — 6 questions on Retrieval-Augmented Generation (RAG), with explanations. No sign-up. Full 12-question mixed test →

Question 1 of 6 · Retrieval-Augmented Generation (RAG)
An enterprise is building a RAG pipeline in watsonx.ai over a corpus of 50 million enterprise documents and requires sub-second retrieval latency at query time. Which vector store configuration should be used?
HNSW (Hierarchical Navigable Small World) is an approximate nearest neighbor index that scales efficiently to tens of millions of vectors while maintaining sub-second query latency. Milvus with HNSW is a supported, production-grade vector store pattern for large-scale watsonx.ai RAG deployments.
Question 2 of 6 · Retrieval-Augmented Generation (RAG)
A knowledge base contains technical PDFs with multi-page tables. During ingestion into a watsonx.ai RAG pipeline, an engineer notices retrieval quality drops because tables are being split across multiple chunks, losing row/column relationships. Which chunking approach should be configured to fix this?
Structure-aware (document-layout-aware) chunking recognizes tables as discrete structural elements and keeps them intact as single chunks, preserving row/column relationships needed for accurate retrieval and grounding.
Question 3 of 6 · Retrieval-Augmented Generation (RAG)
An engineer needs an embedding model in the watsonx.ai foundation model catalog that is specifically optimized for retrieval tasks in a RAG pipeline, rather than a general-purpose text generation model. Which model should be selected?
ibm/slate-125m-english-rtrvr is an IBM embedding model explicitly trained with a retriever objective (the 'rtrvr' suffix denotes retrieval-tuning) and is designed to produce dense vectors optimized for semantic search and retrieval in RAG pipelines.
Question 4 of 6 · Retrieval-Augmented Generation (RAG)
During evaluation of a RAG pipeline, the team finds that generated answers score well on relevance to the user's question, but auditors discover the model is stating facts that do not appear anywhere in the retrieved context. Which evaluation metric should the team prioritize improving?
Faithfulness measures whether the claims in the generated answer are actually entailed by (grounded in) the retrieved context, directly catching hallucinated content that is not supported by the source documents.
Question 5 of 6 · Retrieval-Augmented Generation (RAG)
A RAG system retrieves both the 2023 and 2026 versions of a company's expense reimbursement policy for a single query, and the generated response merges outdated and current rules, producing an incorrect answer. What is the BEST fix at the retrieval layer?
Metadata filtering (e.g., filtering on effective_date, version, or status=active fields attached to indexed chunks) excludes obsolete policy versions from the retrieved set before they ever reach the LLM, directly resolving the version-conflict grounding issue.
Question 6 of 6 · Retrieval-Augmented Generation (RAG)
In a watsonx.ai RAG architecture, why would an engineer configure hybrid search combining BM25 sparse retrieval with dense vector retrieval, rather than relying on dense retrieval alone?
Dense embeddings capture semantic similarity well but can underperform on exact lexical matches such as product codes, part numbers, or error IDs. BM25's sparse lexical scoring complements dense retrieval by catching these exact-term matches, improving overall recall — the standard justification for hybrid retrieval in production RAG pipelines.
Ready for the real thing?

The full course has two full-length practice tests, video lessons for every exam domain, hands-on labs and detailed answer explanations.

Start my full course on Udemy →