grok14ENGINEERING FIELD SCHOOL
Build / Week 8

Retrieval-augmented generation

Connect answers to source evidence and diagnose retrieval separately from generation.

5 lesson sections16-hour study & practice planModule 07 or equivalent experience

See the system

Connect answers to source evidence and diagnose retrieval separately from generation.

Trusted identity scopes retrieval before evidence enters model context. Generation uses the selected revisions, and the answer links supported claims to sources or abstains.
A retrieval request with an access boundary. Trusted identity scopes retrieval before evidence enters model context. Generation uses the selected revisions, and the answer links supported claims to sources or abstains.
Module 08 / Lesson 01

Preserve document identity and structure

Parsing should preserve meaningful structure: headings, tables, source identifiers, revisions, and effective dates. Bad extraction turns an authoritative document into unreliable model context. Retain a source-to-chunk relationship so citations can be resolved and updates can be propagated.

Chunk size is a hypothesis to test. Tiny chunks may lose conditions and exceptions; huge chunks may dilute relevance and increase cost. Start with semantic boundaries and evaluate against actual questions.

Apply the ideaSplit a policy containing an exception and explain how you preserve the exception’s context.
Module 08 / Lesson 02

Compare lexical and semantic retrieval

Lexical retrieval matches terms and can work well for exact identifiers. Vector retrieval compares embeddings and can find related language without identical words. Hybrid approaches combine signals; reranking can improve ordering at additional latency and cost.

No method dominates every workload. Measure retrieval on representative questions, including codes, names, paraphrases, and conflicting documents. Use a simple baseline before adding more moving parts.

Apply the ideaCreate one exact-identifier question and one paraphrase question and predict which approach helps.
Module 08 / Lesson 03

Assemble trustworthy context

Filter by authorization before content enters model context. Prefer current, relevant evidence and make conflicting sources visible. The answer should point to the exact source revision used. A citation that merely points to a vaguely related document is insufficient.

Define abstention behavior when evidence is absent or contradictory. Do not fill gaps with plausible policy details. A useful assistant can say what it could establish and what requires a reviewer.

Apply the ideaWrite an answer format with cited claims and a clearly labeled unresolved question.
Module 08 / Lesson 04

Evaluate retrieval independently

Retrieval evaluation asks whether the required evidence appears in the returned results. Recall@k measures the fraction of relevant items retrieved within k under the dataset’s labeling rules. Answer evaluation asks whether the final response is correct and supported.

A generator cannot reliably repair missing evidence. Conversely, a retriever can return the right document while the model ignores a critical exception. Keep labels, denominators, and failure categories separate so fixes target the right component.

Apply the ideaLabel a small query set with expected document IDs and compute retrieval success independently.
Module 08 / Lesson 05

Maintain freshness and deletion

A changing corpus needs a lifecycle: detect a change, parse it, update chunks and embeddings, synchronize indexes, and verify removal of obsolete content. Cache invalidation belongs to this lifecycle too.

Record the version of evidence associated with an answer. Test a revoked document and an updated policy. If removal is delayed, the application needs an explicit behavior rather than silently continuing to serve old information.

Apply the ideaDefine a test proving that a removed policy no longer appears in results, answers, or caches.

Worked scenario

An assistant answers a leave-policy question from an obsolete revision because it contains the exact wording of the query. A relevance score alone cannot determine authority. Add revision/effective-date handling and a test where the old document is a stronger keyword match than the current policy.

Practical assignment

Use the downloadable local reference lab where indicated. It uses synthetic data and Python’s standard library. The managed Databricks extension requires separate workspace verification.
  1. Run the local lexical retrieval reference.
  2. Inspect source IDs and department restrictions.
  3. Add a contradictory or outdated fixture.
  4. Create answerable and unanswerable questions.
  5. Measure retrieval against expected document IDs.
  6. Write a change/delete propagation checklist for a managed index.

What to submit

Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.

Review dimensionSubmission evidence
CorrectnessShow the expected behavior and a meaningful counterexample.
ReproducibilityState setup, inputs, versions, and what was actually executed.
Delivery judgmentExplain the client impact, alternative, and unresolved assumption.
Operational boundaryIdentify permissions, failure behavior, and any resource cleanup.

Knowledge check

1. What does retrieval evaluation establish?
2. What should happen when evidence is missing?
3. Why retain document revisions?

Answer guide
  1. Whether required evidence is returned. Final answer quality and operational effects require separate evaluation.
  2. Abstain or escalate appropriately. The system should respect the evidence boundary.
  3. To explain which evidence supported an answer. Versioned evidence supports reproducibility and audit.

References & next step

Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.

Open the official reference library

Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.