See the system
Connect answers to source evidence and diagnose retrieval separately from generation.
Preserve document identity and structure
Parsing should preserve meaningful structure: headings, tables, source identifiers, revisions, and effective dates. Bad extraction turns an authoritative document into unreliable model context. Retain a source-to-chunk relationship so citations can be resolved and updates can be propagated.
Chunk size is a hypothesis to test. Tiny chunks may lose conditions and exceptions; huge chunks may dilute relevance and increase cost. Start with semantic boundaries and evaluate against actual questions.
Compare lexical and semantic retrieval
Lexical retrieval matches terms and can work well for exact identifiers. Vector retrieval compares embeddings and can find related language without identical words. Hybrid approaches combine signals; reranking can improve ordering at additional latency and cost.
No method dominates every workload. Measure retrieval on representative questions, including codes, names, paraphrases, and conflicting documents. Use a simple baseline before adding more moving parts.
Assemble trustworthy context
Filter by authorization before content enters model context. Prefer current, relevant evidence and make conflicting sources visible. The answer should point to the exact source revision used. A citation that merely points to a vaguely related document is insufficient.
Define abstention behavior when evidence is absent or contradictory. Do not fill gaps with plausible policy details. A useful assistant can say what it could establish and what requires a reviewer.
Evaluate retrieval independently
Retrieval evaluation asks whether the required evidence appears in the returned results. Recall@k measures the fraction of relevant items retrieved within k under the dataset’s labeling rules. Answer evaluation asks whether the final response is correct and supported.
A generator cannot reliably repair missing evidence. Conversely, a retriever can return the right document while the model ignores a critical exception. Keep labels, denominators, and failure categories separate so fixes target the right component.
Maintain freshness and deletion
A changing corpus needs a lifecycle: detect a change, parse it, update chunks and embeddings, synchronize indexes, and verify removal of obsolete content. Cache invalidation belongs to this lifecycle too.
Record the version of evidence associated with an answer. Test a revoked document and an updated policy. If removal is delayed, the application needs an explicit behavior rather than silently continuing to serve old information.
Worked scenario
An assistant answers a leave-policy question from an obsolete revision because it contains the exact wording of the query. A relevance score alone cannot determine authority. Add revision/effective-date handling and a test where the old document is a stronger keyword match than the current policy.
Practical assignment
- Run the local lexical retrieval reference.
- Inspect source IDs and department restrictions.
- Add a contradictory or outdated fixture.
- Create answerable and unanswerable questions.
- Measure retrieval against expected document IDs.
- Write a change/delete propagation checklist for a managed index.
What to submit
Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.
| Review dimension | Submission evidence |
|---|---|
| Correctness | Show the expected behavior and a meaningful counterexample. |
| Reproducibility | State setup, inputs, versions, and what was actually executed. |
| Delivery judgment | Explain the client impact, alternative, and unresolved assumption. |
| Operational boundary | Identify permissions, failure behavior, and any resource cleanup. |
Knowledge check
Answer guide
- Whether required evidence is returned. Final answer quality and operational effects require separate evaluation.
- Abstain or escalate appropriately. The system should respect the evidence boundary.
- To explain which evidence supported an answer. Versioned evidence supports reproducibility and audit.
References & next step
Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.
Open the official reference library
Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.