See the system
Build a repeatable evaluation loop that connects application traces, human judgment, and release decisions.
Trace the system, not only the model
A trace represents one application request; spans describe important operations within it. Capture retrieval, model calls, tool validation, approvals, and integration outcomes with safe metadata. This reveals whether failures originate in missing evidence, wrong reasoning, or a broken downstream service.
MLflow 3 provides a current workflow for GenAI tracing and evaluation. Distinguish open-source/local capabilities from managed workspace features. Do not expose sensitive payloads just to improve debugging: log the minimum necessary and control trace access.
Build a representative evaluation set
Start from the actual task distribution and failure risks. Include normal requests, ambiguous questions, missing evidence, access-denied cases, stale content, and tool failures. Keep reference answers or expected outcomes with clear provenance and review ownership.
Separate development examples from held-out evaluation. Duplicated or near-identical cases can inflate confidence. Synthetic cases are useful for edge conditions but do not establish the operational distribution; label their origin.
Match scorers to the claim
Deterministic checks work well for schema validity, known identifiers, access rules, and exact tool outcomes. Human reviewers can assess nuanced task correctness. LLM judges can assist at scale, but are fallible and sensitive to instructions, evidence, and model changes.
Calibrate a judge against human-labeled examples and inspect disagreement. Do not treat a single aggregate score as universal truth. Report per-category results, sample size, and important failures.
Compare versions under the same conditions
Keep dataset version, prompt version, model identifier, retrieval configuration, code revision, and relevant environment together. Evaluate candidate and baseline on the same workload. Otherwise a change in cases may look like an application improvement.
Look beyond averages: a higher overall score can conceal a regression for one department or a catastrophic access failure. Separate critical gates from weighted quality metrics. Set the release policy before tuning to the final benchmark.
Use production feedback responsibly
User feedback, failed tasks, and incident traces can reveal gaps in the offline set. Curate those cases, remove unnecessary personal data, and assign review ownership before adding them. A thumbs-up is a useful signal but not necessarily a correctness label.
Online monitoring features may be preview or environment-dependent. Verify support and costs. Reuse relevant offline metrics when meaningful, while recognizing that production lacks a reference answer for every request.
Worked scenario
A candidate improves answer scores from 82/100 to 88/100 cases but leaks a restricted policy in one test. The release fails the critical access gate even though the average quality improved. Investigate that trace, fix the boundary, and rerun the affected suite before comparing quality again.
Practical assignment
- Create a versioned evaluation dataset with provenance.
- Run deterministic checks in the local reference lab.
- Design MLflow spans and metrics for a managed implementation.
- Compare two controlled configurations on the same cases.
- Report category scores, latency, cost assumptions, and failures.
- Distinguish executed local tests from unexecuted managed evaluation.
What to submit
Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.
| Review dimension | Submission evidence |
|---|---|
| Correctness | Show the expected behavior and a meaningful counterexample. |
| Reproducibility | State setup, inputs, versions, and what was actually executed. |
| Delivery judgment | Explain the client impact, alternative, and unresolved assumption. |
| Operational boundary | Identify permissions, failure behavior, and any resource cleanup. |
Knowledge check
Answer guide
- The relevant application steps and safe diagnostic metadata. The path explains whether retrieval, generation, authorization, or an integration failed.
- No. Critical boundaries must be separate release gates.
- With calibration and disagreement review. Judges can assist but may be biased or wrong.
References & next step
Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.
Open the official reference library
Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.