grok14ENGINEERING FIELD SCHOOL
Evaluate / Week 11

MLflow evaluation & tracing

Build a repeatable evaluation loop that connects application traces, human judgment, and release decisions.

5 lesson sections16-hour study & practice planModule 10 or equivalent experience

See the system

Build a repeatable evaluation loop that connects application traces, human judgment, and release decisions.

A fixed case set exercises a versioned candidate. Inspect scores and traces, apply critical release gates, and improve the measured failure before rerunning the candidate.
The evaluation improvement loop. A fixed case set exercises a versioned candidate. Inspect scores and traces, apply critical release gates, and improve the measured failure before rerunning the candidate.
Module 11 / Lesson 01

Trace the system, not only the model

A trace represents one application request; spans describe important operations within it. Capture retrieval, model calls, tool validation, approvals, and integration outcomes with safe metadata. This reveals whether failures originate in missing evidence, wrong reasoning, or a broken downstream service.

MLflow 3 provides a current workflow for GenAI tracing and evaluation. Distinguish open-source/local capabilities from managed workspace features. Do not expose sensitive payloads just to improve debugging: log the minimum necessary and control trace access.

Apply the ideaDesign span names and safe attributes for a request that proposes a ticket update.
Module 11 / Lesson 02

Build a representative evaluation set

Start from the actual task distribution and failure risks. Include normal requests, ambiguous questions, missing evidence, access-denied cases, stale content, and tool failures. Keep reference answers or expected outcomes with clear provenance and review ownership.

Separate development examples from held-out evaluation. Duplicated or near-identical cases can inflate confidence. Synthetic cases are useful for edge conditions but do not establish the operational distribution; label their origin.

Apply the ideaCreate ten labeled cases spanning at least four different failure categories.
Module 11 / Lesson 03

Match scorers to the claim

Deterministic checks work well for schema validity, known identifiers, access rules, and exact tool outcomes. Human reviewers can assess nuanced task correctness. LLM judges can assist at scale, but are fallible and sensitive to instructions, evidence, and model changes.

Calibrate a judge against human-labeled examples and inspect disagreement. Do not treat a single aggregate score as universal truth. Report per-category results, sample size, and important failures.

Apply the ideaChoose a scorer for citation existence, policy correctness, and unauthorized write prevention.
Module 11 / Lesson 04

Compare versions under the same conditions

Keep dataset version, prompt version, model identifier, retrieval configuration, code revision, and relevant environment together. Evaluate candidate and baseline on the same workload. Otherwise a change in cases may look like an application improvement.

Look beyond averages: a higher overall score can conceal a regression for one department or a catastrophic access failure. Separate critical gates from weighted quality metrics. Set the release policy before tuning to the final benchmark.

Apply the ideaDefine a gate with a quality threshold and a separate mandatory access-control suite.
Module 11 / Lesson 05

Use production feedback responsibly

User feedback, failed tasks, and incident traces can reveal gaps in the offline set. Curate those cases, remove unnecessary personal data, and assign review ownership before adding them. A thumbs-up is a useful signal but not necessarily a correctness label.

Online monitoring features may be preview or environment-dependent. Verify support and costs. Reuse relevant offline metrics when meaningful, while recognizing that production lacks a reference answer for every request.

Apply the ideaWrite a feedback-to-evaluation workflow including review, redaction, versioning, and regression testing.

Worked scenario

A candidate improves answer scores from 82/100 to 88/100 cases but leaks a restricted policy in one test. The release fails the critical access gate even though the average quality improved. Investigate that trace, fix the boundary, and rerun the affected suite before comparing quality again.

Practical assignment

Use the downloadable local reference lab where indicated. It uses synthetic data and Python’s standard library. The managed Databricks extension requires separate workspace verification.
  1. Create a versioned evaluation dataset with provenance.
  2. Run deterministic checks in the local reference lab.
  3. Design MLflow spans and metrics for a managed implementation.
  4. Compare two controlled configurations on the same cases.
  5. Report category scores, latency, cost assumptions, and failures.
  6. Distinguish executed local tests from unexecuted managed evaluation.

What to submit

Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.

Review dimensionSubmission evidence
CorrectnessShow the expected behavior and a meaningful counterexample.
ReproducibilityState setup, inputs, versions, and what was actually executed.
Delivery judgmentExplain the client impact, alternative, and unresolved assumption.
Operational boundaryIdentify permissions, failure behavior, and any resource cleanup.

Knowledge check

1. What should a trace show?
2. Can a higher average override an unauthorized-access failure?
3. How should LLM judges be used?

Answer guide
  1. The relevant application steps and safe diagnostic metadata. The path explains whether retrieval, generation, authorization, or an integration failed.
  2. No. Critical boundaries must be separate release gates.
  3. With calibration and disagreement review. Judges can assist but may be biased or wrong.

References & next step

Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.

Open the official reference library

Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.