grok14ENGINEERING FIELD SCHOOL
Evaluate / Week 16

Quality, latency & cost

Measure the complete task and compare architecture choices under declared, repeatable conditions.

5 lesson sections16-hour study & practice planModule 15 or equivalent experience

See the system

Measure the complete task and compare architecture choices under declared, repeatable conditions.

A fixed case set exercises a versioned candidate. Inspect scores and traces, apply critical release gates, and improve the measured failure before rerunning the candidate.
The evaluation improvement loop. A fixed case set exercises a versioned candidate. Inspect scores and traces, apply critical release gates, and improve the measured failure before rerunning the candidate.
Module 16 / Lesson 01

Break latency into components

Measure authentication, retrieval, reranking, generation, tool calls, and serialization. End-to-end latency is what the user experiences; a fast model call may hide a slow retrieval path.

Percentiles describe the distribution. p95 is a threshold at or below which roughly 95% of observed requests fall under the stated calculation. Report sample size and load. A percentile from a tiny or unrealistic sample may be misleading.

Apply the ideaCreate a latency budget by component and identify which stage you would measure first.
Module 16 / Lesson 02

Calculate cost per useful outcome

Model token charges are only part of cost. Include relevant search, compute, storage, observability, and integration usage. A cheap answer that requires extensive correction may have a high cost per completed task.

Record pricing source and date, currency, workload assumptions, and resource utilization. Separate actual measured spend from estimates. Do not multiply a small benchmark indefinitely without considering concurrency, quotas, and infrastructure scaling.

Apply the ideaBuild an estimate formula separating variable model usage from persistent platform resources.
Module 16 / Lesson 03

Cache with identity and freshness in mind

Caching can reduce latency and repeated computation. Decide what is cached: embeddings, retrieval results, or final answers. Each has different invalidation and access implications.

An answer cache may need user/entitlement scope, document version, prompt version, and relevant configuration in its identity. Do not share a department-specific answer across unauthorized users. Test policy updates and revocation before relying on the cache.

Apply the ideaSpecify cache keys and invalidation triggers for a restricted policy assistant.
Module 16 / Lesson 04

Compare candidate configurations fairly

Use the same evaluation set and equivalent operating conditions. Change one important factor at a time when you need causal understanding. Record warmup, concurrency, input sizes, region, model configuration, and repetitions.

A cost-quality tradeoff may justify routing simple requests to a smaller model, but routing itself needs evaluation and safe fallback behavior. Do not compare the best run of one candidate with an average run of another.

Apply the ideaDesign a comparison of lexical retrieval, hybrid retrieval, and reranking with fixed cases.
Module 16 / Lesson 05

Document the decision and its limits

An architecture decision should state the chosen configuration, alternatives, evidence, constraints, and conditions that would trigger reconsideration. Prefer explicit tradeoffs to declaring one model “best.”

Monitor after release because traffic mix, data, providers, and costs can change. A configuration that works for short internal questions may fail under long multi-turn conversations or burst traffic.

Apply the ideaWrite a decision record that names a quality floor, latency target, and cost budget.

Worked scenario

Configuration A is faster on average but fails many exact policy-code questions. Configuration B costs more per request but improves completed-task accuracy. Compare cost per successful task and category-specific errors before deciding. Do not allow access failures to be traded away for speed.

Practical assignment

This is a practical design or implementation assignment. Use synthetic data. Where managed services are required, verify account access, costs, supported features, and cleanup before provisioning.
  1. Define workload and success criteria.
  2. Collect local timing for a fixed fixture set.
  3. Write a current-price formula for a managed variant.
  4. Compare three proposed configurations honestly.
  5. Document cache authorization and invalidation.
  6. Submit a decision with measured and estimated fields clearly separated.

What to submit

Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.

Review dimensionSubmission evidence
CorrectnessShow the expected behavior and a meaningful counterexample.
ReproducibilityState setup, inputs, versions, and what was actually executed.
Delivery judgmentExplain the client impact, alternative, and unresolved assumption.
Operational boundaryIdentify permissions, failure behavior, and any resource cleanup.

Knowledge check

1. Which latency matters to the user?
2. Can final-answer caches ignore user permissions?
3. How should estimated costs be presented?

Answer guide
  1. End-to-end task latency. Users experience the full application path.
  2. No. Cached answers can leak data unless scoped and invalidated appropriately.
  3. With dated prices and workload assumptions. Costs depend on usage, configuration, region, and current pricing.

References & next step

Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.

Open the official reference library

Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.