See the system
Measure the complete task and compare architecture choices under declared, repeatable conditions.
Break latency into components
Measure authentication, retrieval, reranking, generation, tool calls, and serialization. End-to-end latency is what the user experiences; a fast model call may hide a slow retrieval path.
Percentiles describe the distribution. p95 is a threshold at or below which roughly 95% of observed requests fall under the stated calculation. Report sample size and load. A percentile from a tiny or unrealistic sample may be misleading.
Calculate cost per useful outcome
Model token charges are only part of cost. Include relevant search, compute, storage, observability, and integration usage. A cheap answer that requires extensive correction may have a high cost per completed task.
Record pricing source and date, currency, workload assumptions, and resource utilization. Separate actual measured spend from estimates. Do not multiply a small benchmark indefinitely without considering concurrency, quotas, and infrastructure scaling.
Cache with identity and freshness in mind
Caching can reduce latency and repeated computation. Decide what is cached: embeddings, retrieval results, or final answers. Each has different invalidation and access implications.
An answer cache may need user/entitlement scope, document version, prompt version, and relevant configuration in its identity. Do not share a department-specific answer across unauthorized users. Test policy updates and revocation before relying on the cache.
Compare candidate configurations fairly
Use the same evaluation set and equivalent operating conditions. Change one important factor at a time when you need causal understanding. Record warmup, concurrency, input sizes, region, model configuration, and repetitions.
A cost-quality tradeoff may justify routing simple requests to a smaller model, but routing itself needs evaluation and safe fallback behavior. Do not compare the best run of one candidate with an average run of another.
Document the decision and its limits
An architecture decision should state the chosen configuration, alternatives, evidence, constraints, and conditions that would trigger reconsideration. Prefer explicit tradeoffs to declaring one model “best.”
Monitor after release because traffic mix, data, providers, and costs can change. A configuration that works for short internal questions may fail under long multi-turn conversations or burst traffic.
Worked scenario
Configuration A is faster on average but fails many exact policy-code questions. Configuration B costs more per request but improves completed-task accuracy. Compare cost per successful task and category-specific errors before deciding. Do not allow access failures to be traded away for speed.
Practical assignment
- Define workload and success criteria.
- Collect local timing for a fixed fixture set.
- Write a current-price formula for a managed variant.
- Compare three proposed configurations honestly.
- Document cache authorization and invalidation.
- Submit a decision with measured and estimated fields clearly separated.
What to submit
Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.
| Review dimension | Submission evidence |
|---|---|
| Correctness | Show the expected behavior and a meaningful counterexample. |
| Reproducibility | State setup, inputs, versions, and what was actually executed. |
| Delivery judgment | Explain the client impact, alternative, and unresolved assumption. |
| Operational boundary | Identify permissions, failure behavior, and any resource cleanup. |
Knowledge check
Answer guide
- End-to-end task latency. Users experience the full application path.
- No. Cached answers can leak data unless scoped and invalidated appropriately.
- With dated prices and workload assumptions. Costs depend on usage, configuration, region, and current pricing.
References & next step
Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.
Open the official reference library
Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.