grok14ENGINEERING FIELD SCHOOL
Build / Week 7

LLM application engineering

Choose models and prompts through task evidence, validate outputs, and design for uncertainty.

5 lesson sections16-hour study & practice planModule 06 or equivalent experience

See the system

Choose models and prompts through task evidence, validate outputs, and design for uncertainty.

The application boundary owns identity, authorization, action state, and budgets. Data retrieval supplies authorized evidence. Model generation proposes answers/actions. Trusted integration code enforces approval and records writes. Arrows show logical responsibility, not a cloud network specification.
Governed support copilot: logical architecture. The application boundary owns identity, authorization, action state, and budgets. Data retrieval supplies authorized evidence. Model generation proposes answers/actions. Trusted integration code enforces approval and records writes. Arrows show logical responsibility, not a cloud network specification.
Module 07 / Lesson 01

Understand the model boundary

A language model generates outputs conditioned on its input and learned parameters. It is not an authoritative database or an authorization engine. Context limits, sampling behavior, provider changes, and ambiguous instructions affect outputs.

Make the task small and explicit. Supply the necessary evidence and expected response shape. Decide which claims require external verification and which tasks can tolerate variation. A fluent response is not evidence that a fact is true.

Apply the ideaList which parts of a ticket summary can be verified directly against the source ticket.
Module 07 / Lesson 02

Use structured output carefully

Structured generation can make outputs easier to consume, but syntax and domain validity are different. A valid JSON object can still contain a wrong customer ID or an unauthorized action. Validate types, allowed values, references, and business rules before use.

Define behavior for refusal, missing fields, truncation, and invalid output. Avoid endless “repair” loops. A bounded retry or a clear fallback is easier to operate and evaluate.

Apply the ideaDesign a summary schema with required fields and a separate uncertainty field.
Module 07 / Lesson 03

Treat prompts as versioned application code

A prompt encodes task instructions, output requirements, and assumptions. Version it with the application configuration and evaluation results. Changes that look stylistic can alter task behavior.

Separate trusted instructions from untrusted source material. Delimiters help organization but do not by themselves prevent prompt injection. Explain when the assistant should abstain and ensure consequential actions remain controlled by trusted code.

Apply the ideaWrite a prompt change and a regression case that could reveal its unintended effect.
Module 07 / Lesson 04

Select with a representative workload

Compare candidate models or configurations on the same held-out cases. Measure correctness, refusal/abstention behavior, latency, and cost. Include difficult and unsupported cases, not only examples that demonstrate the intended feature.

The largest model is not automatically the best operational choice. Verify model availability, data handling, supported interfaces, and limits in the chosen environment. Retain a rule or conventional software baseline to test whether model complexity is justified.

Apply the ideaCreate a five-row comparison table with quality, cost, latency, limitations, and availability.
Module 07 / Lesson 05

Choose the right adaptation strategy

Prompting is useful for clear tasks with adequate context. Retrieval supplies changing external knowledge. Tools perform bounded actions or queries. Fine-tuning may adapt behavior, but introduces data, evaluation, and maintenance requirements; it does not automatically make private knowledge current.

Use the smallest intervention that addresses a measured failure. The posting mentions DBRX; treat it as a model example to understand, while choosing supported models based on current access and task evidence.

Apply the ideaClassify three failures as missing knowledge, wrong behavior, or missing integration, and choose a response.

Worked scenario

A ticket classifier returns valid JSON with category “urgent-refund,” but the application only supports billing, technical, and account. Schema validation should reject the unsupported category rather than silently introducing a new routing path. Compare a rules baseline with the model on ambiguous cases.

A small example

This example isolates one concept. Read its boundary conditions before adapting it to an application.

python
ALLOWED = {"billing", "technical", "account"}

def validate_category(result):
    if not isinstance(result, dict):
        raise ValueError("Expected an object")
    if result.get("category") not in ALLOWED:
        raise ValueError("Unsupported category")
    return result["category"]

print(validate_category({"category": "billing"}))

Practical assignment

This is a practical design or implementation assignment. Use synthetic data. Where managed services are required, verify account access, costs, supported features, and cleanup before provisioning.
  1. Define a classification or summarization contract.
  2. Create normal, ambiguous, and unsupported fixtures.
  3. Implement a deterministic baseline.
  4. Design or run a budgeted model comparison if credentials are available.
  5. Validate outputs and record failure behavior.
  6. Label live measurements separately from simulated results.

What to submit

Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.

Review dimensionSubmission evidence
CorrectnessShow the expected behavior and a meaningful counterexample.
ReproducibilityState setup, inputs, versions, and what was actually executed.
Delivery judgmentExplain the client impact, alternative, and unresolved assumption.
Operational boundaryIdentify permissions, failure behavior, and any resource cleanup.

Knowledge check

1. Does valid JSON prove domain correctness?
2. What is the best reason to add retrieval?
3. What should a prompt change trigger?

Answer guide
  1. No. Identifiers, values, claims, and permissions still need verification.
  2. External knowledge must be supplied or updated. Retrieval supplies evidence but still requires quality and access controls.
  3. Relevant regression evaluation. Small wording changes can change application behavior.

References & next step

Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.

Open the official reference library

Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.