grok14ENGINEERING FIELD SCHOOL
Build / Week 6

Orchestration & streaming

Coordinate data work, handle late events, and recover interrupted pipelines with evidence of correctness.

5 lesson sections16-hour study & practice planModule 05 or equivalent experience

See the system

Coordinate data work, handle late events, and recover interrupted pipelines with evidence of correctness.

Source records pass through validation and curated tables. Transformations establish usable entities before search indexing. Reconciliation checks completeness, freshness, updates, and deletions.
From source records to reliable evidence. Source records pass through validation and curated tables. Transformations establish usable entities before search indexing. Reconciliation checks completeness, freshness, updates, and deletions.
Module 06 / Lesson 01

Model dependencies and outcomes

An orchestrator schedules work and tracks dependencies; it does not make incorrect transformations correct. Define each task’s input, output, retry semantics, and success criteria. Downstream consumers need trustworthy data, not just a green job icon.

Use a dependency graph to distinguish independent tasks from those requiring a completed source. Keep resource cleanup and failure notification in the design, and avoid a single opaque task that makes recovery unnecessarily broad.

Apply the ideaDraw ingest, validate, transform, index, and evaluate tasks with their dependencies.
Module 06 / Lesson 02

Choose a transformation and orchestration approach

Managed declarative pipelines, scheduled jobs, Airflow, and dbt solve overlapping but different problems. Choose around the existing environment, transformations, dependencies, operational ownership, and verified capabilities.

Do not add all tools to demonstrate familiarity. A simple scheduled job may be enough. When teaching a managed feature, distinguish what the platform handles from what the application still must specify: quality rules, source semantics, permissions, and recovery decisions.

Apply the ideaWrite a short decision record selecting one approach for the support corpus.
Module 06 / Lesson 03

Distinguish event time from processing time

An event may occur at 10:00 and arrive at 10:07. Event time describes the business event; processing time describes when the system handles it. Late data can change a previously calculated result.

Watermarks help bound retained state under the engine’s documented semantics. They are not a promise that all late events will be incorporated. Define acceptable lateness, update behavior, and a route for records outside the normal window.

Apply the ideaCreate an example with out-of-order events and state what the final aggregate should mean.
Module 06 / Lesson 04

Recover with checkpoints and replay

Checkpoints persist progress and state for supported processing patterns. They do not remove the need to understand source retention, sink transactions, and duplicate delivery. End-to-end guarantees depend on the entire path.

Replay and backfill should have an explicit range, versioned transformation, and reconciliation. Test the boundary where a sink committed but the orchestrator did not record success. Idempotent output semantics make this failure recoverable.

Apply the ideaWrite a backfill plan specifying source range, target behavior, and validation.
Module 06 / Lesson 05

Measure freshness and operational health

A pipeline can run successfully while publishing old data. Measure the age of source events and the age of the newest usable downstream record. Define who responds to missed freshness targets.

Track rejected records, lag, retries, and processing duration. Route invalid data with enough context to investigate while respecting data access rules. Alerts should tell an operator what condition matters and where to begin.

Apply the ideaDefine an alert for stale policy search that avoids paging on every harmless retry.

Worked scenario

The ingestion job is green, but the search index has not synchronized the last three policy revisions. The user-facing freshness indicator should follow the usable indexed content, not only the upstream job. Add a version/freshness reconciliation between source, transformed table, and index.

Practical assignment

This is a practical design or implementation assignment. Use synthetic data. Where managed services are required, verify account access, costs, supported features, and cleanup before provisioning.
  1. Create duplicate and late synthetic events.
  2. Define a processing graph and source/sink guarantees.
  3. Simulate failure between write and checkpoint.
  4. Replay a bounded range.
  5. Verify no unintended logical duplicates.
  6. Document freshness checks and recovery instructions.

What to submit

Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.

Review dimensionSubmission evidence
CorrectnessShow the expected behavior and a meaningful counterexample.
ReproducibilityState setup, inputs, versions, and what was actually executed.
Delivery judgmentExplain the client impact, alternative, and unresolved assumption.
Operational boundaryIdentify permissions, failure behavior, and any resource cleanup.

Knowledge check

1. Does a green job prove fresh search results?
2. What is event time?
3. Where are exactly-once guarantees established?

Answer guide
  1. No. Downstream synchronization may fail or lag after the source job completes.
  2. When a business event occurred. It differs from when processing happens.
  3. Across documented source, processing, and sink semantics. A component guarantee cannot be assumed for the entire workflow.

References & next step

Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.

Open the official reference library

Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.