grok14ENGINEERING FIELD SCHOOL
Build / Week 4

Cloud & Databricks foundations

Understand storage, compute, identity, and catalog boundaries before provisioning an AI application.

5 lesson sections16-hour study & practice planModule 03 or equivalent experience

See the system

Understand storage, compute, identity, and catalog boundaries before provisioning an AI application.

Authentication identifies the caller. Entitlements constrain data retrieval and action execution. Citations, cached responses, and logs must preserve the same access boundary.
Access is enforced beyond the interface. Authentication identifies the caller. Entitlements constrain data retrieval and action execution. Citations, cached responses, and logs must preserve the same access boundary.
Module 04 / Lesson 01

Separate storage, compute, and control

Cloud object storage holds durable files. Compute reads and transforms data. The platform’s control features coordinate resources, permissions, jobs, and metadata. These are related but distinct layers with different failure and cost characteristics.

Draw which identity accesses which resource through which network path. A successful notebook query does not prove that the eventual application service identity can access the same table, nor that an end user should receive every row.

Apply the ideaDraw the path from user to application to compute to data, labeling identities.
Module 04 / Lesson 02

Learn the workspace vocabulary

A workspace organizes development and operational resources. Catalogs and schemas organize governed data objects. Tables represent structured data; volumes support governed file access in suitable configurations. Compute options vary by workload and platform availability.

Learn concepts before memorizing console clicks. Document the exact cloud, region, account features, and runtime used for a lab. Platform menus and feature names can change while the underlying access and lifecycle concepts remain useful.

Apply the ideaCreate an environment inventory with workspace, cloud, region, and required features.
Module 04 / Lesson 03

Design least-privilege access

A person, application, and deployment pipeline should not all share a powerful credential. Assign distinct identities and the minimum documented permissions for their tasks. Verify both allowed and denied operations.

Unity Catalog governance is part of the data boundary, but the application still needs a deliberate strategy for representing user entitlements. Avoid assuming the application’s broad service access automatically becomes per-user authorization.

Apply the ideaWrite an access matrix for analyst, application, pipeline, and reviewer roles.
Module 04 / Lesson 04

Keep environments and credentials separate

Development, test, and production should have explicit configuration and appropriate data boundaries. A name prefix alone does not isolate permissions. Review who can deploy, who can read secrets, and which systems an environment can reach.

Use supported identity flows and secret handling. Never put access tokens in notebooks, downloadable assets, or screenshots. Establish credential rotation and revocation expectations with the platform owner.

Apply the ideaExplain how a leaked development credential would be revoked and what it could access.
Module 04 / Lesson 05

Control lifecycle and cost

Compute, search resources, model calls, and storage can all contribute to cost. A notebook finishing does not imply every provisioned resource stopped billing. Before a lab, identify billable resources, expected usage, and cleanup responsibilities.

Free or trial environments may not support all required features. Offer a local conceptual exercise, then label managed-platform verification incomplete until it is actually performed. Reproducibility includes teardown, not just setup.

Apply the ideaCreate a preflight and teardown checklist with an owner for each resource.

Worked scenario

An analyst can query a table in a notebook, but the deployed app fails. The app runs under a separate identity with different grants. Diagnose the actual identity, target catalog/schema, documented privileges, and network path instead of copying the analyst’s personal credential into the service.

Practical assignment

This is a practical design or implementation assignment. Use synthetic data. Where managed services are required, verify account access, costs, supported features, and cleanup before provisioning.
  1. Choose one cloud and record environment availability.
  2. List required privileges before provisioning.
  3. Create a synthetic governed dataset where access permits.
  4. Verify one allowed and one denied operation per role.
  5. Capture the architecture and resource inventory.
  6. Run cleanup and record managed checks not performed.

What to submit

Submit the artifacts named above, a short explanation of your decisions, and evidence of the checks you performed. Distinguish measured results from estimates and designs from executed integrations.

Review dimensionSubmission evidence
CorrectnessShow the expected behavior and a meaningful counterexample.
ReproducibilityState setup, inputs, versions, and what was actually executed.
Delivery judgmentExplain the client impact, alternative, and unresolved assumption.
Operational boundaryIdentify permissions, failure behavior, and any resource cleanup.

Knowledge check

1. A notebook query succeeds. Does that prove the deployed app has access?
2. What should a lab state before provisioning?
3. Which credential design is preferable?

Answer guide
  1. No. The app may use a different identity, network path, and permissions.
  2. Account requirements, billable resources, and cleanup. Availability and charges depend on the selected environment.
  3. Distinct least-privilege identities. Separate identities support constrained permissions, audit, and revocation.

References & next step

Platform examples are environment-dependent. Start with the official documentation in the reference library and verify the exact cloud, region, privileges, and versions you use.

Open the official reference library

Editorial edition: 5 October 2026. The local reference lab is executed locally; this course does not claim a live Databricks deployment.