Unity Catalog
Unity Catalog is the governance layer of Azure Databricks: one place that decides who may read and write which tables, volumes and storage paths. The Azure Enterprise Baseline provides its regional metastore root, and the Azure Data and ETL Blueprint builds one governed catalog per stage on top of it.
What it does
A metastore is the top-level container, one per region, shared by the workspaces attached to it. Below it, catalogs hold schemas, which hold tables and volumes (governed file areas). Unity Catalog reaches storage only through a storage credential, on Azure a Databricks access connector with a managed identity; an external location binds a credential to a storage path. Grants give principals privileges such as SELECT, MODIFY, READ_FILES and WRITE_FILES on these objects. A catalog's managed storage holds its managed tables, and path-based access to it is blocked.
How BuiltForProd uses it
In the Baseline. The uc-metastore unit in acme-core-artifacts creates the metastore root of the region: the ADLS Gen2 account stacmeeus2ucmetastore with the container metastore, zone-redundant, Entra-only (shared keys off, OAuth by default), 30-day soft delete, protected from deletion in code. Its access connector <prefix>-uc-dbac holds Storage Blob Data Contributor on the account and is always admitted by a resource instance rule; until allowed_subnet_ids lists the stages' Databricks subnets, the account's network default stays Allow, with Entra authentication still required. Reads and writes are logged to the audit workspace. The metastore object is created once per region in the Databricks account against this root, and every stage workspace of the region attaches to it. Storage costs ~$0.02/GB-month; the connector is free.
In each stage. The databricks-catalog unit builds the stage's governed objects on the data lake:
| Object | Name and setting |
|---|---|
| Storage credential | <prefix>-lake, the lake's access connector <prefix>-dbac; bound to this workspace only |
| External locations | <prefix>-raw, -processed, -curated, -managed, one per container; bound to this workspace |
| Catalog | acme_<stage>, managed storage in the managed container; bound to this workspace |
| Schemas | raw, processed, curated |
| Volume | scripts, external, on processed/scripts/: the job's PySpark script |
| Principal | On the catalog | Zone paths |
|---|---|---|
| ETL trigger identity (the job's run-as) | USE_CATALOG, USE_SCHEMA, CREATE_TABLE, MODIFY, SELECT, READ_VOLUME, WRITE_VOLUME | Read raw; read and write processed |
ACME_LeadETLEngineers | ALL_PRIVILEGES | Read and write every zone |
ACME_ETLEngineers | USE_CATALOG, USE_SCHEMA, SELECT, READ_VOLUME | Read curated |
The two Entra groups are addressed by display name through the Databricks account's automatic identity management; grant_entra_groups = false leaves them out.
Why the zones are external locations. The pipeline reads raw/input/*.json and writes processed/output/ by path, and Unity Catalog blocks path access to managed storage. So the zone containers are external locations with READ_FILES and WRITE_FILES grants, the schemas inherit the catalog's managed location in managed, and the Delta table the job writes in catalog mode lands there as a managed table.
Terms you will see
| Term | Meaning |
|---|---|
| Metastore | The regional top of the hierarchy; its root storage is in the Baseline. |
| Storage credential | The access connector identity Unity Catalog uses to reach storage. |
| External location | A storage path governed through a credential, here one per lake container. |
| Managed storage | Where managed tables live; not reachable by path. |
| Workspace binding | ISOLATED mode: the object is usable from this stage's workspace only. |
Where to read more
- Azure Data and ETL Blueprint overview and Azure Enterprise Baseline overview.
- Azure Databricks for the workspace and the job.
- Azure Data Lake Storage for the zones.