Azure Databricks
Azure Databricks is a managed Apache Spark platform run jointly by Microsoft and Databricks: the control plane is Databricks', the clusters are virtual machines in your subscription. The Azure Data and ETL Blueprint runs one workspace per stage and one job in it that transforms every batch of landed files.
What it does
A workspace is the unit of users, notebooks, jobs and clusters; its managed resource group holds the clusters' resources and the DBFS root storage. VNet injection places the clusters in two subnets of your own virtual network, and secure cluster connectivity gives cluster nodes no public IP: they call out to the control plane. A job runs tasks, such as a Python file, on a job cluster created for the run and deleted after it, or on serverless compute. Compute is billed as VM time plus Databricks units (DBUs).
How BuiltForProd uses it
The workspace. The databricks-workspace unit creates <prefix>-dbw in the stage subscription:
| Setting | Value |
|---|---|
| SKU | Premium, which includes Unity Catalog |
| Network | VNet injection into the spoke's snet-dbx-host and snet-dbx-container, whose network security groups the Baseline creates; no public IP on any cluster node |
| DBFS root | Storage account dbsacmeeus2<stage>, zone-redundant, with infrastructure (double) encryption |
| Front end | The workspace UI and REST API keep a public endpoint with Microsoft Entra authentication |
| Logs | Diagnostic categories to the stage application workspace |
| Protection | CanNotDelete lock on the resource group in prod (delete_locks) |
Cluster egress follows the hub's egress mode, and nothing in the workspace depends on it. With the hub firewall, the spoke route table sends traffic to the firewall, whose platform rules admit the AzureDatabricks service tag and the control plane's metastore and log storage. In the spoke NAT mode, a NAT gateway on both subnets carries it.
The job. The databricks-job unit defines <prefix>-etl, one Python task that runs /Volumes/acme_<stage>/processed/scripts/etl_transform.py, a PySpark script the code repository uploads to the lake.
| Setting | Value (stage file) |
|---|---|
| Compute | A job cluster per run on the latest LTS runtime: Standard_D4ds_v5 driver plus 2 workers, single-user access mode |
| Spot | Workers on spot with on-demand fallback in dev and staging, on-demand in prod; the driver is always on-demand |
| Limits | 3 concurrent runs, later runs queue; 60-minute timeout; one retry after a minute |
| Identity | Runs as the ETL trigger's managed identity, registered as a workspace service principal with Can Manage Run |
| People | ACME_LeadETLEngineers can manage the job, ACME_ETLEngineers can view it |
| Cost (comments) | VM time plus spark_serverless ( |
The ETL trigger reads the job ID from the contract vault secret etl--job-id and starts one run per batch with the files as the input_path parameter. The run appends partitioned Parquet under processed/output/, or, with use_catalog = true, writes the managed Delta table acme_<stage>.processed.etl_output. What the job may read and write is bounded by its Unity Catalog grants. The Databricks objects are managed with the databricks/databricks ~> 1.135 provider.
Terms you will see
| Term | Meaning |
|---|---|
| VNet injection | Clusters in the spoke's own Databricks subnets. |
| Secure cluster connectivity | Cluster nodes without public IPs, connecting out to the control plane. |
| Job cluster | A cluster created for one run and deleted after it. |
| DBU | Databricks unit, the compute charge on top of VM time. |
| Run as | The principal whose permissions a job run uses: the trigger identity. |
Where to read more
- Azure Data and ETL Blueprint overview and Azure Enterprise Baseline overview.
- Unity Catalog for the grants that bound the job.
- Azure Event Grid for how a run starts.