Skip to main content

Azure Databricks

Azure Databricks is a managed Apache Spark platform run jointly by Microsoft and Databricks: the control plane is Databricks', the clusters are virtual machines in your subscription. The Azure Data and ETL Blueprint runs one workspace per stage and one job in it that transforms every batch of landed files.

What it does​

A workspace is the unit of users, notebooks, jobs and clusters; its managed resource group holds the clusters' resources and the DBFS root storage. VNet injection places the clusters in two subnets of your own virtual network, and secure cluster connectivity gives cluster nodes no public IP: they call out to the control plane. A job runs tasks, such as a Python file, on a job cluster created for the run and deleted after it, or on serverless compute. Compute is billed as VM time plus Databricks units (DBUs).

How BuiltForProd uses it​

The workspace. The databricks-workspace unit creates <prefix>-dbw in the stage subscription:

SettingValue
SKUPremium, which includes Unity Catalog
NetworkVNet injection into the spoke's snet-dbx-host and snet-dbx-container, whose network security groups the Baseline creates; no public IP on any cluster node
DBFS rootStorage account dbsacmeeus2<stage>, zone-redundant, with infrastructure (double) encryption
Front endThe workspace UI and REST API keep a public endpoint with Microsoft Entra authentication
LogsDiagnostic categories to the stage application workspace
ProtectionCanNotDelete lock on the resource group in prod (delete_locks)

Cluster egress follows the hub's egress mode, and nothing in the workspace depends on it. With the hub firewall, the spoke route table sends traffic to the firewall, whose platform rules admit the AzureDatabricks service tag and the control plane's metastore and log storage. In the spoke NAT mode, a NAT gateway on both subnets carries it.

The job. The databricks-job unit defines <prefix>-etl, one Python task that runs /Volumes/acme_<stage>/processed/scripts/etl_transform.py, a PySpark script the code repository uploads to the lake.

SettingValue (stage file)
ComputeA job cluster per run on the latest LTS runtime: Standard_D4ds_v5 driver plus 2 workers, single-user access mode
SpotWorkers on spot with on-demand fallback in dev and staging, on-demand in prod; the driver is always on-demand
Limits3 concurrent runs, later runs queue; 60-minute timeout; one retry after a minute
IdentityRuns as the ETL trigger's managed identity, registered as a workspace service principal with Can Manage Run
PeopleACME_LeadETLEngineers can manage the job, ACME_ETLEngineers can view it
Cost (comments)VM time plus $0.30/DBU-hour for Premium jobs compute; spark_serverless ($0.35/DBU-hour) is the no-cluster option

The ETL trigger reads the job ID from the contract vault secret etl--job-id and starts one run per batch with the files as the input_path parameter. The run appends partitioned Parquet under processed/output/, or, with use_catalog = true, writes the managed Delta table acme_<stage>.processed.etl_output. What the job may read and write is bounded by its Unity Catalog grants. The Databricks objects are managed with the databricks/databricks ~> 1.135 provider.

Terms you will see​

TermMeaning
VNet injectionClusters in the spoke's own Databricks subnets.
Secure cluster connectivityCluster nodes without public IPs, connecting out to the control plane.
Job clusterA cluster created for one run and deleted after it.
DBUDatabricks unit, the compute charge on top of VM time.
Run asThe principal whose permissions a job run uses: the trigger identity.

Where to read more​