Azure Data and ETL Blueprint
The Data and ETL Blueprint for Azure is an event-driven data platform deployed on the Azure Enterprise Baseline: a JSON file dropped into the raw zone of a stage's data lake becomes partitioned Parquet, or a Delta table in Unity Catalog, through an Azure Databricks job that an event-driven Azure Container Apps job starts. This overview section is public.
What it deploys
The blueprint is two repositories, deployed into the dev, staging and prod platform subscriptions of the Baseline.
| Repository | What it holds |
|---|---|
acme-azure-blueprint-etl-infra | Per stage: an Azure Data Lake Storage (ADLS) Gen2 lake with raw, processed, curated and managed zones behind private endpoints; an Azure Databricks Premium workspace injected into the stage network with secure cluster connectivity; Unity Catalog with a storage credential, external locations, a catalog per stage and its schemas; the Databricks job; and the trigger: an Event Grid system topic, a Storage queue and an event-driven Container Apps job |
acme-azure-blueprint-etl-code | The PySpark transform the Databricks job runs and the Python trigger container, with GitHub Actions pipelines that build the trigger image once, tag it immutably in Azure Container Registry, upload the transform and promote both from dev to staging and prod |
Every identity authenticates with Microsoft Entra tokens: no storage account key, shared access signature or Databricks personal access token exists. The blueprints page describes how blueprints build on the Baseline.
The complete Data and ETL Blueprint documentation for Azure is available to customers who hold it. Read the documentation overview for the concepts behind it, then sign in from the navigation bar to open the full Data and ETL Blueprint documentation, or contact BuiltForProd to purchase it.