Azure Event Grid
Azure Event Grid is Microsoft's event routing service: Azure resources publish events, and subscriptions deliver the ones that match a filter to a handler. The Azure Data and ETL Blueprint uses it to turn every JSON file that lands in the raw zone of a stage's data lake into a message the ETL trigger consumes.
What it does
A system topic publishes the events of one Azure resource, such as a storage account's Microsoft.Storage.BlobCreated. An event subscription on the topic filters events by type, by subject prefix and suffix and by advanced filters on event data, then delivers them to an endpoint such as a Storage queue. Delivery can authenticate with the topic's managed identity and is retried on failure until a maximum attempt count or event lifetime is reached.
How BuiltForProd uses it
The data-lake unit creates the system topic <prefix>-datalake-events on the stage's lake account, with a system-assigned identity that holds Storage Queue Data Message Sender on the account, and the queue etl-trigger in the same account. One subscription, raw-input-json, does the rest:
| Setting | Value |
|---|---|
| Event type | Microsoft.Storage.BlobCreated |
| Subject | Begins with /blobServices/default/containers/raw/blobs/input/, ends with .json, case-insensitive |
| Advanced filter | data.api is PutBlob, PutBlockList, FlushWithClose or CopyBlob |
| Destination | Storage queue etl-trigger, delivered with the topic's managed identity; messages live 24 hours |
| Retry | Up to 30 attempts within 1,440 minutes |
The advanced filter matters on a hierarchical-namespace account: a Data Lake upload raises BlobCreated twice, once for the empty file and once when it is flushed and closed. Only the second, complete file starts the pipeline; blob uploads and copies raise one event anyway.
The lake's firewall denies public traffic by default and lets trusted Azure services through, so Event Grid's delivery to the queue passes with its identity and no key; see Azure Data Lake Storage. From the queue, the ETL trigger, a Container Apps job, scales on the backlog, reads up to 32 messages per execution and starts one Databricks run per batch. Storage queues have no dead-letter queue: the trigger logs and removes a message after too many failed receives, and the 24-hour message lifetime bounds the rest.
Terms you will see
| Term | Meaning |
|---|---|
| System topic | The event source for one Azure resource, here the lake account. |
| Event subscription | The filter and destination for a topic's events. |
| Subject filter | The path prefix and suffix an event's subject must match. |
| Advanced filter | A match on a field of the event data, here the upload API. |
FlushWithClose | The Data Lake call that completes a file upload. |
Where to read more
- Azure Data and ETL Blueprint overview for the pipeline end to end.
- Azure Data Lake Storage for the lake and its zones.
- Azure Container Apps for the trigger job.