Architecture overview
The Data and ETL Blueprint runs an event-driven data pipeline in each platform project of the GCP Enterprise Baseline. A JSON file lands under input/ in the raw bucket, Eventarc delivers the event to an internal Cloud Run service, and the service starts one Dataproc Serverless batch that writes partitioned Parquet to the processed bucket, or rows of a BigQuery table in catalog mode. Dataplex Universal Catalog publishes the files as BigQuery tables on a schedule.
How it works
The pipeline has no always-on compute. The trigger service scales from zero, takes one event per instance and ignores every object that is not input/*.json. For a matching object it derives a batch ID from the bucket, the object name and its generation, so a redelivered event never starts a second batch, and it waits for a free slot when three batches are already pending or running. The batch runs the PySpark script gs://acme-usw1-<stage>-processed-data/scripts/etl_transform.py on the Dataproc Serverless 2.3 runtime, as the Spark service account, in the stage's data subnet. The data flow follows one file through every identity.
Three zones
| Zone | Bucket | Holds | Dataplex zone |
|---|---|---|---|
| Raw | acme-usw1-<stage>-raw-data | JSON files as producers deliver them, under input/ | raw (RAW) |
| Processed | acme-usw1-<stage>-processed-data | Partitioned Parquet under output/, and the Spark script under scripts/ | curated (CURATED) |
| Curated | acme-usw1-<stage>-curated-data | Business-ready data, aggregated and validated | curated (CURATED) |
Each bucket has uniform bucket-level access, public access prevention, versioning, soft delete and lifecycle tiering to Nearline and Coldline. Access is IAM only: no access control lists, no keys and no public objects.
Cataloged by Dataplex, queried in BigQuery
The Dataplex lake acme-usw1-<stage>-lake has two zones and one asset per bucket. Discovery runs hourly and publishes the files as tables of the BigQuery datasets raw and curated in the stage project, with the year, month and day partitions of the Parquet output kept current. Two more datasets belong to the blueprint itself: acme_<stage>_processed, where catalog mode writes the table etl_output, and acme_<stage>_curated for business tables. A policy-tag taxonomy with the tags pii and confidential restricts tagged columns to the lead ETL engineers.
Five units per stage
Each stage is built from the same template of five Terragrunt units:
unit "spark-identity" {
source = find_in_parent_folders("units/spark-identity")
path = "shared-infra/spark-identity"
}
unit "data-lake" {
source = find_in_parent_folders("units/data-lake")
path = "shared-infra/data-lake"
values = {
location = local.bucket_location
soft_delete_days = local.soft_delete_days
kms_key_name = local.kms_key_name
producer_members = local.producer_members
delete_protection = local.delete_protection
}
}
# --- the ETL pipeline ---------------------------------------------------------
unit "bigquery-datasets" {
source = find_in_parent_folders("units/bigquery-datasets")
path = "blueprint-etl/bigquery"
values = {
location = local.bucket_location
kms_key_name = local.kms_key_name
delete_protection = local.delete_protection
}
}
unit "dataplex" {
source = find_in_parent_folders("units/dataplex")
path = "blueprint-etl/dataplex"
values = {
discovery_schedule = local.discovery_schedule
bucket_location = local.bucket_location
asset_read_access_mode = local.asset_read_access_mode
}
}
unit "etl-trigger" {
source = find_in_parent_folders("units/etl-trigger")
path = "blueprint-etl/etl-trigger"
values = {
memory = local.trigger_memory
timeout_seconds = local.trigger_timeout
spark_runtime_version = local.spark_runtime_version
spark_executors = local.spark_executors
spark_executor_cores = local.spark_executor_cores
max_concurrent_batches = local.spark_max_concurrent
use_catalog = local.use_catalog
bucket_location = local.bucket_location
delete_protection = local.delete_protection
}
}
spark-identity and data-lake are the shared layer of the stage; bigquery-datasets, dataplex and etl-trigger are the pipeline. The stages differ in one value: prod sets delete_protection, so a non-empty bucket or dataset and the trigger service cannot be destroyed.
Delivery chain
The trigger image is built once, on merge, and never rebuilt. A release adds its version tag to the same image, deploys it to staging and, after a reviewer approves, to prod. Each deployment also uploads the Spark script to the stage's processed bucket and records the deployed tag in the stage's Parameter Manager parameter etl-trigger--image-tag, which the infrastructure reads back so a plan never rolls a deployment back. The pattern is described in immutable artifacts and environments and promotion.
What it builds on
The blueprint creates no network and no project. It reads the Shared VPC network, the stage's data subnet and the Artifact Registry path from the values the Baseline publishes in each stage project, pulls the trigger image from the Baseline's registry in acme-core-artifacts, and signs its pipelines in through the Baseline's Workload Identity Federation. Everything authenticates with IAM, so no service account key and no secret exists.
For the full list of services, see the service inventory. For what a customer receives, see what you receive.