Service inventory
The Data and ETL Blueprint uses eleven Google Cloud services in each stage project, created by five units, and relies on four more that the GCP Enterprise Baseline runs in its core projects. Everything in the first table is created by the blueprint's own code, except the Dataproc Serverless batches, which the trigger service submits at run time.
How the pieces group
Google Cloud services
| Service | Unit | What it does in the blueprint |
|---|---|---|
| Cloud Storage | data-lake | The buckets acme-usw1-<stage>-raw-data, -processed-data and -curated-data: uniform bucket-level access, public access prevention, versioning, soft delete, lifecycle tiering, optional CMEK |
| Dataplex Universal Catalog | dataplex | The lake acme-usw1-<stage>-lake, the zones raw (RAW) and curated (CURATED), one asset per bucket, hourly discovery and the zone IAM |
| BigQuery | bigquery-datasets, dataplex | The datasets acme_<stage>_processed and acme_<stage>_curated, and the datasets raw and curated that Dataplex discovery publishes the files into |
| BigQuery policy tags (Data Catalog taxonomy) | bigquery-datasets | The taxonomy acme-usw1-<stage>-sensitivity with the tags pii and confidential for column-level access |
| Dataproc Serverless | spark-identity, etl-trigger | One batch per input file on runtime 2.3: two executors of four cores, at most three batches at once, one-hour time to live, in the stage's data subnet |
| Eventarc | etl-trigger | The trigger acme-usw1-<stage>-raw-finalized: the raw bucket's object.finalized events, delivered to the trigger service |
| Pub/Sub | data-lake | The transport Eventarc uses for Cloud Storage events; the project's Cloud Storage service agent may publish |
| Cloud Run | etl-trigger | The service acme-usw1-<stage>-etl-trigger: internal ingress, no unauthenticated access, one request per instance, zero to three instances, direct VPC egress through the data subnet |
| Parameter Manager | etl-trigger | Reads three landing-zone values of the Baseline; holds etl-trigger--image-tag, the trigger image tag deployed in the stage |
| IAM and service accounts | spark-identity, etl-trigger, every unit | Three service accounts per stage (Spark, trigger, Eventarc) and the resource-level grants for them, the ETL engineer groups, the producers and the code pipeline |
| Cloud Logging | etl-trigger | The trigger's structured JSON log lines (object, batch ID, operation); Dataproc Serverless writes each batch's output there as well |
Cloud KMS is used only when a stage sets a customer-managed key for the buckets and datasets; by default Google-managed encryption applies.
Used from the Baseline
| Service | Where it lives | Use |
|---|---|---|
| Artifact Registry | acme-core-artifacts | The repository acme-gcp-blueprint-etl-trigger, which holds the trigger image with immutable tags |
| Workload Identity Federation | acme-core-auto | Short-lived credentials for every workflow of both repositories, with no stored keys |
| Shared VPC | acme-core-network | The stage's data subnet, where the trigger's egress and every batch run, with Private Google Access |
| Parameter Manager contract | Each stage project (pm-publish) | The network, data subnet and registry path the etl-trigger unit reads at plan time |
The Baseline also enables the APIs the blueprint needs in each stage project and grants the stage's Cloud Run and Dataproc service agents the right to use the data subnet.
Delivery tooling
The infrastructure is written with OpenTofu and Terragrunt Stacks. The trigger is packaged as a Docker image on python:3.12-slim, and every pipeline runs on GitHub Actions. Checkov, Trivy and tflint scan the infrastructure code in pre-commit hooks, the plan workflow runs tflint and Checkov again, Ruff lints the Python code, and Trivy scans every trigger image before a pull request can merge.
The exact versions are on the versions page. The full catalog of services across BuiltForProd products is in Google Cloud services used.