Skip to main content

Service inventory

The Data and ETL Blueprint uses eleven Google Cloud services in each stage project, created by five units, and relies on four more that the GCP Enterprise Baseline runs in its core projects. Everything in the first table is created by the blueprint's own code, except the Dataproc Serverless batches, which the trigger service submits at run time.

How the pieces group​

Google Cloud services​

ServiceUnitWhat it does in the blueprint
Cloud Storagedata-lakeThe buckets acme-usw1-<stage>-raw-data, -processed-data and -curated-data: uniform bucket-level access, public access prevention, versioning, soft delete, lifecycle tiering, optional CMEK
Dataplex Universal CatalogdataplexThe lake acme-usw1-<stage>-lake, the zones raw (RAW) and curated (CURATED), one asset per bucket, hourly discovery and the zone IAM
BigQuerybigquery-datasets, dataplexThe datasets acme_<stage>_processed and acme_<stage>_curated, and the datasets raw and curated that Dataplex discovery publishes the files into
BigQuery policy tags (Data Catalog taxonomy)bigquery-datasetsThe taxonomy acme-usw1-<stage>-sensitivity with the tags pii and confidential for column-level access
Dataproc Serverlessspark-identity, etl-triggerOne batch per input file on runtime 2.3: two executors of four cores, at most three batches at once, one-hour time to live, in the stage's data subnet
Eventarcetl-triggerThe trigger acme-usw1-<stage>-raw-finalized: the raw bucket's object.finalized events, delivered to the trigger service
Pub/Subdata-lakeThe transport Eventarc uses for Cloud Storage events; the project's Cloud Storage service agent may publish
Cloud Runetl-triggerThe service acme-usw1-<stage>-etl-trigger: internal ingress, no unauthenticated access, one request per instance, zero to three instances, direct VPC egress through the data subnet
Parameter Manageretl-triggerReads three landing-zone values of the Baseline; holds etl-trigger--image-tag, the trigger image tag deployed in the stage
IAM and service accountsspark-identity, etl-trigger, every unitThree service accounts per stage (Spark, trigger, Eventarc) and the resource-level grants for them, the ETL engineer groups, the producers and the code pipeline
Cloud Loggingetl-triggerThe trigger's structured JSON log lines (object, batch ID, operation); Dataproc Serverless writes each batch's output there as well

Cloud KMS is used only when a stage sets a customer-managed key for the buckets and datasets; by default Google-managed encryption applies.

Used from the Baseline​

ServiceWhere it livesUse
Artifact Registryacme-core-artifactsThe repository acme-gcp-blueprint-etl-trigger, which holds the trigger image with immutable tags
Workload Identity Federationacme-core-autoShort-lived credentials for every workflow of both repositories, with no stored keys
Shared VPCacme-core-networkThe stage's data subnet, where the trigger's egress and every batch run, with Private Google Access
Parameter Manager contractEach stage project (pm-publish)The network, data subnet and registry path the etl-trigger unit reads at plan time

The Baseline also enables the APIs the blueprint needs in each stage project and grants the stage's Cloud Run and Dataproc service agents the right to use the data subnet.

Delivery tooling​

The infrastructure is written with OpenTofu and Terragrunt Stacks. The trigger is packaged as a Docker image on python:3.12-slim, and every pipeline runs on GitHub Actions. Checkov, Trivy and tflint scan the infrastructure code in pre-commit hooks, the plan workflow runs tflint and Checkov again, Ruff lints the Python code, and Trivy scans every trigger image before a pull request can merge.

The exact versions are on the versions page. The full catalog of services across BuiltForProd products is in Google Cloud services used.