GCP Data and ETL Blueprint
The Data and ETL Blueprint for Google Cloud is an event-driven data platform deployed on the GCP Enterprise Baseline: every JSON file that lands in the raw bucket of a stage becomes partitioned Parquet, or a BigQuery table, through one Dataproc Serverless batch that an Eventarc-driven Cloud Run service starts. This overview section is public.
What it deploys
The blueprint is two repositories, deployed into the dev, staging and prod platform projects of the Baseline.
| Repository | What it holds |
|---|---|
acme-gcp-blueprint-etl-infra | Per stage: raw, processed and curated Cloud Storage buckets with uniform bucket-level access, versioning and lifecycle tiering; a Dataplex Universal Catalog lake whose discovery publishes the files as BigQuery tables; BigQuery datasets with a policy-tag taxonomy for column-level access; the Spark service account; and the trigger: Eventarc on the raw bucket and an internal Cloud Run service that starts one Dataproc Serverless batch per input file |
acme-gcp-blueprint-etl-code | The PySpark transform each batch runs and the Python trigger service, with GitHub Actions pipelines that build the trigger image once, tag it immutably in Artifact Registry, upload the transform and promote both from dev to staging and prod |
Everything authenticates with IAM: no service account key and no secret exists, and the Secret Manager of each stage stays empty by design. The blueprints page describes how blueprints build on the Baseline.
The complete Data and ETL Blueprint documentation for Google Cloud is available to customers who hold it. Read the documentation overview for the concepts behind it, then sign in from the navigation bar to open the full Data and ETL Blueprint documentation, or contact BuiltForProd to purchase it.