Skip to main content

What you receive

You receive a working, event-driven data pipeline and a governed data lake, deployed into your dev, staging and prod projects on top of the GCP Enterprise Baseline, together with the two Git repositories that define it, customized for your organization, and the documentation to run and change it. After handover your engineers own all of it.

How delivery works​

The blueprint is sold with the Baseline in Full Deployment with Blueprints, or on its own as a Blueprint Deployment for customers who already run the GCP Enterprise Baseline. In both cases the BuiltForProd team deploys it, and the first deployment of each stage is performed once by that team. See blueprints for how the blueprints fit together and purchasing and licensing for the terms.

The repositories​

RepositoryYou receive
acme-gcp-blueprint-etl-infraFive OpenTofu modules, five Terragrunt units, one stack template, a stack file per stage, four guard scripts, pre-commit hooks, and the plan, apply and drift-detection workflows
acme-gcp-blueprint-etl-codeThe PySpark transform, the Cloud Run trigger service and its Dockerfile, unit tests for both, the deploy script scripts/deploy-stage.sh, pre-commit hooks and three GitHub Actions workflows

acme in the names is your namespace. Your domain, your GitHub organization and the project IDs of the Baseline are set in the code before deployment. Every value a deployment must set carries a TODO: marker, and every switch you may change later carries an @optional: marker, so git grep "@optional: " in either repository lists the choices. Each repository carries the PolyForm Internal Use License 1.0.0 in LICENSE.md.

The running pipeline​

In each of dev, staging and prod:

  • Three Cloud Storage buckets, acme-usw1-<stage>-raw-data, -processed-data and -curated-data, with uniform bucket-level access, public access prevention, versioning, a soft delete window and lifecycle tiering.
  • A Dataplex Universal Catalog lake with a raw and a curated zone, one asset per bucket and hourly discovery that publishes the files as BigQuery tables.
  • Two BigQuery datasets, acme_<stage>_processed and acme_<stage>_curated, and the policy-tag taxonomy acme-usw1-<stage>-sensitivity with the tags pii and confidential.
  • The Spark service account sa-acme-usw1-<stage>-spark that every batch runs as.
  • The internal Cloud Run service acme-usw1-<stage>-etl-trigger, its Eventarc trigger on the raw bucket, and the Parameter Manager parameter etl-trigger--image-tag that records the deployed image tag.
  • The IAM grants that give the lead ETL engineers read and write access, the ETL engineers read access, and your input producers create-only access to the raw bucket.

Upload a JSON file under input/ in the raw bucket and the pipeline runs end to end: one Dataproc Serverless batch per file, partitioned Parquet in the processed bucket, and a BigQuery table once discovery has run. The sample transform flattens nested fields, parses timestamps, keeps one row per id and adds processing columns.

The delivery pipeline​

  • Pull requests to the code repository run Ruff, the trigger tests on Python 3.12, the Spark tests on Python 3.11 with Java 17 and PySpark 3.5.3, a container build and a Trivy scan. This workflow never signs in to Google Cloud.
  • A merge to main builds the trigger image once as main-<sha>, then uploads the Spark script, deploys the trigger revision and records the tag in dev.
  • A GitHub Release adds the version tag to the same image and deploys it and the script to staging and, after a reviewer approves, to prod.
  • Pull requests to the infrastructure repository run lint, guard scripts, tflint and Checkov, then plan all three stages; a merge applies them, with prod waiting for a reviewer. A nightly run reports drift as an issue.

The documentation​

This documentation set: the architecture, the data lake, Dataproc Serverless, the ETL trigger, Dataplex governance, the ETL code, the pipelines, querying, security, performance tuning, operations, how-to guides for day-2 work such as granting data access and releasing, and a reference of every unit, module, stack value and environment variable.

What stays with the Baseline​

The organization, the projects, the Shared VPC, the CI identities, the Google groups, the guardrails, the security services and Artifact Registry come from the GCP Enterprise Baseline. The blueprint reads what it needs from Parameter Manager in each stage project and does not change the Baseline. The architecture overview shows how the two connect.