Skip to main content

Cloud Storage

Cloud Storage is Google Cloud's object store: objects in buckets, with access, retention and lifecycle set per bucket. The GCP Enterprise Baseline keeps the OpenTofu state and the audit archive in Cloud Storage, and the GCP Data and ETL Blueprint builds its data lake from three buckets per stage.

What it does​

A bucket has a location and a default storage class (Standard, Nearline, Coldline, Archive); lifecycle rules move or delete objects by age or state. Versioning keeps noncurrent versions of overwritten or deleted objects, and soft delete keeps deleted objects restorable for a set number of days. A retention policy forbids deleting an object before it reaches an age, and locking the policy makes that permanent. Uniform bucket-level access turns object ACLs off so IAM alone decides access, and public access prevention refuses any public grant. A bucket can encrypt with a customer-managed key from Cloud KMS.

How BuiltForProd uses it​

BucketProjectHoldsProtection and lifecycle
acme-usw1-root-tfstateacme-core-root (seed)OpenTofu state of every repositoryVersioning, 7-day soft delete, noncurrent versions deleted after 90 days, optional CMEK, prevent_destroy; writers scoped to their own prefix
acme-usw1-audit-logsacme-core-auditThe audit archive of the organization365-day retention policy (lock switch), CMEK from the kms-audit key ring, Nearline after 90 days, Coldline after 365, 7-day soft delete
<prefix>-raw-data, -processed-data, -curated-dataeach ETL stage projectInput files, Spark output, curated dataVersioning, soft delete (soft_delete_days, 7), Nearline at 90 days and Coldline at 180 for input/ and output/, noncurrent versions deleted after 30 days, optional CMEK

State. Every repository's root.hcl writes state to the same bucket with native locking: the Baseline under one prefix per unit, the Web App Blueprint under apps/app-blueprint/ and the Data and ETL Blueprint under apps/etl-blueprint/. Bucket IAM with prefix conditions lets each repository identity write only its own prefix.

Audit archive. The archive sink of Cloud Logging writes here. Current objects are never deleted by a rule; lock_archive_bucket_retention in security.hcl locks the retention policy, which cannot be undone. The key comes from Cloud KMS, and an IAM deny policy blocks deleting it.

The data lake. Bucket names are <prefix>-<zone>-data, for example acme-usw1-prd-raw-data. Only a few identities hold bucket roles: the Spark service account reads raw and writes processed and curated; lead ETL engineers read and write every zone and may upload to raw; ETL engineers read every zone; producers listed in producer_members may only create objects in raw; the ETL code pipeline may write only under scripts/ of processed, through an IAM condition. Everyone else reaches the data through Dataplex Universal Catalog and BigQuery. In prod, delete_protection keeps a non-empty bucket from being destroyed. The kms_key_name stage value turns on CMEK for the buckets and datasets together. A new object in raw raises the event that starts the pipeline; see Eventarc.

Controls everywhere. The organization policies storage.uniformBucketLevelAccess and storage.publicAccessPrevention apply to every project; only acme-core-public has a project-level exception to public access prevention. Access is IAM only, with no ACLs and no keys, and Data Access audit logs record object reads.

Terms you will see​

TermMeaning
Uniform bucket-level accessIAM alone controls access; object ACLs are off.
Public access preventionNo public grant can be added to the bucket.
Soft deleteDeleted objects stay restorable for a number of days.
Retention lockA permanent minimum age before objects can be deleted.
Prefix conditionAn IAM condition that limits a grant to object names under a prefix.

Where to read more​