BigQuery
BigQuery is Google Cloud's serverless data warehouse: tables in datasets, queried with SQL and billed by storage and by the data each query reads. The GCP Data and ETL Blueprint keeps its processed and curated tables there, with column-level access through policy tags, and the GCP Enterprise Baseline exposes its central audit log bucket to BigQuery as a linked dataset.
What it does
A dataset groups tables in one location and carries IAM roles such as Data Viewer and Data Editor. A policy tag from a Data Catalog taxonomy attached to a column restricts reading that column to principals holding Fine-Grained Reader on the tag, whatever their dataset role. A linked dataset makes a Cloud Logging bucket queryable from BigQuery without copying the logs. Dataplex discovery can also publish files in Cloud Storage as BigQuery tables.
How BuiltForProd uses it
Blueprint datasets. The bigquery-datasets unit creates two datasets in every ETL stage project, in the lake's location, named after the stage slug (dev, stg, prd):
| Dataset | Holds | Access |
|---|---|---|
acme_prd_processed | The batch's catalog-mode output table etl_output and other transformed tables | Lead ETL engineers and the Spark identity edit; ETL engineers read |
acme_prd_curated | Business-ready, aggregated and validated tables | The same |
Tables have no default expiration. In prod, delete_protection keeps a dataset that still holds tables from being destroyed; kms_key_name turns on CMEK for the datasets and the buckets together.
Catalog mode. By default the Spark batch writes partitioned Parquet to the processed bucket. With use_catalog = true in a stage file, the batch writes its result to etl_output in the stage's processed dataset through the BigQuery connector instead; see Dataproc Serverless.
Column-level access. The unit creates the taxonomy <prefix>-sensitivity with fine-grained access control and two policy tags, pii and confidential. A column tagged with either is readable only by the lead ETL engineers group, which holds Fine-Grained Reader on both tags, so a dataset viewer sees the rest of the table but not those columns.
Discovered tables. Dataplex Universal Catalog publishes the files of the lake as tables in datasets named after its zones, raw and curated; those datasets belong to Dataplex, not to this unit.
Logs in BigQuery. On the Baseline side, the central log bucket acme-audit in acme-core-audit has Log Analytics on and a linked dataset acme_audit_analytics (enable_linked_dataset, on by default). SQL over a year of audit logs runs from BigQuery with no export and no storage charge; queries are billed at BigQuery prices. See Cloud Logging.
Terms you will see
| Term | Meaning |
|---|---|
| Dataset | A container of tables with its own location and IAM. |
| Policy tag | A label on a column that restricts who can read it. |
| Fine-Grained Reader | The role that lets a principal read columns under a policy tag. |
| Catalog mode | The batch writes to etl_output in BigQuery instead of Parquet files. |
| Linked dataset | A BigQuery view of a Cloud Logging bucket. |
Where to read more
- GCP Data and ETL Blueprint overview and the GCP Enterprise Baseline overview.
- Dataplex Universal Catalog for the discovered tables.
- Cloud Logging for the bucket behind the linked dataset.