Skip to main content

Dataplex Universal Catalog

Dataplex Universal Catalog catalogs and governs data spread across Cloud Storage and BigQuery. The GCP Data and ETL Blueprint uses it to organize each stage's data lake into a raw and a curated zone, discover the files that land there and publish them as BigQuery tables.

What it does​

A lake groups the data of one domain. Zones inside it carry a type: a RAW zone accepts data in any format as it arrives, a CURATED zone expects analytics-ready formats such as Parquet. An asset attaches a Cloud Storage bucket or a BigQuery dataset to a zone. Discovery scans assets on a schedule, infers schemas and partitions, and publishes tables to BigQuery. With managed read access the published tables are BigLake tables governed by the zone's IAM roles; with direct access readers use the bucket's own permissions.

How BuiltForProd uses it​

The dataplex unit builds one lake per ETL stage, <prefix>-lake (for example acme-usw1-prd-lake), over the three Cloud Storage buckets of the stage:

AssetZoneDiscovery
rawrawJSON inputs as delivered (UTF-8)
processedcuratedThe batch's Parquet under output/year=YYYY/month=M/day=D/
curatedcuratedBusiness-ready data

Discovery runs on discovery_schedule, hourly (0 * * * *, UTC) in every stage file; the first run happens at the next schedule, not when the lake is created. It publishes tables in BigQuery datasets named after the zones and keeps the Hive-style date partitions of the batch output current. More frequent runs mean more discovery jobs.

Zone IAM. The lead ETL engineers and ETL engineers groups get roles/dataplex.dataReader on both zones; the Spark identity and the lead ETL engineers get roles/dataplex.dataWriter on the curated zone. asset_read_access_mode is DIRECT by default, so the bucket bindings of the data-lake unit enforce access and the zone bindings record the intended access; MANAGED publishes BigLake tables that the zone IAM governs instead.

Location. A regional bucket_location such as us-west1 attaches the buckets as single-region assets; the US multi-region switches them to multi-region. The lake and zones stay in the stage's region.

Terms you will see​

TermMeaning
LakeThe top-level Dataplex container for one stage's data.
RAW zoneA zone for data as it lands, any format.
CURATED zoneA zone for analytics-ready formats such as Parquet.
AssetA bucket or dataset attached to a zone.
DiscoveryThe scheduled scan that infers schemas and publishes BigQuery tables.
Read access modeDIRECT (bucket IAM) or MANAGED (BigLake tables under zone IAM).

Where to read more​