Skip to main content

Dataproc Serverless

Dataproc Serverless runs Apache Spark workloads without a cluster: each batch gets its compute when it starts and releases it when it ends. The GCP Data and ETL Blueprint submits one batch for every input file that lands in a stage's raw bucket.

What it does​

A batch names a script, a runtime version (which fixes the Spark, Java, Scala and Python versions), the executors and cores it may use, a service account to run as and the subnet its workers use. Google provisions the workers in that subnet, runs the job and tears them down. Compute is billed in Data Compute Units (DCU) per hour of the batch's run time.

How BuiltForProd uses it​

Who submits. The Cloud Run service <prefix>-etl-trigger receives the object-finalized event from Eventarc, ignores anything but input/*.json, and submits one batch per object. It holds roles/dataproc.editor (submit and list batches) and roles/iam.serviceAccountUser on the Spark identity, the only identity the blueprint lets act as it.

What runs. The batch runs etl_transform.py from gs://<prefix>-processed-data/scripts/, which the ETL code pipeline uploads with every deploy, so the script and the trigger image move between stages together. The batch reads the JSON object and writes either partitioned Parquet under output/year=YYYY/month=M/day=D/ in the processed bucket (the default) or, with use_catalog = true, the table etl_output in the stage's processed BigQuery dataset.

As whom and where. The Spark identity sa-<prefix>-spark (for example sa-acme-usw1-prd-spark) holds roles/dataproc.worker on the project and the bucket, dataset and curated-zone grants the other units give it. Workers run in the stage's data subnet of the Shared VPC and reach Google APIs through Private Google Access.

Sizing. Each stage file sets the batch values, the same in dev, staging and prod:

ValueSettingMeaning
spark_runtime_version2.3The Dataproc Serverless runtime
spark_executors2Executors per batch; raise for larger inputs (billed per DCU-hour)
spark_executor_cores4Cores per executor
spark_max_concurrent_batches3Batches that may run at once

The concurrency limit is enforced by the trigger, which lists the running batches and waits for a slot within its 300-second request timeout; Cloud Run never runs more trigger instances than that limit, and Eventarc retries an event whose request timed out.

Terms you will see​

TermMeaning
BatchOne Spark job with its own short-lived compute.
Runtime versionThe pinned Spark and language versions of a batch (2.3).
DCUData Compute Unit, the billing unit of batch compute.
Spark identitysa-<prefix>-spark, the service account every batch runs as.
Catalog modeWriting the result to BigQuery instead of Parquet (use_catalog).

Where to read more​