Dataproc Serverless
Dataproc Serverless runs Apache Spark workloads without a cluster: each batch gets its compute when it starts and releases it when it ends. The GCP Data and ETL Blueprint submits one batch for every input file that lands in a stage's raw bucket.
What it does
A batch names a script, a runtime version (which fixes the Spark, Java, Scala and Python versions), the executors and cores it may use, a service account to run as and the subnet its workers use. Google provisions the workers in that subnet, runs the job and tears them down. Compute is billed in Data Compute Units (DCU) per hour of the batch's run time.
How BuiltForProd uses it
Who submits. The Cloud Run service <prefix>-etl-trigger receives the object-finalized event from Eventarc, ignores anything but input/*.json, and submits one batch per object. It holds roles/dataproc.editor (submit and list batches) and roles/iam.serviceAccountUser on the Spark identity, the only identity the blueprint lets act as it.
What runs. The batch runs etl_transform.py from gs://<prefix>-processed-data/scripts/, which the ETL code pipeline uploads with every deploy, so the script and the trigger image move between stages together. The batch reads the JSON object and writes either partitioned Parquet under output/year=YYYY/month=M/day=D/ in the processed bucket (the default) or, with use_catalog = true, the table etl_output in the stage's processed BigQuery dataset.
As whom and where. The Spark identity sa-<prefix>-spark (for example sa-acme-usw1-prd-spark) holds roles/dataproc.worker on the project and the bucket, dataset and curated-zone grants the other units give it. Workers run in the stage's data subnet of the Shared VPC and reach Google APIs through Private Google Access.
Sizing. Each stage file sets the batch values, the same in dev, staging and prod:
| Value | Setting | Meaning |
|---|---|---|
spark_runtime_version | 2.3 | The Dataproc Serverless runtime |
spark_executors | 2 | Executors per batch; raise for larger inputs (billed per DCU-hour) |
spark_executor_cores | 4 | Cores per executor |
spark_max_concurrent_batches | 3 | Batches that may run at once |
The concurrency limit is enforced by the trigger, which lists the running batches and waits for a slot within its 300-second request timeout; Cloud Run never runs more trigger instances than that limit, and Eventarc retries an event whose request timed out.
Terms you will see
| Term | Meaning |
|---|---|
| Batch | One Spark job with its own short-lived compute. |
| Runtime version | The pinned Spark and language versions of a batch (2.3). |
| DCU | Data Compute Unit, the billing unit of batch compute. |
| Spark identity | sa-<prefix>-spark, the service account every batch runs as. |
| Catalog mode | Writing the result to BigQuery instead of Parquet (use_catalog). |
Where to read more
- GCP Data and ETL Blueprint overview for the pipeline end to end.
- Eventarc for how a landed file starts the trigger.
- Dataplex Universal Catalog for how the output becomes tables.