Google Kubernetes Engine
Google Kubernetes Engine (GKE) is Google's managed Kubernetes: Google runs the control plane, and the nodes are Compute Engine virtual machines in your project. The GCP Web App Blueprint runs one private regional cluster per stage, <prefix>-gke such as acme-usw1-prd-gke, and the GCP Enterprise Baseline can run a small Autopilot cluster for self-hosted CI runners.
What it does
A Standard cluster lets you manage node pools; an Autopilot cluster manages nodes for you and bills per pod. A regional cluster spreads the control plane and nodes over several zones. Release channels keep the version current. Private nodes have no external addresses. The DNS-based control-plane endpoint accepts kubectl calls through Google APIs, authenticated by IAM, without any network path to the cluster. Dataplane V2 is the eBPF data plane that enforces NetworkPolicy. Node auto-provisioning creates node pools sized for pending pods. Workload Identity Federation for GKE lets pods call Google APIs as IAM principals without keys.
How BuiltForProd uses it
The gke unit builds each stage cluster from the landing-zone contract and the stage file:
| Aspect | Setting |
|---|---|
| Type and version | Standard, regional, REGULAR release channel, minimum control-plane version 1.35 at creation |
| Network | VPC-native on the stage's nodes subnet of the Shared VPC, pods and services in its secondary ranges; Dataplane V2; intranode visibility on |
| Control plane | DNS-based endpoint (IAM tokens) and the internal IP endpoint in the contract's control-plane range; external IP endpoint off; no client certificates |
| Nodes | Private, shielded (secure boot, integrity monitoring), Container-Optimized OS, GKE metadata server; egress through Cloud NAT |
| System pool | system, e2-standard-2 (~$49/month per node), 1 to 2 nodes per zone, tainted CriticalAddonsOnly; dev in 2 zones, staging and prod in 3 |
| Application nodes | Node auto-provisioning up to nap_cpu_limit vCPUs and nap_memory_limit GiB: 32 and 128 in dev, 64 and 256 in staging, 128 and 512 in prod |
| Features | Workload Identity Federation for GKE, Gateway API (standard channel), NodeLocal DNSCache, security posture (basic), Managed Service for Prometheus |
| Maintenance | Weekly window Tuesday to Thursday, 07:00 to 11:00 UTC |
| Protection | gke_deletion_protection on in prod |
resource "google_container_node_pool" "system" {
project = var.project_id
name = "system"
location = var.gcp_region
cluster = google_container_cluster.this.name
node_locations = local.zones
initial_node_count = var.min_per_zone
autoscaling {
min_node_count = var.min_per_zone
max_node_count = var.max_per_zone
location_policy = "BALANCED"
}
management {
auto_repair = true
auto_upgrade = true
}
upgrade_settings {
max_surge = 1
max_unavailable = 0
}
node_config {
machine_type = var.system_machine_type
image_type = "COS_CONTAINERD"
disk_type = "pd-balanced"
disk_size_gb = 50
service_account = google_service_account.nodes.email
oauth_scopes = ["https://www.googleapis.com/auth/cloud-platform"]
resource_labels = var.labels
labels = { "node-role" = "system" }
metadata = { "disable-legacy-endpoints" = "true" }
shielded_instance_config {
enable_secure_boot = true
enable_integrity_monitoring = true
}
workload_metadata_config {
mode = "GKE_METADATA"
}
taint {
key = "CriticalAddonsOnly"
value = "true"
effect = "NO_SCHEDULE"
}
}
lifecycle {
# The autoscaler owns the node count; zones are chosen once with the cluster.
ignore_changes = [initial_node_count, node_locations]
}
}
Why a tainted system pool. Every add-on the blueprint installs tolerates the taint and runs on the system pool; the application chart deliberately does not, so its pods land on auto-provisioned nodes, which the autoscaler adds and removes. The NAP limits are the upper bound of the application bill. See Kubernetes for what runs where.
Nodes and images. The node service account sa-acme-gke-nodes writes logs and metrics and nothing else in the project; the Baseline's Artifact Registry unit makes it a reader of the API image repository for the stages listed in webapp_gke_stages.
Access. OpenTofu, Helm and engineers reach the cluster through the DNS-based endpoint with IAM access tokens; access to the cluster is IAM only. Optional settings in the unit file: Binary Authorization enforcement (enable_binary_authorization), turning off Managed Service for Prometheus, and authorized networks for runners that use the internal endpoint.
Runner cluster (Baseline, optional). With enable_github_runners = true in units/github-runners/terragrunt.hcl, off by default, the Baseline creates a private GKE Autopilot cluster <prefix>-runners in acme-core-auto on the hub's runner subnet, with its own DNS-based endpoint, for Actions Runner Controller (~$73/month cluster fee beyond the free tier credit, plus pod resources while jobs run). See GitHub Actions.
Terms you will see
| Term | Meaning |
|---|---|
| System pool | The fixed, tainted node pool for the platform add-ons. |
| Node auto-provisioning | GKE creating node pools sized for pending pods, within CPU and memory limits. |
| DNS-based endpoint | The control-plane address reached through Google APIs with IAM tokens. |
| Dataplane V2 | The eBPF data plane that enforces NetworkPolicy. |
webapp_gke_stages | The Baseline list of stages whose node accounts may pull the API image. |
Where to read more
- GCP Web App Blueprint overview for the request path end to end.
- Kubernetes for what runs inside the cluster.
- GKE Gateway for how traffic enters it.