Skip to main content

Google Kubernetes Engine

Google Kubernetes Engine (GKE) is Google's managed Kubernetes: Google runs the control plane, and the nodes are Compute Engine virtual machines in your project. The GCP Web App Blueprint runs one private regional cluster per stage, <prefix>-gke such as acme-usw1-prd-gke, and the GCP Enterprise Baseline can run a small Autopilot cluster for self-hosted CI runners.

What it does​

A Standard cluster lets you manage node pools; an Autopilot cluster manages nodes for you and bills per pod. A regional cluster spreads the control plane and nodes over several zones. Release channels keep the version current. Private nodes have no external addresses. The DNS-based control-plane endpoint accepts kubectl calls through Google APIs, authenticated by IAM, without any network path to the cluster. Dataplane V2 is the eBPF data plane that enforces NetworkPolicy. Node auto-provisioning creates node pools sized for pending pods. Workload Identity Federation for GKE lets pods call Google APIs as IAM principals without keys.

How BuiltForProd uses it​

The gke unit builds each stage cluster from the landing-zone contract and the stage file:

AspectSetting
Type and versionStandard, regional, REGULAR release channel, minimum control-plane version 1.35 at creation
NetworkVPC-native on the stage's nodes subnet of the Shared VPC, pods and services in its secondary ranges; Dataplane V2; intranode visibility on
Control planeDNS-based endpoint (IAM tokens) and the internal IP endpoint in the contract's control-plane range; external IP endpoint off; no client certificates
NodesPrivate, shielded (secure boot, integrity monitoring), Container-Optimized OS, GKE metadata server; egress through Cloud NAT
System poolsystem, e2-standard-2 (~$49/month per node), 1 to 2 nodes per zone, tainted CriticalAddonsOnly; dev in 2 zones, staging and prod in 3
Application nodesNode auto-provisioning up to nap_cpu_limit vCPUs and nap_memory_limit GiB: 32 and 128 in dev, 64 and 256 in staging, 128 and 512 in prod
FeaturesWorkload Identity Federation for GKE, Gateway API (standard channel), NodeLocal DNSCache, security posture (basic), Managed Service for Prometheus
MaintenanceWeekly window Tuesday to Thursday, 07:00 to 11:00 UTC
Protectiongke_deletion_protection on in prod
GCP/acme-gcp-blueprint-webapp-infra/modules/gke/nodes.tf (lines 18-74)
resource "google_container_node_pool" "system" {
project = var.project_id
name = "system"
location = var.gcp_region
cluster = google_container_cluster.this.name
node_locations = local.zones

initial_node_count = var.min_per_zone

autoscaling {
min_node_count = var.min_per_zone
max_node_count = var.max_per_zone
location_policy = "BALANCED"
}

management {
auto_repair = true
auto_upgrade = true
}

upgrade_settings {
max_surge = 1
max_unavailable = 0
}

node_config {
machine_type = var.system_machine_type
image_type = "COS_CONTAINERD"
disk_type = "pd-balanced"
disk_size_gb = 50
service_account = google_service_account.nodes.email
oauth_scopes = ["https://www.googleapis.com/auth/cloud-platform"]
resource_labels = var.labels
labels = { "node-role" = "system" }
metadata = { "disable-legacy-endpoints" = "true" }

shielded_instance_config {
enable_secure_boot = true
enable_integrity_monitoring = true
}

workload_metadata_config {
mode = "GKE_METADATA"
}

taint {
key = "CriticalAddonsOnly"
value = "true"
effect = "NO_SCHEDULE"
}
}

lifecycle {
# The autoscaler owns the node count; zones are chosen once with the cluster.
ignore_changes = [initial_node_count, node_locations]
}
}

Why a tainted system pool. Every add-on the blueprint installs tolerates the taint and runs on the system pool; the application chart deliberately does not, so its pods land on auto-provisioned nodes, which the autoscaler adds and removes. The NAP limits are the upper bound of the application bill. See Kubernetes for what runs where.

Nodes and images. The node service account sa-acme-gke-nodes writes logs and metrics and nothing else in the project; the Baseline's Artifact Registry unit makes it a reader of the API image repository for the stages listed in webapp_gke_stages.

Access. OpenTofu, Helm and engineers reach the cluster through the DNS-based endpoint with IAM access tokens; access to the cluster is IAM only. Optional settings in the unit file: Binary Authorization enforcement (enable_binary_authorization), turning off Managed Service for Prometheus, and authorized networks for runners that use the internal endpoint.

Runner cluster (Baseline, optional). With enable_github_runners = true in units/github-runners/terragrunt.hcl, off by default, the Baseline creates a private GKE Autopilot cluster <prefix>-runners in acme-core-auto on the hub's runner subnet, with its own DNS-based endpoint, for Actions Runner Controller (~$73/month cluster fee beyond the free tier credit, plus pod resources while jobs run). See GitHub Actions.

Terms you will see​

TermMeaning
System poolThe fixed, tainted node pool for the platform add-ons.
Node auto-provisioningGKE creating node pools sized for pending pods, within CPU and memory limits.
DNS-based endpointThe control-plane address reached through Google APIs with IAM tokens.
Dataplane V2The eBPF data plane that enforces NetworkPolicy.
webapp_gke_stagesThe Baseline list of stages whose node accounts may pull the API image.

Where to read more​