Node Auto Provisioning
Node Auto Provisioning (NAP) is the AKS mode that creates nodes of the right size for pods waiting to be scheduled and removes them when they are no longer needed. The Azure Web App Blueprint gives its application pods no fixed node pool: every application node comes from NAP.
What it does
NAP is built on the open-source Karpenter project and is run by AKS. It watches for pending pods, picks a virtual machine size that fits them from the sizes a NodePool allows, launches the node and lets the scheduler place the pods. It also consolidates: it removes empty nodes and replaces underused ones with fewer or smaller nodes, within a disruption budget. An AKSNodeClass sets the node image, the OS disk and the tags. A NodePool's limits cap the total CPU it may create.
How BuiltForProd uses it
The AKS cluster runs NAP with its built-in pools off, so the two objects the node-pools unit (module nap-node-pools) writes are the only application capacity:
| Object | Setting |
|---|---|
AKSNodeClass default | Azure Linux image, 128 GB OS disk, the platform tags |
NodePool default | amd64 Linux; capacity types spot and on-demand, spot preferred with on-demand fallback |
| VM families D and E, at most 32 vCPU per node; no taints | |
Consolidation WhenEmptyOrUnderutilized after 1 minute; at most 10% of nodes disrupted at once | |
| Nodes expire after 720 hours (30 days) and are replaced | |
CPU limit from the stage file's nap_cpu_limit |
| Stage | nap_cpu_limit |
|---|---|
| dev | 32 vCPU |
| staging | 64 vCPU |
| prod | 128 vCPU |
The limit is the bill's upper bound for application nodes: once the pool holds that many vCPUs, further pods stay pending instead of buying more capacity.
Where pods land. The system pool is tainted CriticalAddonsOnly, and the platform add-ons (ArgoCD, cert-manager, the External Secrets Operator) tolerate the taint. The application chart sets no toleration, no node selector and no affinity, so its pods cannot land on the system pool and NAP launches a node for them. The chart's topology spread across zones and nodes is soft (ScheduleAnyway), so NAP adds the nodes the spread needs rather than leaving pods stuck. NAP nodes carry the label capacity.acme.io/source=nap.
Switches. The unit file has two @optional: lines: capacity_types = ["on-demand"] removes spot capacity (on-demand costs roughly 60% more per node), and sku_families = ["D", "E", "F"] adds the compute-optimized F family.
Terms you will see
| Term | Meaning |
|---|---|
| NodePool | The rules for the nodes NAP may create: families, capacity types, limits, disruption. |
| AKSNodeClass | The node image, OS disk and tags of those nodes. |
| Consolidation | Removing or replacing nodes so the pods run on less capacity. |
| Disruption budget | The share of nodes NAP may take down at the same time, here 10%. |
| Capacity type | Spot or on-demand virtual machines. |
| CPU limit | The total vCPUs the NodePool may hold, set per stage. |
Where to read more
- Azure Web App Blueprint overview for the cluster in context.
- Kubernetes for how the application pods are scheduled and spread.
- Scalable in the BuiltForProd Standard.