Skip to content

MachinePool

The ghcr.io/captf-io/azure-machinepool image implements the machinepool role on Azure: one uniform Linux virtual machine scale set of worker nodes per MachinePool, in the cluster’s resource group, with an Azure Autoscale setting that holds its capacity. See Machine Pools for the objects to create.

What it creates

Resource Purpose When
azurerm_linux_virtual_machine_scale_set.pool_scale_set The workers: manual upgrades, no overprovisioning, no single placement group, in the worker subnet, security groups and identity, with termination notifications Always
azurerm_monitor_autoscale_setting.pool_autoscale_setting Holds the scale set’s capacity: pinned to replicas, or between the autoscaling bounds with CPU rules Always
terraform_data.pool_default_zones The cluster’s zones as of the first apply, kept for a pool without its own failure domains Always (used without failure_domains)

Without exports (an externally managed cluster and no external_cluster_exports) the scale set and autoscale setting are not created and a precondition fails.

Inputs

Contract inputs used: captf_cluster_outputs (the cluster’s exports), captf_tags, machinepool_name, replicas, bootstrap_data, bootstrap_format, failure_domains, cluster_failure_domains, kubernetes_version, node_labels and autoscaling. captf_contract is validated; captf_cluster and captf_object are not used.

User variables, set with spec.variables on the TerraformMachinePool (variables.tf):

Name Type Default Description
accelerated_networking bool true Accelerated networking on the instances’ NICs
additional_tags map(string) {} Extra Azure tags on the scale set and autoscale setting. Keys starting with captf.io_ or captf.io/ (any case) are rejected; at most 44
autoscaling_scale_in_cpu_percent number 25 With autoscaling, scale in by one instance below this average CPU over 10 minutes; must be below the scale-out threshold
autoscaling_scale_out_cpu_percent number 75 With autoscaling, scale out by one instance above this average CPU over 10 minutes
boot_diagnostics bool true Keep the serial console logs in Azure-managed storage
encryption_at_host bool false Encrypt temporary disks and caches on the host too; needs the EncryptionAtHost feature
external_cluster_exports any null Exports (schema captf.io/azure-cluster/v1) for an externally managed TerraformCluster
image_id string null Required. A managed image, Compute Gallery image (version), or community or shared gallery image (version) ID. {version} and {semver} become the pool’s version, v1.31.4 and 1.31.4, without any +suffix
ip_forwarding bool false IP forwarding on the NICs, for CNIs that route pod addresses natively
os_disk_size_gib number 128 OS disk size per instance, 30 to 4095 GiB
os_disk_storage_account_type string "Premium_LRS" Standard_LRS, StandardSSD_LRS, StandardSSD_ZRS, Premium_LRS or Premium_ZRS
spot bool false Spot instances, deleted on eviction, at most the on-demand price
trusted_launch bool false Secure boot and vTPM; needs a generation 2 image built for trusted launch
vm_size string "Standard_D4s_v5" Azure VM size of the instances

Outputs

Output Value
provider_id azure:///subscriptions/<subscription>/resourceGroups/<group, lowercase>/providers/Microsoft.Compute/virtualMachineScaleSets/<scale set>; changes when a new generation replaces the scale set
provider_id_list Every instance the scale set lists, whatever its power state, as <provider_id>/virtualMachines/<instance ID>, the format cloud-provider-azure writes to the Node
replicas The scale set’s capacity as last refreshed, which Azure Autoscale sets; null once the scale set is gone
instances Per instance: provider_id, instance_id, addresses (InternalIP, Hostname), failure_domain (its zone) and state
health See Health
autoscale_setting_id Extra: the autoscale setting’s ARM ID
dropped_node_labels Extra: node_labels keys left out because the kubelet may not set them on itself
scale_set_id, scale_set_name Extra: the current scale set’s ARM ID and name

The scale set is named <machinepool_name, made valid, at most 41 characters>-<8 hex characters>, and its instances’ hostnames start with that name; cloud-provider-azure finds an instance by its hostname, so name the Nodes after it (nodeRegistration.name: '{{ ds.meta_data["local_hostname"] }}').

Health

The module lists the scale set first and reads its instances only when the listing finds it. Each instance’s power state maps like a machine’s; the pool follows the shared order:

Azure state Contract state Reason
Scale set gone (a refresh dropped it) terminated ScaleSetNotFound
Capacity 0 running, healthy none
Not listed yet (first apply, new generation) or no instances pending NoMembers; provider_id_list is [] until the next refresh
Some instance stopped, deallocated or in an unknown state the worst of degraded, stopped, unknown <reason>:<hostname> per instance
Instances starting, or the count off the capacity running, not healthy PowerState/starting:<hostname> per instance, ScalingInProgress
Every instance running at capacity running, healthy none

Azure lists an instance until it is deleted, so no instance maps to terminated: a deleted instance leaves provider_id_list.

Lifecycle

Change Effect
bootstrap_data (a token rotation, about every 7.5 minutes), node_labels, image_id, vm_size, os_disk_size_gib, accelerated_networking, ip_forwarding, boot_diagnostics, encryption_at_host, the exports, tags The scale set’s model updates in place; new instances use it, running ones are left alone
replicas, autoscaling disabled The autoscale setting pins the new capacity; Azure Autoscale applies it within about a minute
autoscaling, autoscaling_scale_*_cpu_percent The autoscale setting gets the new bounds and rules
kubernetes_version, compared verbatim (a +rke2rN bump included) A new scale set at the current capacity, then the old one is deleted
failure_domains, spot, trusted_launch, os_disk_storage_account_type A new scale set, as for a version change: Azure cannot change these on a scale set
cluster_failure_domains Nothing: a pool without its own failure_domains keeps the cluster’s zones as of its first apply

The provider turns off roll_instances_when_required and reimage_on_manual_upgrade: with azurerm’s defaults every token rotation would reimage every instance. So nothing rolls by itself, and a version change creates a new scale set with create_before_destroy; the pool briefly needs twice its quota. There is no drain: pools have no Machines. Scheduled Events announce every deletion 5 minutes ahead, for a termination handler that drains.

Bootstrap

  • Delivery. As for the machine role: custom data, cloud-config in a MIME multipart message after a boot hook, and Ignition unchanged (gzipped Ignition fails a precondition). The boot hook writes /etc/kubernetes/azure.json for the worker identity when the file is absent and renders the node labels as the shared behavior describes; Ignition with labels left to render fails a precondition.
  • Size limit. 65,535 bytes of custom data, 87,380 base64 characters; a precondition stops a larger message.
  • Who can read it. As for the machine role: not through the instance metadata service or a read of the scale set; on the node only root; and the pool’s state Secret on the management cluster.

Limitations

Keep pools to about 200 instances

A scale set holds at most 1,000 instances, but every refresh reads each instance’s NICs with one Azure call; keep pools to about 200 instances, or raise spec.membershipRefreshIntervalSeconds.

  • Workers only: the control plane uses the machine role.
  • Changes reach new instances only; a version change is the way to roll the pool.
  • Capacity changes go through Azure Autoscale and take about a minute to reach the scale set; replicas follows at the next refresh.
  • The scale set read fails if an instance disappears between listing it and reading its NICs; the next refresh succeeds.
  • Azure public cloud only.

Exceptions

  • pool/autoscaling-ignore-changes, a tfcapi-lint warning, is allowed: the scale set ignores changes to instances, its desired count, which the check’s pattern does not know, and Azure Autoscale holds the capacity in both modes.
  • Version rolls by generation. The roll is a new scale set name with create_before_destroy, not terraform_data.kubernetes_version_roll: a replacement under the same name would collide with the old scale set.
  • Other replacements. failure_domains, spot, trusted_launch and os_disk_storage_account_type replace the scale set, because azurerm 5.7.0 cannot update them in place.
  • replicas is null once the scale set is gone: the capacity is then unknown, and an empty provider_id_list with an unknown capacity keeps Cluster API’s guard against deleting every Node engaged.
  • The membership_excludes_terminated test asserts that stopped and deallocated instances stay members, since Azure has no terminated instance state; the ScaleSetNotFound reading has no test.

Example

terraformmachinepool.yaml
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
metadata:
  name: demo-pool-0
  namespace: team-a
  labels:
    cluster.x-k8s.io/cluster-name: demo
spec:
  source:
    image: ghcr.io/captf-io/azure-machinepool:v0.1.0-opentofu
  variables:
    image_id: /communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver}
    autoscaling_scale_out_cpu_percent: 70