MachinePool¶
The ghcr.io/captf-io/azure-machinepool image implements the machinepool role on Azure: one uniform Linux virtual machine scale set of worker nodes per MachinePool, in the cluster’s resource group, with an Azure Autoscale setting that holds its capacity. See Machine Pools for the objects to create.
What it creates¶
| Resource | Purpose | When |
|---|---|---|
azurerm_linux_virtual_machine_scale_set.pool_scale_set | The workers: manual upgrades, no overprovisioning, no single placement group, in the worker subnet, security groups and identity, with termination notifications | Always |
azurerm_monitor_autoscale_setting.pool_autoscale_setting | Holds the scale set’s capacity: pinned to replicas, or between the autoscaling bounds with CPU rules | Always |
terraform_data.pool_default_zones | The cluster’s zones as of the first apply, kept for a pool without its own failure domains | Always (used without failure_domains) |
Without exports (an externally managed cluster and no external_cluster_exports) the scale set and autoscale setting are not created and a precondition fails.
Inputs¶
Contract inputs used: captf_cluster_outputs (the cluster’s exports), captf_tags, machinepool_name, replicas, bootstrap_data, bootstrap_format, failure_domains, cluster_failure_domains, kubernetes_version, node_labels and autoscaling. captf_contract is validated; captf_cluster and captf_object are not used.
User variables, set with spec.variables on the TerraformMachinePool (variables.tf):
| Name | Type | Default | Description |
|---|---|---|---|
accelerated_networking | bool | true | Accelerated networking on the instances’ NICs |
additional_tags | map(string) | {} | Extra Azure tags on the scale set and autoscale setting. Keys starting with captf.io_ or captf.io/ (any case) are rejected; at most 44 |
autoscaling_scale_in_cpu_percent | number | 25 | With autoscaling, scale in by one instance below this average CPU over 10 minutes; must be below the scale-out threshold |
autoscaling_scale_out_cpu_percent | number | 75 | With autoscaling, scale out by one instance above this average CPU over 10 minutes |
boot_diagnostics | bool | true | Keep the serial console logs in Azure-managed storage |
encryption_at_host | bool | false | Encrypt temporary disks and caches on the host too; needs the EncryptionAtHost feature |
external_cluster_exports | any | null | Exports (schema captf.io/azure-cluster/v1) for an externally managed TerraformCluster |
image_id | string | null | Required. A managed image, Compute Gallery image (version), or community or shared gallery image (version) ID. {version} and {semver} become the pool’s version, v1.31.4 and 1.31.4, without any +suffix |
ip_forwarding | bool | false | IP forwarding on the NICs, for CNIs that route pod addresses natively |
os_disk_size_gib | number | 128 | OS disk size per instance, 30 to 4095 GiB |
os_disk_storage_account_type | string | "Premium_LRS" | Standard_LRS, StandardSSD_LRS, StandardSSD_ZRS, Premium_LRS or Premium_ZRS |
spot | bool | false | Spot instances, deleted on eviction, at most the on-demand price |
trusted_launch | bool | false | Secure boot and vTPM; needs a generation 2 image built for trusted launch |
vm_size | string | "Standard_D4s_v5" | Azure VM size of the instances |
Outputs¶
| Output | Value |
|---|---|
provider_id | azure:///subscriptions/<subscription>/resourceGroups/<group, lowercase>/providers/Microsoft.Compute/virtualMachineScaleSets/<scale set>; changes when a new generation replaces the scale set |
provider_id_list | Every instance the scale set lists, whatever its power state, as <provider_id>/virtualMachines/<instance ID>, the format cloud-provider-azure writes to the Node |
replicas | The scale set’s capacity as last refreshed, which Azure Autoscale sets; null once the scale set is gone |
instances | Per instance: provider_id, instance_id, addresses (InternalIP, Hostname), failure_domain (its zone) and state |
health | See Health |
autoscale_setting_id | Extra: the autoscale setting’s ARM ID |
dropped_node_labels | Extra: node_labels keys left out because the kubelet may not set them on itself |
scale_set_id, scale_set_name | Extra: the current scale set’s ARM ID and name |
The scale set is named <machinepool_name, made valid, at most 41 characters>-<8 hex characters>, and its instances’ hostnames start with that name; cloud-provider-azure finds an instance by its hostname, so name the Nodes after it (nodeRegistration.name: '{{ ds.meta_data["local_hostname"] }}').
Health¶
The module lists the scale set first and reads its instances only when the listing finds it. Each instance’s power state maps like a machine’s; the pool follows the shared order:
| Azure state | Contract state | Reason |
|---|---|---|
| Scale set gone (a refresh dropped it) | terminated | ScaleSetNotFound |
| Capacity 0 | running, healthy | none |
| Not listed yet (first apply, new generation) or no instances | pending | NoMembers; provider_id_list is [] until the next refresh |
| Some instance stopped, deallocated or in an unknown state | the worst of degraded, stopped, unknown | <reason>:<hostname> per instance |
| Instances starting, or the count off the capacity | running, not healthy | PowerState/starting:<hostname> per instance, ScalingInProgress |
| Every instance running at capacity | running, healthy | none |
Azure lists an instance until it is deleted, so no instance maps to terminated: a deleted instance leaves provider_id_list.
Lifecycle¶
| Change | Effect |
|---|---|
bootstrap_data (a token rotation, about every 7.5 minutes), node_labels, image_id, vm_size, os_disk_size_gib, accelerated_networking, ip_forwarding, boot_diagnostics, encryption_at_host, the exports, tags | The scale set’s model updates in place; new instances use it, running ones are left alone |
replicas, autoscaling disabled | The autoscale setting pins the new capacity; Azure Autoscale applies it within about a minute |
autoscaling, autoscaling_scale_*_cpu_percent | The autoscale setting gets the new bounds and rules |
kubernetes_version, compared verbatim (a +rke2rN bump included) | A new scale set at the current capacity, then the old one is deleted |
failure_domains, spot, trusted_launch, os_disk_storage_account_type | A new scale set, as for a version change: Azure cannot change these on a scale set |
cluster_failure_domains | Nothing: a pool without its own failure_domains keeps the cluster’s zones as of its first apply |
The provider turns off roll_instances_when_required and reimage_on_manual_upgrade: with azurerm’s defaults every token rotation would reimage every instance. So nothing rolls by itself, and a version change creates a new scale set with create_before_destroy; the pool briefly needs twice its quota. There is no drain: pools have no Machines. Scheduled Events announce every deletion 5 minutes ahead, for a termination handler that drains.
Bootstrap¶
- Delivery. As for the machine role: custom data,
cloud-configin a MIME multipart message after a boot hook, and Ignition unchanged (gzipped Ignition fails a precondition). The boot hook writes/etc/kubernetes/azure.jsonfor the worker identity when the file is absent and renders the node labels as the shared behavior describes; Ignition with labels left to render fails a precondition. - Size limit. 65,535 bytes of custom data, 87,380 base64 characters; a precondition stops a larger message.
- Who can read it. As for the machine role: not through the instance metadata service or a read of the scale set; on the node only root; and the pool’s state Secret on the management cluster.
Limitations¶
Keep pools to about 200 instances
A scale set holds at most 1,000 instances, but every refresh reads each instance’s NICs with one Azure call; keep pools to about 200 instances, or raise spec.membershipRefreshIntervalSeconds.
- Workers only: the control plane uses the machine role.
- Changes reach new instances only; a version change is the way to roll the pool.
- Capacity changes go through Azure Autoscale and take about a minute to reach the scale set;
replicasfollows at the next refresh. - The scale set read fails if an instance disappears between listing it and reading its NICs; the next refresh succeeds.
- Azure public cloud only.
Exceptions¶
pool/autoscaling-ignore-changes, atfcapi-lintwarning, is allowed: the scale set ignores changes toinstances, its desired count, which the check’s pattern does not know, and Azure Autoscale holds the capacity in both modes.- Version rolls by generation. The roll is a new scale set name with
create_before_destroy, notterraform_data.kubernetes_version_roll: a replacement under the same name would collide with the old scale set. - Other replacements.
failure_domains,spot,trusted_launchandos_disk_storage_account_typereplace the scale set, because azurerm 5.7.0 cannot update them in place. replicasisnullonce the scale set is gone: the capacity is then unknown, and an emptyprovider_id_listwith an unknown capacity keeps Cluster API’s guard against deleting every Node engaged.- The
membership_excludes_terminatedtest asserts that stopped and deallocated instances stay members, since Azure has no terminated instance state; theScaleSetNotFoundreading has no test.
Example¶
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
metadata:
name: demo-pool-0
namespace: team-a
labels:
cluster.x-k8s.io/cluster-name: demo
spec:
source:
image: ghcr.io/captf-io/azure-machinepool:v0.1.0-opentofu
variables:
image_id: /communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver}
autoscaling_scale_out_cpu_percent: 70