# Machine Pools

This page shows you how to create a `MachinePool` backed by a `TerraformMachinePool`: a group of nodes CAPTF provisions and scales as one native cloud scaling group (an autoscaling group, a scale set, an instance group, or similar), rather than as individual `TerraformMachine`s. It covers fixed and autoscaled replica counts, the group’s reported members, and deletion. See [The Kinds](<https://captf.io/docs/concepts/kinds/index.md>) for how a `TerraformMachinePool` compares to a `TerraformMachine`, and the [machinepool contract](<https://captf.io/docs/module-author/contract/v1alpha1/machinepool/index.md>) for everything the module role must implement.

> [!NOTE]
>
> **Before you begin**
>
> - The provider installed, and a `TerraformCluster` provisioned or being provisioned (see [Installation](<https://captf.io/docs/operator-guide/installation/index.md>)).
> - A `TerraformClusterIdentity` allowed in your namespace (see [Identities and Credentials](<https://captf.io/docs/user-guide/identities/index.md>)).
> - A machinepool-role module image: one that implements the [machinepool contract](<https://captf.io/docs/module-author/contract/v1alpha1/machinepool/index.md>), managing one scaling group and reporting its provider IDs, desired capacity and members.

## Create a MachinePool

A `MachinePool` has a single infrastructure object for its whole group, not one per member: `MachinePool.spec.template.spec.infrastructureRef` names one `TerraformMachinePool` directly, by kind and name, the same way `Cluster.spec.infrastructureRef` names one `TerraformCluster`. There is no per-replica cloning, so you create the `TerraformMachinePool` yourself rather than pointing at a `TerraformMachinePoolTemplate`; a `TerraformMachinePoolTemplate` exists only for a `MachinePool` a ClusterClass topology manages, covered in [Templates and ClusterClass](<https://captf.io/docs/user-guide/clusterclass/index.md>).

The `TerraformMachinePool` must carry the cluster’s `cluster.x-k8s.io/cluster-name` label: CAPTF looks up the owning `Cluster` by that label, not by the `MachinePool`’s `spec.clusterName`. Cluster API’s `MachinePool` controller patches this label on from `spec.clusterName` once the `TerraformMachinePool` exists, but only on its own next reconcile; setting the label yourself avoids that wait.

For example:

```yaml
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachinePool
metadata:
  name: my-cluster-workers
  namespace: team-a
spec:
  clusterName: my-cluster
  replicas: 3
  template:
    spec:
      clusterName: my-cluster
      version: v1.31.4
      bootstrap:
        configRef:
          apiGroup: bootstrap.cluster.x-k8s.io
          kind: KubeadmConfig
          name: my-cluster-workers
      infrastructureRef:
        apiGroup: infrastructure.cluster.x-k8s.io
        kind: TerraformMachinePool
        name: my-cluster-workers
---
apiVersion: bootstrap.cluster.x-k8s.io/v1beta2
kind: KubeadmConfig
metadata:
  name: my-cluster-workers
  namespace: team-a
spec:
  joinConfiguration: {}
---
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
metadata:
  name: my-cluster-workers
  namespace: team-a
  labels:
    cluster.x-k8s.io/cluster-name: my-cluster
spec:
  source:
    image: ghcr.io/example/machinepool-module:v0.1.0
  identityRef:
    name: aws-prod
```

The `KubeadmConfig` is the bootstrap provider’s own object (one per pool, for the same reason: a pool has no per-member Machine to hold one each); its content is the bootstrap provider’s concern, not CAPTF’s. With `spec.replicas` unset, Cluster API defaults it to 1; set it to the fixed size you want, as above. Every field of `TerraformMachinePoolSpec` is mutable: changing `spec.source`, `spec.identityRef`, `spec.variables` or any other field re-applies the module on the next reconcile, and there is no `spec.applyPolicy`. Only an apply that renders a change of the cluster’s exports is guarded (see [When the cluster’s exports change](<#when-the-clusters-exports-change>) and [The Kinds](<https://captf.io/docs/concepts/kinds/index.md>)). Pass module-specific configuration through `spec.variables` or `spec.variablesFrom` as for any other kind (see [Module Variables](<https://captf.io/docs/user-guide/variables/index.md>)); `spec.jobs` tunes the Job the same way it does for a `TerraformMachine` (see [Tuning Jobs](<https://captf.io/docs/user-guide/job-tuning/index.md>)). A pool with none of its own inherits `spec.identityRef`, `spec.jobs` and `spec.drift.intervalSeconds` from the owning `TerraformCluster`’s `spec.defaults`.

`node_labels` (rendered from `spec.template.metadata.labels`, not the `TerraformMachinePool`’s own metadata) and the other inputs a module sees are listed in the [machinepool contract](<https://captf.io/docs/module-author/contract/v1alpha1/machinepool/#inputs>) and the [environment reference](<https://captf.io/docs/reference/environment/index.md>); this page does not restate them.

## Choose fixed replicas or autoscaling

With no autoscaler annotations, `MachinePool.spec.replicas` is the sole source of desired capacity, exactly as in the example above: change it to resize the group.

To let the module’s own native autoscaling policy own the desired count instead, set both annotations Cluster API’s autoscaler contract defines, on the `MachinePool`:

```yaml
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachinePool
metadata:
  name: my-cluster-workers
  namespace: team-a
  annotations:
    cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "3"
    cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "10"
spec:
  clusterName: my-cluster
  template:
    spec:
      clusterName: my-cluster
      version: v1.31.4
      bootstrap:
        configRef:
          apiGroup: bootstrap.cluster.x-k8s.io
          kind: KubeadmConfig
          name: my-cluster-workers
      infrastructureRef:
        apiGroup: infrastructure.cluster.x-k8s.io
        kind: TerraformMachinePool
        name: my-cluster-workers
```

Both annotations must be present, parse as non-negative integers, and satisfy min ≤ max, or the pool reports [`AutoscalingActive=False/AutoscalingAnnotationsInvalid`](<https://captf.io/docs/reference/conditions/#autoscalingactive>) and applies without autoscaling. Valid, the pool reports `AutoscalingActive=True/ReplicasManagedByModule`: the module owns the group’s desired count and its own scaling policy (target tracking, scheduled, or whatever it implements), and the controller claims the `cluster.x-k8s.io/replicas-managed-by` annotation on the `MachinePool` so Cluster API stops treating `spec.replicas` as authoritative. On every reconcile the controller then writes the group’s observed desired capacity back to `MachinePool.spec.replicas`, emitting a [`ReplicasWrittenBack`](<https://captf.io/docs/reference/events/index.md>) event when it changes; see [Annotations, Labels and Finalizers](<https://captf.io/docs/reference/annotations-labels/#cluster-api-and-clusterctl-keys>) for both keys. Leave `spec.replicas` unset in this mode: Cluster API defaults and clamps it from the annotations for the first apply, and the write-back takes over from there. Removing both annotations returns `spec.replicas` to being authoritative and releases `replicas-managed-by`.

If another controller already owns `cluster.x-k8s.io/replicas-managed-by` on the `MachinePool` (its value is something other than the one CAPTF claims), CAPTF leaves the annotation alone and stops writing observed replicas back. The pool reports `AutoscalingActive=False`/`ReplicasManagedExternally`, with a `Warning` event, and `spec.replicas` stays under that other controller’s control. See [`AutoscalingActive`](<https://captf.io/docs/reference/conditions/#autoscalingactive>).

> [!WARNING]
>
> **The Kubernetes Cluster Autoscaler does not drive these pools**
>
> Its `clusterapi` cloud provider requires MachinePool Machines, which CAPTF does not implement. Running it against a CAPTF pool is unsupported, since its `spec.replicas` patches would be overwritten by the write-back above.

## Add an autoscaled pool to a generated cluster

This walks through adding an autoscaled `MachinePool` to a cluster generated from the default flavor ([Templates and ClusterClass](<https://captf.io/docs/user-guide/clusterclass/index.md>)), since none of the shipped flavors creates one on their own.

Generate the default flavor with no workers; a `MachineDeployment` is still created, with `spec.replicas: 0`, rather than omitted:

```sh
clusterctl generate cluster my-cluster --infrastructure terraform \
  --target-namespace team-a \
  --kubernetes-version v1.31.4 \
  --control-plane-machine-count 1 --worker-machine-count 0 \
  | kubectl apply -f -
```

Add the pool alongside it: a `MachinePool` with the autoscaler annotations, its `KubeadmConfig` (the flavor’s bootstrap provider is kubeadm), and the `TerraformMachinePool`:

```yaml
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachinePool
metadata:
  name: my-cluster-workers
  namespace: team-a
  annotations:
    cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "2"
    cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "5"
spec:
  clusterName: my-cluster
  template:
    spec:
      clusterName: my-cluster
      version: v1.31.4
      bootstrap:
        configRef:
          apiGroup: bootstrap.cluster.x-k8s.io
          kind: KubeadmConfig
          name: my-cluster-workers
      infrastructureRef:
        apiGroup: infrastructure.cluster.x-k8s.io
        kind: TerraformMachinePool
        name: my-cluster-workers
---
apiVersion: bootstrap.cluster.x-k8s.io/v1beta2
kind: KubeadmConfig
metadata:
  name: my-cluster-workers
  namespace: team-a
spec:
  joinConfiguration: {}
---
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
metadata:
  name: my-cluster-workers
  namespace: team-a
  labels:
    cluster.x-k8s.io/cluster-name: my-cluster
spec:
  source:
    image: ghcr.io/example/machinepool-module:v0.1.0
  identityRef:
    name: aws-prod
```

`spec.replicas` is left unset on the `MachinePool`: with both autoscaler annotations present and valid, Cluster API defaults and clamps it from `min-size`/`max-size` for the first apply, and the pool’s own write-back takes over from there (see [Choose fixed replicas or autoscaling](<#choose-fixed-replicas-or-autoscaling>) above). Apply the three objects, then confirm as in [Confirm it worked](<#confirm-it-worked>) below.

## Set the membership refresh interval

Between applies, the controller runs a refresh to pick up members joining or leaving the group, on `spec.membershipRefreshIntervalSeconds` (15–86400 seconds; unset or 0 means 60):

```yaml
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
spec:
  membershipRefreshIntervalSeconds: 30
```

A new member is unschedulable until its provider ID reaches `spec.providerIDList`, so a shorter interval gets new nodes ready for workloads sooner, at the cost of more frequent Jobs. The controller also refreshes right after every apply and, while the group has not converged (`spec.providerIDList`’s length differs from `status.replicas`), every 30 seconds, or the configured interval instead when it is shorter.

## Check the group’s members

```sh
kubectl get terraformmachinepool <name> -n <namespace>
```

shows the owning cluster, `MachinePool`, desired replicas and the `Ready` condition. `spec.providerIDList` is every non-terminated member, sorted and deduplicated; `status.replicas` is the desired capacity as of the last refresh; `status.instances` is the module’s own per-member detail (provider ID, an optional instance ID, addresses, failure domain and health state), capped at 1000 entries:

```sh
kubectl get terraformmachinepool <name> -n <namespace> \
  -o jsonpath='{.status.instances}'
```

See [TerraformMachinePool](<https://captf.io/docs/reference/resources/terraformmachinepool/#status>) for every status field. A pool reports provisioned, and the `Ready` condition (mirrored onto the `MachinePool`’s `InfrastructureReady`) true, once its state carries a successful apply and its module reports a health state other than pending; unlike a fixed-replica machine, that latch does not depend on `spec.providerIDList` being non-empty, since a pool may legitimately scale to zero.

## Drift on a pool

> [!NOTE]
>
> **A pool’s drift check cannot be disabled**
>
> `spec.drift.intervalSeconds: 0` falls back to the manager’s default interval rather than turning checks off.

The reason is that the drift Job’s own refresh is what feeds a plan; without it, a cloud-side scaling change would never register as drift. With `spec.drift.action: Remediate`, a detected difference re-applies the pool’s current inputs; with autoscaling enabled the module is responsible for excluding its own desired-count attribute from that plan, or every cloud-side scale reports as drift. See [Drift](<https://captf.io/docs/user-guide/drift/index.md>) for setting the interval and action, and [Drift and Health](<https://captf.io/docs/concepts/drift-and-health/index.md>) for how a check runs and feeds health.

## When the cluster’s exports change

A pool’s inputs include the cluster’s exports (`captf_cluster_outputs`). When they change, the pool applies the change, and that apply is guarded: if its plan deletes or replaces anything, it stops before the apply step and the change is **held**. Nothing else is guarded: the first apply, bootstrap rotations, version rolls, replica changes and spec edits apply as usual while the exports are unchanged.

While a change is held, the pool keeps applying everything else with the exports of its last successful apply, and `ApplyJobSucceeded` shows `False`/`DestructivePlanBlocked` (so the pool’s `Ready` is `False`) with what the plan would delete or replace and the pool’s approval hash. After you have read the plan, approve it:

```sh
kubectl annotate terraformmachinepool <name> -n <namespace> \
  captf.io/approve-destructive-plan=<approval-hash> --overwrite
```

The approval hash is the inputs hash without `bootstrap_data`, so it survives bootstrap rotations and changes on any other input change. Anyone who may patch the pool may approve it. The annotation is removed after the apply it approved succeeds. If the exports return to the applied ones, the change is withdrawn and its approval removed; if they move to another change, the old approval is removed. After a guarded apply fails part-way, every apply of the pool is guarded until one succeeds. See [The destructive-plan guard](<https://captf.io/docs/concepts/approvals/destructive-guard/#machine-pools>) for the full behavior and its limits.

## Delete a MachinePool

Deleting the `MachinePool` deletes its `KubeadmConfig` and `TerraformMachinePool` with it. Nothing blocks a `TerraformMachinePool`’s deletion: it destroys the group from its durable inputs and removes its finalizer once the destroy Job succeeds, whether or not the owning `MachinePool` or `Cluster` still exist. See [the reconcile lifecycle](<https://captf.io/docs/concepts/lifecycle/index.md>) for how deletion and finalizers work across every kind.

## Confirm it worked

> [!TIP]
>
> ```sh
> kubectl get terraformmachinepool <name> -n <namespace>
> ```
>
> `Ready` reads `True` once the group is provisioned, and `Replicas` shows the desired capacity. `kubectl get machinepool <name> -n <namespace>` shows the same replica count and `InfrastructureReady=True` once Cluster API has copied `spec.providerIDList` and `status.replicas` across, which happens only once the workload cluster is reachable.

> [!NOTE]
>
> **See also**
>
> - [The Kinds](<https://captf.io/docs/concepts/kinds/index.md>) for how a `TerraformMachinePool`’s mutability differs from a `TerraformMachine`’s.
> - [The machinepool contract](<https://captf.io/docs/module-author/contract/v1alpha1/machinepool/index.md>) for every input and output a module role must implement.
> - [Drift](<https://captf.io/docs/user-guide/drift/index.md>) and [Drift and Health](<https://captf.io/docs/concepts/drift-and-health/index.md>).
> - [Templates and ClusterClass](<https://captf.io/docs/user-guide/clusterclass/index.md>) for a `TerraformMachinePoolTemplate` used through a ClusterClass topology.
> - [Module Variables](<https://captf.io/docs/user-guide/variables/index.md>) and [Tuning Jobs](<https://captf.io/docs/user-guide/job-tuning/index.md>).
> - [TerraformMachinePool](<https://captf.io/docs/reference/resources/terraformmachinepool/index.md>), [Conditions](<https://captf.io/docs/reference/conditions/#autoscalingactive>), [Events](<https://captf.io/docs/reference/events/index.md>) and [Annotations, Labels and Finalizers](<https://captf.io/docs/reference/annotations-labels/#cluster-api-and-clusterctl-keys>).
