Machine Remediation
A TerraformMachine reports its instance’s health on the
InfrastructureHealthy condition. By itself that only ever turns the
owning Machine’s InfrastructureReady condition False; nothing acts on
it unless something is watching. spec.remediation optionally asks
Cluster API to replace an unhealthy instance, through a
MachineHealthCheck and the Machine’s owner. This page covers
configuring it, how it reaches a replacement, and what it does to the
Machine along the way. See Drift and
Health for how health itself is
computed.
Before you begin
- A provisioned
TerraformMachinewhose owning Machine is selected by aMachineHealthCheckwith anunhealthyMachineConditionscheck onInfrastructureReady. kubectlaccess to edit theTerraformMachineand read its owning Machine.
Enable remediation
spec.remediation.annotateMachine defaults to false, meaning CAPTF
never touches the Machine. Set it to true to have CAPTF request
remediation once an instance has been unhealthy for long enough:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
remediation:
annotateMachine: true
spec.remediation is operational policy: it stays mutable for the
machine’s whole life, unlike spec.source and spec.identityRef, which
are fixed at creation.
Tune the threshold and sampling interval
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
remediation:
annotateMachine: true
unhealthyThreshold: 5
healthCheckIntervalSeconds: 120
unhealthyThreshold(1-100, default 3) is how many consecutive unhealthy health samples are needed before CAPTF signals remediation. A sample is one completed refresh or drift Job; a single transient reading is never enough on its own, since once aMachineHealthCheckacts on the signal, the Machine is replaced. A terminated instance skips the threshold; see Terminated instances below.healthCheckIntervalSeconds(60-86400, default 300) is how often a provisioned instance is refreshed to resample its health whileannotateMachineistrue, independent ofspec.drift.intervalSeconds. WithannotateMachineset tofalseit has no effect, and health is then re-read only at the drift (or refresh) cadence set byspec.drift.intervalSeconds; see Drift for that setting, including howintervalSeconds: 0stops health sampling entirely onceannotateMachineisfalse.
How it works with a MachineHealthCheck
- Health outside
running/healthy, read after provisioning, turnsInfrastructureHealthyFalse, which turnsReadyFalse, which Cluster API mirrors into the owning Machine’sInfrastructureReadycondition. - A
MachineHealthCheckwithspec.checks.unhealthyMachineConditions: [{type: InfrastructureReady, status: "False", timeoutSeconds: <N>}]selecting the Machine starts its own timeout onceInfrastructureReadyturnsFalse. - When the timeout (or, with
annotateMachine, the annotation below) marks the Machine unhealthy,MachineHealthChecksets the Machine’sOwnerRemediatedcondition. Only the Machine’s owner acts on that condition: aMachineDeployment’sMachineSet, or a control-plane provider that implements remediation, such asKubeadmControlPlaneorRKE2ControlPlane. A single-replica control plane refuses to remediate itself. - The owner that acts replaces the Machine (and, through its deletion,
the
TerraformMachine) with a new one rather than repairing it in place:spec.sourceis immutable, so there is nothing for CAPTF to patch.
annotateMachine does not replace this flow; it feeds it a faster, more
specific signal than the timeout alone (see the next section).
The remediate-machine annotation
With annotateMachine: true, once unhealthyThreshold consecutive
samples are unhealthy, degraded or stopped (or the instance is
terminated), CAPTF sets cluster.x-k8s.io/remediate-machine on the
owning Machine. This asks the MachineHealthCheck reconciler to treat
the Machine as unhealthy at once, bypassing unhealthyMachineConditions’
timeout; remediation still goes through the owner as described above.
CAPTF marks the annotation as its own with a second annotation,
captf.io/remediation-requested, whose value is why it was set, for
example InstanceUnhealthy for 5 consecutive samples or the instance is terminated.
CAPTF removes both annotations once the instance reads Healthy again
and the Machine is not being deleted, withdrawing a request that has not
yet been acted on. It never removes cluster.x-k8s.io/remediate-machine
when captf.io/remediation-requested is absent: an annotation set by
someone else is left alone.
See Annotations, Labels and Finalizers for both annotations’ full definitions.
Terminated instances
A terminated instance skips unhealthyThreshold entirely: CAPTF sets the
annotations on the very first sample that reads the instance terminated,
since there is nothing to wait for and no risk of a transient reading.
Confirm it worked
kubectl get machine <name> -n <namespace> -o jsonpath='{.metadata.annotations}'showscluster.x-k8s.io/remediate-machineandcaptf.io/remediation-requestedonce a request is made, and neither once the instance recovers or is replaced.- A
RemediationRequestedevent on theTerraformMachinemarks a new request, andRemediationWithdrawnmarks a withdrawal; see Events. captf_remediation_requests_total{action="requested"|"withdrawn"}counts both; see Metrics.
See also
- Drift and Health for how health samples are produced.
- Drift for
spec.drift.intervalSeconds, which paces health sampling whenannotateMachineisfalse. - API Reference for
MachineRemediation’s full field list. - Conditions for
InfrastructureHealthy’s reasons.