Machine Remediation¶
A TerraformMachine reports its instance’s health on the InfrastructureHealthy condition. By itself that only ever turns the owning Machine’s InfrastructureReady condition False; nothing acts on it unless something is watching. spec.remediation optionally asks Cluster API to replace an unhealthy instance, through a MachineHealthCheck and the Machine’s owner. This page covers configuring it, how it reaches a replacement, and what it does to the Machine along the way. See Drift and Health for how health itself is computed.
Before you begin
- A provisioned
TerraformMachinewhose owning Machine is selected by aMachineHealthCheckwith anunhealthyMachineConditionscheck onInfrastructureReady. kubectlaccess to edit theTerraformMachineand read its owning Machine.
Enable remediation¶
spec.remediation.annotateMachine defaults to false, meaning CAPTF never touches the Machine. Set it to true to have CAPTF request remediation once an instance has been unhealthy for long enough:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
remediation:
annotateMachine: true
spec.remediation is operational policy: it stays mutable for the machine’s whole life, unlike spec.source and spec.identityRef, which are fixed at creation.
Tune the threshold and sampling interval¶
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
remediation:
annotateMachine: true
unhealthyThreshold: 5
healthCheckIntervalSeconds: 120
unhealthyThreshold(1-100, default 3) is how many consecutive unhealthy health samples are needed before CAPTF signals remediation. A sample is one completed refresh or drift Job; a single transient reading is never enough on its own, since once aMachineHealthCheckacts on the signal, the Machine is replaced. A terminated instance skips the threshold; see Terminated instances below.healthCheckIntervalSeconds(60-86400, default 300) is how often a provisioned instance is refreshed to resample its health whileannotateMachineistrue, independent ofspec.drift.intervalSeconds. WithannotateMachineset tofalseit has no effect, and health is then re-read only at the drift (or refresh) cadence set byspec.drift.intervalSeconds; see Drift for that setting, including howintervalSeconds: 0stops health sampling entirely onceannotateMachineisfalse.
How it works with a MachineHealthCheck¶
- Health outside
running/healthy, read after provisioning, turnsInfrastructureHealthyFalse, which turnsReadyFalse, which Cluster API mirrors into the owning Machine’sInfrastructureReadycondition. - A
MachineHealthCheckwithspec.checks.unhealthyMachineConditions: [{type: InfrastructureReady, status: "False", timeoutSeconds: <N>}]selecting the Machine starts its own timeout onceInfrastructureReadyturnsFalse. - When the timeout (or, with
annotateMachine, the annotation below) marks the Machine unhealthy,MachineHealthChecksets the Machine’sOwnerRemediatedcondition. Only the Machine’s owner acts on that condition: aMachineDeployment’sMachineSet, or a control-plane provider that implements remediation, such asKubeadmControlPlaneorRKE2ControlPlane. A single-replica control plane refuses to remediate itself. - The owner that acts replaces the Machine (and, through its deletion, the
TerraformMachine) with a new one rather than repairing it in place:spec.sourceis immutable, so there is nothing for CAPTF to patch.
annotateMachine feeds the flow, it does not replace it
It gives the flow a faster, more specific signal than the timeout alone (see the next section).
The remediate-machine annotation¶
With annotateMachine: true, once unhealthyThreshold consecutive samples are unhealthy, degraded or stopped (or the instance is terminated), CAPTF sets cluster.x-k8s.io/remediate-machine on the owning Machine. This asks the MachineHealthCheck reconciler to treat the Machine as unhealthy at once, bypassing unhealthyMachineConditions’ timeout; remediation still goes through the owner as described above. CAPTF marks the annotation as its own with a second annotation, captf.io/remediation-requested, whose value is why it was set, for example InstanceUnhealthy for 5 consecutive samples or the instance is terminated.
CAPTF removes both annotations once the instance reads Healthy again and the Machine is not being deleted, withdrawing a request that has not yet been acted on. It never removes cluster.x-k8s.io/remediate-machine when captf.io/remediation-requested is absent: an annotation set by someone else is left alone.
See Annotations, Labels and Finalizers for both annotations’ full definitions.
Terminated instances¶
A terminated instance skips the threshold
CAPTF sets the annotations on the very first sample that reads the instance terminated, since there is nothing to wait for and no risk of a transient reading.
The exception is a provider_id output that goes missing: the first missing sample only sets InfrastructureHealthy=Unknown/ProviderIDMissing, and the second consecutive one is read as terminated and requests remediation.
Confirm it worked¶
kubectl get machine <name> -n <namespace> -o jsonpath='{.metadata.annotations}'showscluster.x-k8s.io/remediate-machineandcaptf.io/remediation-requestedonce a request is made, and neither once the instance recovers or is replaced.- A
RemediationRequestedevent on theTerraformMachinemarks a new request, andRemediationWithdrawnmarks a withdrawal; see Events. captf_remediation_requests_total{action="requested"|"withdrawn"}counts both; see Metrics.