Machine Role¶
Implements the CAPI InfraMachine contract (infra-machine.md) for a TerraformMachine. One workspace/state per TerraformMachine. The module creates exactly one node instance and reports its provider ID, addresses, and health. The machine is immutable (spec.source and spec.identityRef cannot change): apply once, refresh/drift for health, destroy on delete.
Inputs¶
Common inputs apply (common.md), including captf_cluster_outputs.
| Name | Type | Required | Value source |
|---|---|---|---|
machine_name | string | yes | Owning CAPI Machine.metadata.name (the TerraformMachine name may differ) |
bootstrap_data | string, sensitive | yes | Base64 of the bootstrap Secret’s value key (below) |
bootstrap_format | string | yes | cloud-config or ignition, from the bootstrap Secret’s format key (below) |
failure_domain | string or null | yes (may be null) | Machine.spec.failureDomain (below) |
kubernetes_version | string or null | yes (may be null) | Machine.spec.version (below) |
control_plane | bool | yes | true when the owning Machine carries the control-plane label (below) |
bootstrap_data¶
Base64 (standard alphabet, padded) of the raw bytes of the value key of the Secret named by Machine.spec.bootstrap.dataSecretName. The bootstrap contract requires a single key value (bootstrap-config.md “BootstrapConfig: data secret”; dataSecretName in machine_types.go Bootstrap).
The controller base64-encodes value when rendering tfvars, always and for every bootstrap provider: CAPRKE2 with gzipUserData: true writes value as raw gzip bytes, which are not valid UTF-8 and so cannot be a JSON/HCL string (see the RKE2ControlPlane guide); encoding unconditionally keeps one rule instead of sniffing content. The input stays sensitive = true.
Modules use one of two patterns:
- pass it unchanged to a
user_data_base64-style argument (for example AWSaws_instance.user_data_base64oraws_launch_template.user_data, which both take base64); or base64decode(var.bootstrap_data)where the argument wants the plain payload — only valid when the payload is UTF-8, that is not gzipped:base64decodefails on non-UTF-8 bytes, so a binary payload must go to a base64-taking argument instead.
Apply waits until the Secret exists (the contract workflow also exits while dataSecretName is nil; infra-machine.md “Typical InfraMachine reconciliation workflow”); dataSecretName: "" (the field allows MinLength=0, machine_types.go) is treated exactly like nil (WaitingForBootstrapData).
Control-plane bootstrap data embeds cluster private keys
For control-plane Machines the payload embeds the cluster CA and service account private keys uncompressed (controlplane_init.go, controlplane_join.go), so it is both larger and more sensitive than a worker’s. Size limits (for example AWS’s 16 KiB user-data) and keeping keys out of readable instance metadata are module business: gzip the payload, or stage it in a secret store and pass only a small stub through bootstrap_data/user-data.
bootstrap_format¶
cloud-config or ignition, from the bootstrap Secret’s format key; cloud-config when absent.
format is not part of the bootstrap contract, which specifies only value (bootstrap-config.md). It is written by the kubeadm bootstrap provider (kubeadmconfig_controller.go writes value and format; enum cloud-config;ignition in kubeadmconfig_types.go Format). Other bootstrap providers need not write it. CAPRKE2 also always writes both keys, with the same enum (rke2config_controller.go; see the RKE2ControlPlane guide). format describes the decoded payload; with CAPRKE2 gzipUserData: true the decoded bytes are gzip of that format.
failure_domain (input)¶
Machine.spec.failureDomain (1–256 chars; machine_types.go). The InfraMachine MUST be placed in this failure domain (infra-machine.md “InfraMachine: failure domain”).
kubernetes_version (input)¶
Machine.spec.version (optional, 1–256 chars; machine_types.go). The value may carry a control-plane-provider-specific distro suffix: with RKE2 it is vX.Y.Z+rke2rN (for example v1.31.4+rke2r1), copied verbatim from RKE2ControlPlane.spec.version to Machine.spec.version (see the RKE2ControlPlane guide). The controller passes this input to the module verbatim, unstripped; a module that uses it for an image lookup or a semver comparison MUST strip the +… build-metadata suffix itself (see the RKE2ControlPlane guide).
control_plane (input)¶
true when the owning Machine carries the cluster.x-k8s.io/control-plane label (MachineControlPlaneLabel, machine_types.go; CAPI’s own util.IsControlPlaneMachine checks only for this label’s presence, util.go). There is no role string in MachineSpec.
Outputs¶
Common outputs apply (health).
| Name | Type | Required | Maps to |
|---|---|---|---|
provider_id | string or null | yes (may be null until known) | TerraformMachine.spec.providerID → Machine.spec.providerID (below) |
addresses | list(object({type=string, address=string})) | yes (may be []) | TerraformMachine.status.addresses → Machine.status.addresses (below) |
failure_domain | string or null | yes (may be null) | TerraformMachine.status.failureDomain → Machine.status.failureDomain (below) |
interruptible | bool or null | declared by every machine module; value optional (null = false) | TerraformMachine.status.interruptible (below) |
provider_id (output)¶
Maps to TerraformMachine.spec.providerID → Machine.spec.providerID. Must equal the Node’s spec.providerID exactly (see “Node providerID matching” below). 1–512 chars (infra-machine.md “InfraMachine: provider ID”; MinLength=1, MaxLength=512 markers in machine_types.go). "" is treated as null.
Written once; a later non-null change is OutputsValid=False/ProviderIDChanged. If it turns null after provisioning, one sample is not trusted: the first reports InfrastructureHealthy=Unknown/ProviderIDMissing, and only a second consecutive missing sample treats the instance as terminated (see common.md “Out-of-band termination”). Either way spec.providerID is kept. We latch it — never clear it once set — because CAPI does not: the Machine controller recopies Machine.spec.providerID from TerraformMachine.spec.providerID on every reconcile rather than latching it once (machine_controller_phases.go), and clearing our side would blank the Machine’s field and stop address/failure-domain copying from the InfraMachine on that same path.
addresses (output)¶
Maps to TerraformMachine.status.addresses → Machine.status.addresses (infra-machine.md “InfraMachine: addresses”). type is one of Hostname, ExternalIP, InternalIP, ExternalDNS, InternalDNS (MachineAddressType enum, common_types.go); address is 1–256 chars; at most 256 entries (MachineAddress and Machine.status.addresses markers, common_types.go and machine_types.go). The controller validates every mapped output against these CRD markers before patching status; a violation is OutputsValid=False/OutputsInvalid, and nothing is written.
Canonical order. Before writing status the controller sorts the module’s list by type precedence InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname, then by address. CAPI itself copies the list verbatim (machine_controller_phases.go), so an unsorted module output would otherwise reach Machine.status.addresses in whatever order the module emitted it. Rationale: RKE2’s registrationMethod: internal-first takes the first address in list order whose type is InternalIP or ExternalIP (this is what the code does, not “internal preferred” as its own docs claim); internal-only-ips needs an InternalIP present, and external-only-ips needs an ExternalIP present (see the RKE2ControlPlane guide) — a fixed, sorted order makes that first-match deterministic instead of depending on module authoring order.
failure_domain (output)¶
Maps to TerraformMachine.status.failureDomain → Machine.status.failureDomain (the actual placement; infra-machine.md “InfraMachine: failure domain”). When the failure_domain input is non-null the output MUST equal it: the contract says the InfraMachine MUST be placed in the requested domain, and the controller enforces it — a mismatch is OutputsValid=False/FailureDomainMismatch and Ready=False, so KCP/MachineDeployment spreading can never be silently broken. The output adds information only when no domain was requested.
interruptible (output)¶
Maps to TerraformMachine.status.interruptible (bool); true for spot/preemptible instances. CAPI reads status.interruptible from the InfraMachine and, when true, sets the Node label cluster.x-k8s.io/interruptible (InterruptibleLabel; machine_controller_noderef.go), which termination handlers and schedulers key on. The output must be declared because the generated root re-exports every contract output unconditionally; a module without spot support declares output "interruptible" { value = false }. Null is written as false.
Provisioned rule¶
status.initialization.provisioned = true when the state Secret carries the inputs-hash annotation of a successful apply, provider_id is non-null, and health.state != "pending"; derived from state only, never from Job history, so it is rebuilt after clusterctl move (evaluated once on the target, then latched).
Latched. The formula applies only until it first holds; once provisioned is true it stays true for the object’s life, whatever later health or outputs say (CAPI never flips infrastructureProvisioned back; machine_controller_phases.go “should not flip back”). A later null provider_id or non-running health is reported on InfrastructureHealthy (for example InstanceTerminated), never by un-provisioning.
While pending, the controller refreshes regardless of drift.intervalSeconds, 30s after the last refresh and backing off to at most 5m (not a field; see common.md). After a successful apply it refreshes once more unless the apply’s own outputs are valid with a health.state other than pending or unknown. Ready=True additionally requires health.healthy and state == "running".
Ready timeline¶
What CAPI mirrors into Machine.status.conditions[InfrastructureReady], which MachineHealthCheck can act on:
| Phase | Ready | Reason |
|---|---|---|
Waiting for owner Machine, cluster infrastructure, exports, bootstrap data | Unknown | WaitingFor* (on the DependenciesReady condition; Ready reason ReadyUnknown) |
| Apply Job started | False | Provisioning |
| Provisioned, healthy | True | Ready |
Provisioned, health not running/healthy | False | NotReady (detail on InfrastructureHealthy) |
| Deleting | False | Deleting |
Unknown rather than False while waiting matters for MHC
unhealthyMachineConditions timeouts count from the condition’s lastTransitionTime, and a reason change does not reset it. Workers created alongside a slow control plane can wait a long time for bootstrap data; starting the False window only at apply start means an MHC InfrastructureReady=False timeout measures apply time, not control-plane bring-up.
After provisioning, Ready is computed from InfrastructureHealthy and deletion only; drift results, drift-Job failures, identity or RBAC problems never turn it False (they are separate conditions), so a credentials outage cannot trigger fleet-wide remediation.
Lifecycle¶
- Immutable. Changes to
spec.sourceandspec.identityRefafter creation are rejected by the webhook, andspec.providerIDcan only be set once, by the controller. Those changes come via MachineDeployment/KCP rollout.spec.jobs,spec.driftandspec.remediationare operational policy and stay mutable: they change how the controller treats the machine (a Job deadline, drift checks, auto-remediation), never what was applied.captf_cluster_outputschanges do not trigger re-apply: the rendered inputs of the first apply (includingbootstrap_dataandcaptf_cluster_outputs) are stored in the object-ownedcaptf-inputs-<kindshort>-<name>Secret and re-fed unchanged on drift and destroy; the effective image and identity are pinned there too, so a later change toTerraformCluster.spec.defaults(for example a ClusterClass in-place update) never changes what runs against an existing machine. Bootstrap data is read once. The kubeadm bootstrap provider does not rewrite a Machine’s bootstrap Secret after creation — it only extends the join token’s TTL in the workload cluster until the Node joins (kubeadmconfig_controller.gorefreshBootstrapTokenIfNeeded) — so “read once” loses nothing. Pools are different (seemachinepool.md). - Gating. Apply waits for: the owner
Machine(looked up through the Machine ownerRef,util.GetOwnerMachine, then the Cluster throughcluster.x-k8s.io/cluster-name; a TerraformMachine never has a Cluster ownerRef);Cluster.status.initialization.infrastructureProvisioned(cluster_types.go); the cluster’sexportsreadable (WaitingForClusterExports); and the bootstrap Secret (WaitingForBootstrapData; the contract workflow also exits whiledataSecretNameis nil, and""counts as nil). The owning Cluster’sspec.infrastructureRefmust be kindTerraformCluster; otherwiseDependenciesReady=False/ClusterNotTerraformand nothing is done. Once those pass, thespec.variablesFromsources must exist and be labeled (DependenciesReady=False/VariablesSourceNotFound) with valid keys and values (False/VariablesInvalid); seecommon.md“User variables”. While gated,Ready=Unknown(timeline above). - Retry. A state Secret that exists without the inputs-hash annotation (a failed or interrupted first apply) means the apply is retried with backoff even though the spec is immutable; each attempt gets a distinct Job name (attempt counter).
- Refresh/drift. On
spec.drift.intervalSeconds:apply -refresh-onlythenplan -detailed-exitcode. Drift is reported only (DriftDetected=True, a negative-polarity condition kept out ofReady); never auto-applied for machines (MachineDriftPolicyhas noactionfield: there is nothing to remediate against for an immutable machine). The refresh updatesaddressesandhealth. A failed drift Job setsDriftJobSucceeded=Falseand nothing else. -
Health → remediation.
healthoutsiderunning/healthyafter provisioning setsInfrastructureHealthy=False, thenReady=False/NotReady, mirrored toMachine.status.conditions[InfrastructureReady]=False, which aMachineHealthCheckcan act on. See Machine Remediation for configuringspec.remediation, the unhealthy threshold and sampling interval, and how a MachineHealthCheck reaches a replacement.healthis purely cloud-side: the controller does not cross-checkMachine.status.nodeRef/nodeInfo(Node existence,machine_controller_noderef.go); a missing or NotReady Node is MHC’snodeStartupTimeout/Node-condition checks’ job. - Delete. Deletion of a TerraformMachine is normally initiated by the Machine controller after drain and pre-terminate hooks. Ordering is CAPI’s: drain (bounded byMachine.spec.deletion.nodeDrainTimeoutSeconds) and volume detach (nodeVolumeDetachTimeoutSeconds) finish before the TerraformMachine is deleted, sodestroynever races a running drain; the Node object is deleted by core-machine only after the TerraformMachine is gone (bounded bynodeDeletionTimeoutSeconds;machine_types.go,machine_controller.go). No contract input carries these timeouts.The webhook allows a direct delete when the object has no Machine ownerRef at all (an orphan: KCP deletes the InfraMachine itself when KubeadmConfig or Machine creation failed,
helpers.go), when the owning Machine is itself deleting, or when the request carriesclusterctl.cluster.x-k8s.io/delete-for-move; it refuses deletion otherwise, so drain and the KCP etcd-membership hook are never bypassed by deleting the infra object directly.Then a
destroyJob runs, rendered from thecaptf-inputs-<kindshort>-<name>Secret (never from the bootstrap Secret or the Cluster, which may already be gone); on success it drops the state, the Lease and the inputs Secret, and removes the finalizer. If the owning Machine or Cluster is gone the same path runs. If there is no state Secret and no active Job (deleted while still gated), the finalizer is removed without a Job. If destroy fails, the object stays withApplyJobSucceeded=False/DestroyFailed(Ready=False) and retries with backoff; a Lease whose holder Job no longer exists is force-unlocked automatically. There is no skip-destroy annotation: a permanently failing destroy is resolved by cleaning up out of band and removing the finalizer by hand, after backing up or un-owning thetfstate-*Secrets, which are otherwise garbage-collected with the object (see the stuck-destroy runbook). A stuck control-plane machine destroy blocks KCP remediation, scale and upgrade until resolved (remediation.go). - The module MUST acceptbootstrap_dataopaquely: it is base64 of the bootstrap payload (Inputs), passed to a base64-taking user-data argument orbase64decode()d into a plain one; the module MUST NOT need to parse the payload. - Taints and labels.Machine.spec.taintsand Machine labels are applied to the Node by core-machine (machine_types.go;machine_controller_noderef.go), not by the infrastructure; the machine role has nonode_labelsinput (pools do, seemachinepool.md); no role has a taints input. - In-place updates (the CAPIInPlaceUpdatesfeature gate, alpha) are unsupported: no field is mutable, and the webhook rejects the MachineSet’s SSA patch; the rollout path is a new Machine.
Node providerID matching (module authors)¶
CAPI links a Machine to its Node only when Node.spec.providerID equals Machine.spec.providerID exactly (machine_controller_noderef.go). Our provider_id output is copied to the Machine verbatim, so the module MUST emit the same string the Node will carry:
- With a cloud-controller-manager (CCM): the CCM sets
Node.spec.providerIDin its own format (for exampleaws:///<zone>/<instance-id>,azure:///subscriptions/...,openstack:///<uuid>,vsphere://<uuid>). Emit exactly that format. - Without a CCM: the kubelet must be started with
--provider-id=<value>(viaKubeadmConfignodeRegistration.kubeletExtraArgs), and the value must be derivable on the instance at boot (an instance-metadata lookup, or a per-machine value baked into the module’s user-data wrapper), becausebootstrap_datais shared per template and opaque to the module. - Not supported in v1: a controller-side patch of
Node.spec.providerIDonce the Node registers, as CAPD does (machine.go), would need a workload-cluster client and a manager flag; the controller does not implement it. In v1 the CCM or kubelet--provider-idmust set it. - RKE2: CAPRKE2 itself never sets kubelet
provider-id,node-ipornode-name(the fields exist in its config but nothing in the repo writes them; see the RKE2ControlPlane guide). Without a CCM, the image MUST setagentConfig.kubelet.extraArgs: [provider-id=<same format>]from instance metadata so the kubelet-registered value matches ourprovider_idoutput exactly. This matters more under RKE2 than KCP: every RKE2 join (control plane or worker) requiresRCP.status.availableServerIPsto be non-empty, which requires at least one Ready control-plane Machine, and a Machine only becomes Ready once its Node exists with a matchingproviderID— so a providerID mismatch on the first control-plane Machine blocks every subsequent join, not just that one Machine’s own readiness.
Autoscale-from-zero (InfraMachineTemplate.status.capacity / nodeInfo; infra-machine.md “InfraMachineTemplate: support cluster autoscaling from zero”) is supported through the image: instance size is fixed inside the module, so the image declares it with the OCI labels io.captf.capacity and io.captf.node-info (see image-contract.md “OCI labels”), and a TerraformMachineTemplate reconciler copies them into status.capacity/status.nodeInfo. An image without the labels leaves both fields unset; the Cluster Autoscaler’s capacity.cluster-autoscaler.kubernetes.io/* annotations on the MachineDeployment/MachineSet remain the fallback.
Control-plane machines¶
When control_plane = true, the module is responsible for registering the instance in the control-plane load balancer’s backend or target group. The ids it needs (load balancer, target group, backend pool — whatever the cluster module’s cloud exposes) come from captf_cluster_outputs (see cluster.md exports, which describes the publishing side of this same convention). Registration is done in the machine module’s own Terraform state, not the cluster module’s, so destroy on the machine deregisters it as an ordinary part of tearing down that state (infra-machine.md).
Register the instance in the load balancer before kubeadm finishes
Ordering matters: the instance MUST be in the load balancer backend before kubeadm init/kubeadm join finish on it, because the contract requires the endpoint to be reachable through the load balancer during control-plane bring-up (status.go; see the KubeadmControlPlane guide on reachability) — Cluster.status.initialization.controlPlaneInitialized never latches if the first control-plane node can’t be reached at the endpoint it just joined. In practice this means the module’s apply must register-then-boot (or register-then-poll-healthy) rather than boot-then-register as an afterthought.
For worker Machines (control_plane = false) no load balancer registration applies; captf_cluster_outputs is still read for any other cluster-level values the module needs.
Minimal skeleton¶
A block written on one line may hold at most one argument (OneLineBlock; tofu validate reports “Invalid single-argument block definition” on a violation), so a variable that needs both type and sensitive, or type and default, must use the multi-line block form below.
variable "captf_contract" { type = string }
variable "captf_cluster" { type = object({ name = string, namespace = string }) }
variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) }
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" { type = map(string) }
variable "machine_name" { type = string }
variable "bootstrap_data" { # base64 of the bootstrap Secret's `value`; e.g. user_data_base64 = var.bootstrap_data
type = string
sensitive = true
}
variable "bootstrap_format" { type = string }
variable "failure_domain" {
type = string
default = null
}
variable "kubernetes_version" {
type = string
default = null
}
variable "control_plane" { type = bool }
# Tag every cloud resource this module creates with captf_tags (common.md
# "captf_tags"); this stub only has to reference it, not create anything.
resource "terraform_data" "tags" { input = var.captf_tags }
output "provider_id" { value = "noop:///${var.captf_object.namespace}/${var.captf_object.name}" }
output "addresses" { value = [{ type = "InternalIP", address = "10.0.0.1" }] }
output "failure_domain" { value = var.failure_domain }
output "interruptible" { value = false }
output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } }