Introduction
Cluster API Provider Terraform (CAPTF) is a Cluster API infrastructure provider. It provisions and manages cluster infrastructure by running Terraform or OpenTofu modules as Kubernetes Jobs, instead of implementing cloud-specific logic in Go. This page introduces CAPTF’s architecture, then points you at the part of the book you need.
Why CAPTF exists
Cluster API needs one infrastructure provider per cloud, and most teams
already have Terraform or OpenTofu modules that provision that cloud’s
infrastructure. CAPTF turns a module written to its contract into a Cluster
API infrastructure provider directly: you package the module as an OCI
image, and CAPTF runs it, reads its outputs back from state, and reconciles
Cluster, Machine and MachinePool objects against them. There is no Go
controller to write for a new cloud, and no second copy of infrastructure
logic to keep in sync with the module that already exists.
How it fits together
flowchart LR
subgraph capi["CAPI objects"]
Cluster["Cluster, Machine, MachinePool"]
end
subgraph tf["Terraform* objects"]
TFObj["TerraformCluster, TerraformMachine(Pool)"]
end
Manager["manager"]
Job["runner Job"]
Module["module image"]
Cloud["cloud APIs"]
State[("state Secret")]
Cluster --> TFObj
TFObj --> Manager
Manager -- creates --> Job
Job --> Module
Module --> Cloud
Module --> State
State --> Manager
Cluster API’s core objects (Cluster, Machine, MachinePool) reference
CAPTF’s own objects (TerraformCluster, TerraformMachine,
TerraformMachinePool) as their infrastructure. The CAPTF manager
reconciles those objects, renders their inputs, and runs a Kubernetes Job
for each operation. The Job’s runner executes a Terraform or OpenTofu module
image against those inputs, calling the cloud’s own APIs, and its state —
including the outputs the manager reads back — lives in a Kubernetes Secret.
See Architecture for the components behind this
diagram and how one apply flows through them.
Where to go next
| Section | For | Start here |
|---|---|---|
| Getting Started | Anyone trying CAPTF for the first time | Quick Start |
| Concepts | Anyone who needs to understand how CAPTF works | Architecture |
| User Guide | People creating clusters with CAPTF | Identities and Credentials |
| Module Authors | People writing Terraform or OpenTofu modules for CAPTF | Module Contract |
| Operator Guide | People installing and running the CAPTF manager | Installation |
| Reference | Everyone | API Reference |
| Developer Guide | People changing CAPTF itself | Contributing |
Project status
Every kind — TerraformCluster, TerraformClusterTemplate,
TerraformMachine, TerraformMachineTemplate, TerraformMachinePool,
TerraformMachinePoolTemplate and TerraformClusterIdentity — is at API
version v1alpha1. The module contract they implement is also v1alpha1
and provisional: it is frozen for implementation, but may still change
before the first real module has provisioned a cluster with it (see the
contract’s changelog).
CAPTF has no end-to-end tests yet. Its test suite is unit tests that mock the Kubernetes API and Job execution; nothing in it creates a real cluster. See Testing for how the test suite is organized.
Quick Start
This tutorial takes CAPTF from an empty management cluster to a
TerraformCluster and a control-plane TerraformMachine that both report
Ready, using the no-op modules that ship with CAPTF: they create no real
infrastructure, so the tutorial needs no cloud account and no credentials.
This flow has not been run end to end against a live management cluster.
Because the no-op modules provision nothing real, nothing ever boots a
kubelet: no Node ever joins, so KubeadmControlPlane never initializes
and the Machine never reaches its Running phase. What you watch come up
in this tutorial is CAPTF’s own objects finishing their applies, not a
usable Kubernetes cluster. See Your First Module for
writing a module that does create something, and the module
contract for turning one into a real
cloud provider.
Before you begin
- A Kubernetes cluster to use as the management cluster, and
kubectlpointed at it. clusterctl,make, Go andpodmanordocker: CAPTF has no release yet, so this tutorial builds the provider’s and the no-op modules’ images from a clone of this repository instead of fetching them.- A container registry you can push to, and that the management cluster can pull from.
Run every command below from the root of that clone: the make targets
and the templates/... paths are relative to it.
1. Install the provider
CAPTF has not published a release, so clusterctl cannot fetch its
manifest from a URL yet; build one into a local repository instead. Build
and push the manager image, then render the manifest against it:
export IMG=registry.example.com/you/cluster-api-provider-terraform:v0.1.0
make docker-build docker-push IMG="${IMG}"
make manifests-release RELEASE_DIR="${HOME}/local-repository/infrastructure-terraform/v0.1.0" \
RELEASE_IMG="${IMG}" VERSION=v0.1.0
Point a clusterctl config at that directory; the config entry’s name is
terraform, CAPTF’s registered provider name:
# clusterctl.yaml
providers:
- name: terraform
type: InfrastructureProvider
url: file:///home/<you>/local-repository/infrastructure-terraform/v0.1.0/infrastructure-components.yaml
<you> is your home directory’s user name, so the url is the absolute
path of the directory manifests-release just wrote.
clusterctl init --config clusterctl.yaml --infrastructure terraform:v0.1.0
This also installs Cluster API’s core, bootstrap and control-plane
providers, and cert-manager itself if a compatible version is not already
present, since CAPTF’s webhooks need it. See
Installation for what this creates
and how to confirm it, and Installing from a local
repository
for the general form of the local-repository steps above.
2. Apply an identity
Cloud credentials come from a cluster-scoped TerraformClusterIdentity. An
admin applies one per set of credentials, naming the namespaces allowed to
use it:
export TERRAFORM_IDENTITY_NAME=aws-prod NAMESPACE=team-a
clusterctl generate yaml --from templates/identity.yaml | kubectl apply -f -
The no-op modules read no credentials, so the generated Secret’s
placeholder keys can stay as they are; a module that calls a real cloud
provider reads its credentials from the same Secret. kubectl get terraformclusteridentity aws-prod shows Ready=True once the Secret
exists and whoever applied the identity was allowed to get it — the
admission webhook checks. See Identities and
Credentials for creating, rotating and
revoking credentials, and how they reach a Job.
3. Build the no-op module images
Build the cluster- and machine-role no-op images and push them where the management cluster can pull them:
make noop-images NOOP="cluster machine"
export NOOP_REGISTRY=registry.example.com/you
make noop-images-push NOOP="cluster machine" NOOP_REGISTRY="${NOOP_REGISTRY}" VERSION=v0.1.0
This pushes ${NOOP_REGISTRY}/noop-cluster and ${NOOP_REGISTRY}/noop-machine,
each tagged v0.1.0-terraform and v0.1.0-opentofu; either base satisfies
the image contract. This tutorial
uses the terraform tag.
4. Generate and apply a cluster
Cluster objects and everything they own live in a namespace clusterctl
does not create:
kubectl create namespace team-a
Generate the default flavor and apply it. This tutorial asks for one control-plane machine and no workers, since a worker never gets bootstrap data until a real control plane initializes, which the no-op modules never do:
export TERRAFORM_CLUSTER_IMAGE="${NOOP_REGISTRY}/noop-cluster:v0.1.0-terraform"
export TERRAFORM_MACHINE_IMAGE="${NOOP_REGISTRY}/noop-machine:v0.1.0-terraform"
export TERRAFORM_IDENTITY_NAME=aws-prod
clusterctl generate cluster my-cluster --from templates/cluster-template.yaml \
--target-namespace team-a \
--kubernetes-version v1.36.3 \
--control-plane-machine-count 1 --worker-machine-count 0 \
| kubectl apply -f -
--from renders the template file directly, so this command needs no
provider registration and works the same before or after a release exists.
A ClusterClass-based flavor is also available; see Templates and
ClusterClass, which also covers every
variable this template accepts, and clusterctl
variables for their defaults and
built-in safeguards.
5. Watch it come up
kubectl get clusters,machines,machinepools,terraformclusters,terraformmachines,terraformmachinepools -n team-a
machinepools returns nothing: the default flavor creates none. See Add
an autoscaled pool to a generated
cluster
to add one to this cluster.
Each Terraform* kind reports a Ready condition, the only one Cluster
API reads (it is mirrored into the owning Cluster’s or Machine’s
InfrastructureReady). Ready is Unknown while an object waits on
dependencies or on its apply to finish, and True once the module’s apply
succeeds:
kubectl get terraformcluster -n team-a my-cluster \
-o jsonpath='{.status.conditions[?(@.type=="Ready")]}'
Expect TerraformCluster my-cluster and the control-plane
TerraformMachine to both reach Ready=True, and their PHASE columns to
reach Provisioned on Cluster my-cluster and its control-plane Machine
too, since the no-op modules do return a control-plane endpoint and a
provider ID. What you will not see, because nothing real ever boots: the
Machine reaching phase Running (no Node ever registers), and
KubeadmControlPlane reporting itself initialized. That gap is expected
here and is exactly what a module that creates real infrastructure closes.
Watch within about 30 minutes: past that, the control-plane
MachineHealthCheck’s node-startup timeout fires because no Node ever
registers, and KubeadmControlPlane starts remediating the Machine (see
Remediation).
For what each condition type and reason means, see
Conditions; for the reconcile flow behind
these states, see The Reconcile Lifecycle.
6. Clean up
Delete the Cluster first, and wait for it to be gone, rather than
deleting the namespace outright: once a namespace starts terminating, the
API server refuses to create the destroy Jobs each Terraform* object
still needs to run before its own finalizer clears.
kubectl delete cluster my-cluster -n team-a
kubectl wait --for=delete cluster/my-cluster -n team-a --timeout=10m
That wait only returns once every descendant — KubeadmControlPlane, the
MachineDeployment, and both TerraformCluster and the control-plane
TerraformMachine — is gone too, the latter two only after their own
destroy Job finished.
An identity cannot be deleted while a TerraformCluster, TerraformMachine
or TerraformMachinePool still references it, so delete it only once the
wait above returns, then the now-empty namespace:
clusterctl generate yaml --from templates/identity.yaml | kubectl delete -f -
kubectl delete namespace team-a
To remove the provider itself as well:
clusterctl delete --config clusterctl.yaml --infrastructure terraform
See also
- Your First Module to write a module that provisions something real.
- Installation for what
clusterctl initinstalls and how to verify it. - Security Model for the trust boundary a
Terraform*object’s Job operates inside. - The Kinds for how CAPTF’s objects relate to Cluster API’s.
Your First Module
This tutorial writes a small machine-role module from scratch, lints it,
packages it as the OCI image CAPTF runs, lints that image, and references
it from a TerraformMachineTemplate. It ends with a working, lint-clean
module image; it does not provision real infrastructure, since the module
in this tutorial creates no cloud resources. For the full set of rules a
module must follow, see the module contract
and the image contract.
Before you begin
tfcapi-lint, installed as in tfcapi-lint.podmanordocker, to build the image.- A directory to work in. This tutorial calls it
machine/.
Terraform or OpenTofu itself is not required on your machine: the reference
Containerfile below downloads it inside the build, and tfcapi-lint module
parses the module’s files directly.
1. Write the module
A machine-role module implements one Terraform/OpenTofu module that CAPTF
calls once per TerraformMachine. It receives the common contract
inputs plus the machine
role’s inputs, and
must return the machine role’s outputs
plus the common health output. This tutorial’s module returns fixed
values instead of calling a cloud provider, which keeps every input and
output visible and needs no credentials or provider plugins to build or
lint. Swap the fixed values for calls to your cloud’s Terraform/OpenTofu
provider once you understand the shape.
Create four files in machine/.
variables.tf
The contract inputs. captf_contract, captf_cluster, captf_object and
captf_tags apply to every role; captf_cluster_outputs applies to the
machine and machinepool roles only (the cluster role produces it, and so
does not receive it); the rest are specific to the machine role.
# Contract inputs of the machine role, v1alpha1 (https://captf.io/docs/module-author/contract/v1alpha1/common.html
# and machine.html).
variable "captf_contract" {
type = string
}
variable "captf_cluster" {
type = object({
name = string
namespace = string
})
}
variable "captf_object" {
type = object({
kind = string
name = string
namespace = string
})
}
# The cluster module's exports. The controller always sets it; the default
# follows the contract skeleton (machine.md).
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" {
type = map(string)
}
variable "machine_name" {
type = string
}
# Base64 of the bootstrap Secret's value.
variable "bootstrap_data" {
type = string
sensitive = true
}
variable "bootstrap_format" {
type = string
}
variable "failure_domain" {
type = string
default = null
}
variable "kubernetes_version" {
type = string
default = null
}
variable "control_plane" {
type = bool
}
main.tf
The module’s own logic. This tutorial stands in a terraform_data
resource for a real instance, so the module has something to hold its
inputs; a real module replaces this with the resources that create an
instance.
# No-op machine module: implements the v1alpha1 machine role with no cloud.
# The stand-in for an instance. user_data decodes bootstrap_data the way a
# module feeding a plain user-data argument would, which proves the
# controller's base64 encoding round-trips (the value stays sensitive).
resource "terraform_data" "instance" {
input = {
cluster = var.captf_cluster
object = var.captf_object
machine_name = var.machine_name
tags = var.captf_tags
backend_id = try(var.captf_cluster_outputs.backend_id, null)
user_data = base64decode(var.bootstrap_data)
bootstrap_format = var.bootstrap_format
failure_domain = var.failure_domain
kubernetes_version = var.kubernetes_version
control_plane = var.control_plane
}
}
outputs.tf
The contract outputs. provider_id, addresses and failure_domain are
required; interruptible must be declared even when it is always false;
health is the common output every role returns.
# Contract outputs of the machine role, v1alpha1.
# Stable per Machine. With no Node behind it, the e2e suite never expects a
# nodeRef; a real module emits the CCM's or kubelet's format (machine.md).
output "provider_id" {
value = "noop:///${var.captf_object.namespace}/${var.machine_name}"
}
output "addresses" {
value = [{ type = "InternalIP", address = "10.0.0.1" }]
}
output "failure_domain" {
value = var.failure_domain
}
output "interruptible" {
value = false
}
output "health" {
value = { state = "running", healthy = true, message = null, reasons = [] }
}
versions.tf
terraform {
# terraform_data needs Terraform 1.4; the variants' plantimestamp() needs
# 1.5. Every OpenTofu release (1.6+) has both.
required_version = ">= 1.5"
}
This module declares no required_providers: terraform_data ships with
Terraform/OpenTofu itself, so there is no provider to install or mirror.
A module that calls a cloud provider adds a required_providers block
here as usual.
2. Lint the module
tfcapi-lint module ./machine --role machine --strict
A clean module prints an empty finding list and an all-zero summary.
--strict fails the command on a warning as well
as an error, so a lint-clean module here stays lint-clean once you add
real inputs, outputs and user variables. See
tfcapi-lint for the full command
reference and tfcapi-lint CLI for every
check ID.
3. Package it as an image
CAPTF runs a module as one OCI image that bundles the module’s files and
the Terraform or OpenTofu binary; there is no separate module source and
no separate runtime image. Save the reference Terraform-based Containerfile
into machine/:
# Reference source image, Terraform base (docs/book/src/module-author/image-contract.md "Reference:
# Terraform base"). Run from your module's root directory:
#
# podman build -f Containerfile.terraform --build-arg ROLE=cluster \
# --build-arg IMAGE_SOURCE=https://github.com/<org>/<repo> \
# --build-arg IMAGE_REVISION="$(git rev-parse HEAD)" \
# --build-arg IMAGE_VERSION=<tag> -t <registry>/<repo>:<tag> .
#
# The image tag is the module version. Lint first:
# tfcapi-lint module . --role cluster --strict
ARG RUNTIME_VERSION=1.16.4
# The org.opencontainers.image.* labels below (image-contract.md "OCI
# labels"): leave these unset for a local/test build, or pass them from
# your CI pipeline (source repo URL, commit SHA, the image tag).
ARG IMAGE_SOURCE=""
ARG IMAGE_REVISION=""
ARG IMAGE_VERSION=""
# Optional but recommended: hermetic provider mirror for the platforms you
# publish. Needs registry egress at build time; drop this stage (and the
# COPY --from=mirror below) for a non-hermetic image.
FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION} AS mirror
WORKDIR /src
COPY . /src
# get: `providers mirror` refuses a module whose nested local modules are not
# installed; get installs them (no providers), in this stage only.
# mkdir: `providers mirror` does not create the target when the module
# requires no providers, and the COPY --from=mirror below needs it.
RUN terraform get \
&& mkdir -p /captf/providers \
&& terraform providers mirror -platform=linux_amd64 -platform=linux_arm64 /captf/providers
FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION}
ARG ROLE=cluster
ARG RUNTIME_VERSION
ARG IMAGE_SOURCE
ARG IMAGE_REVISION
ARG IMAGE_VERSION
COPY --from=mirror /captf/providers /captf/providers
COPY . /captf/module
RUN ln -s /bin/terraform /captf/runtime \
&& adduser -D -u 65532 captf \
&& chown -R 65532:65532 /captf
USER 65532
LABEL io.captf.contract="v1alpha1" \
io.captf.role="${ROLE}" \
io.captf.runtime="terraform" \
io.captf.runtime.version="${RUNTIME_VERSION}" \
org.opencontainers.image.source="${IMAGE_SOURCE}" \
org.opencontainers.image.revision="${IMAGE_REVISION}" \
org.opencontainers.image.version="${IMAGE_VERSION}"
An OpenTofu-based equivalent is also available
(Containerfile.opentofu);
either runtime satisfies the contract. Build from inside machine/:
export IMAGE=registry.example.com/acme/machine:v1.0.0
podman build -f Containerfile.terraform --build-arg ROLE=machine -t "$IMAGE" .
IMAGE is the registry, repository and tag you push to; ROLE=machine
tags the image with the role it implements. See the image
contract for the fixed paths, OCI
labels and execution environment every image must satisfy.
4. Lint the image
tfcapi-lint image reads a pushed registry reference or a local OCI
layout. Without a registry to push to yet, save the image podman just
built to a local OCI directory and lint that:
podman save --format oci-dir -o /tmp/first-module-image "$IMAGE"
tfcapi-lint image --role machine "oci:/tmp/first-module-image"
Once you push $IMAGE to a registry, lint the pushed reference the same
way: tfcapi-lint image --role machine "$IMAGE".
5. Reference it from a TerraformMachineTemplate
A TerraformMachineTemplate is what a MachineDeployment, a
KubeadmControlPlane or a MachineSet points at to create
TerraformMachines; its spec.template.spec.source.image names the image
you just built:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachineTemplate
metadata:
name: first-module
namespace: default
spec:
template:
spec:
source:
image: registry.example.com/acme/machine:v1.0.0
Apply it:
kubectl apply -f first-module-template.yaml
kubectl get terraformmachinetemplate first-module -n default confirms
the object exists; there is nothing to provision yet, since nothing
references this template as a MachineDeployment, MachineSet or control
plane’s infrastructureRef does. A TerraformMachine created from it
needs an identity to run under: either its own spec.identityRef, or one
inherited from its TerraformCluster’s spec.defaults.identityRef, as
set up in the quick start.
Next steps
- Walk through the quick start to wire a machine template like this one into a full cluster.
- Read the module contract for every rule a production module must follow, and the cluster role and machine pool role for the other two roles a full deployment needs.
- Read Runtime Environment for what a module sees when the runner actually executes it.
Architecture
This page describes CAPTF’s components, what runs where, and how one apply flows through them. It is for anyone who needs to understand how CAPTF works before reading a more specific page.
Components
The manager Deployment is CAPTF’s one Deployment and one container
image. It runs the controllers that reconcile every kind except
TerraformClusterTemplate and TerraformMachinePoolTemplate, which have
no reconciler of their own, and it serves the validating webhooks for all
seven kinds from the same process, on a separate port. It watches every
namespace by default: clusterctl init never sets a namespace
restriction. See Configuration for
how to scope or tune it.
The validating webhooks reject an invalid or disallowed change to a
Terraform* object before it is persisted: an immutable field, a malformed
image reference, a spec.variables name that collides with a contract
input, or (for TerraformClusterIdentity) a Secret the requester cannot
read. There are no mutating webhooks — CAPTF resolves every default at
reconcile time instead of writing it back onto the object. See
Security Model for the trust boundary a webhook
decision sits inside, and RBAC for who can
change what.
A runner Job is a batch/v1 Job the manager creates for one operation
(apply, destroy, drift, refresh, restore or plan) on one Terraform*
object, in that object’s own namespace. Kubernetes never retries a Job pod;
the manager owns retries itself, one attempt per Job name. See
The Reconcile Lifecycle for how the manager decides which
operation runs next.
The module image is the OCI image a Terraform* object’s
spec.source.image names. It bundles a Terraform or OpenTofu module
written to CAPTF’s contract for one role (cluster, machine or
machinepool) together with the runtime binary that runs it, at fixed
paths, and optionally a provider filesystem mirror. It never bundles the
runner. See Image Contract.
The runner binary is copied into each Job’s pod by an init container,
rather than shipped in the module image. The manager’s own image bundles
both /manager and a static /runner binary; the init container runs
/runner copy /captf/bin/runner into an emptyDir the module container
also mounts, so by default the runner every Job executes is exactly the
build of the manager that created the Job; an operator can point it at a
different image instead (see
Configuration). The
module container then runs that copy as its command, against the module
and runtime the image provides. The init container’s resources are fixed
and not configurable, since it only copies one static binary and never
varies with the module.
The Kubernetes state backend is the Terraform/OpenTofu kubernetes
backend, which the manager configures for every generated root module.
State lives in Secrets in the Terraform* object’s own namespace, chunked
when it is large, and labeled by the object’s immutable identity —
namespace, kind and name, never its UID — so state survives a clusterctl move and a changed identity is never mistaken for existing state. See
Terraform State.
What runs where
| Component | Runs as | Namespace |
|---|---|---|
| manager, webhooks | One Deployment, one container | The provider’s own namespace (captf-system by default) |
| runner Job | One batch/v1 Job per operation | The Terraform* object’s own namespace |
| runner binary, module | Init and main containers of the Job’s pod | Same pod as the Job |
| state | Secrets | The Terraform* object’s own namespace |
The data flow of one apply
The manager reconciles a TerraformCluster, TerraformMachine or
TerraformMachinePool through a shared core (externally-managed check,
owner lookup, finalizer, pause) and decides that an apply is the next
operation. It renders that object’s inputs into a generated root module —
main.tf.json, which configures the kubernetes backend and calls the
module image’s role module with exactly the contract inputs, and
terraform.tfvars.json, the input values — and writes them to a durable
inputs Secret and a per-run Secret, hashing the inputs as it does. See
Job Inputs for exactly where each value comes from.
The manager then builds and creates the Job: the init container that copies
the runner binary, and the module container that mounts the per-run Secret
read-only, the identity’s mirrored credentials Secret, and empty scratch
volumes. Once the pod starts, the runner initializes the backend, plans and
applies against the module’s runtime, and — for a TerraformCluster apply —
stops before a plan that deletes or replaces a resource unless that plan
was already approved (see Plan Approval).
The generated root re-exports the module’s contract outputs, so they land
in state alongside everything else the runtime tracks.
The runner reports how the Job finished as its pod’s termination message. The manager reads that message, and — once the Job has finished — the state Secrets themselves, to set the object’s status and conditions and, after a successful apply, to adopt (re-label) the new state Secrets as its own.
What a module author provides versus what an operator provides
A module author writes and builds the OCI image: the role module (and any local modules it calls), pinned to the module contract and the image contract, optionally with a provider mirror baked in. They check it with tfcapi-lint before publishing it.
An operator installs the manager
(Installation), creates the
TerraformClusterIdentity credentials CAPTF’s objects reference
(Identities and Credentials), and creates the
Terraform* objects — directly or through templates and
ClusterClass — that name a module author’s
image. They tune Job resources, deadlines and security contexts per object
(Tuning Jobs), and watch the conditions,
events and alerts the manager and runner produce
(Observability).
The Kinds
If you are learning how CAPTF’s objects relate to Cluster API’s, start
here. CAPTF adds seven kinds to one API group and version,
infrastructure.cluster.x-k8s.io/v1alpha1. Three of them are the objects
that carry the desired state Cluster API’s core controllers drive as a
cluster comes up; three are templates that stamp those out; the seventh
holds the cloud credentials the others use. Every field, default and
validation rule is in the API reference; this page
explains how the seven fit together, and what stays fixed once an object
exists.
The seven kinds
| Kind | Cluster API role | Scope | Referenced by |
|---|---|---|---|
TerraformCluster | InfraCluster | Namespaced | Cluster.spec.infrastructureRef, always directly |
TerraformMachine | InfraMachine | Namespaced | Machine.spec.infrastructureRef, always directly |
TerraformMachinePool | InfraMachinePool | Namespaced | MachinePool.spec.template.spec.infrastructureRef, always directly |
TerraformClusterTemplate | Template | Namespaced | A ClusterClass |
TerraformMachineTemplate | Template | Namespaced | A MachineDeployment, MachineSet, KubeadmControlPlane, another control-plane provider, or a ClusterClass |
TerraformMachinePoolTemplate | Template | Namespaced | A ClusterClass |
TerraformClusterIdentity | None (CAPTF-only) | Cluster | spec.identityRef of a cluster, machine or pool in an allowed namespace |
The three workload kinds
A TerraformCluster, TerraformMachine or TerraformMachinePool runs one
Terraform or OpenTofu module role — cluster, machine or machinepool — as a
Job. What goes into that Job and how the reconcile loop drives it are
covered in Job inputs and
the reconcile lifecycle; this section is about the object
itself.
Each of the three is owned by its Cluster API counterpart, not by CAPTF:
the counterpart’s own core controller sets the owner reference. A
TerraformCluster or TerraformMachinePool can be an object you create
and point the Cluster or MachinePool at directly, or one a
ClusterClass topology expands from a TerraformClusterTemplate or
TerraformMachinePoolTemplate; either way, the resulting object is owned
the same way. A TerraformMachine is normally expanded from a
TerraformMachineTemplate, one per Machine, by a MachineDeployment,
a MachineSet or a control-plane provider, whether or not a ClusterClass
is involved; a Machine can also point at a directly created
TerraformMachine, the same way a Cluster or MachinePool can point
at a directly created TerraformCluster or TerraformMachinePool.
TerraformClusteris owned by aCluster.spec.controlPlaneEndpointis mutable until it has a host — set by you, or once by the controller from the module’s output — and immutable after: every Machine and kubeconfig of the cluster points at it. Everything else, including the module image, can change on a live object; a new image or a changed input re-applies the module against the existing state.spec.identityRefis required.spec.applyPolicyand the destructive-plan guard are covered in plan preview and approval.TerraformMachineis owned by aMachine.spec.source,spec.identityRef,spec.variablesandspec.variablesFromdefine the machine and are immutable after creation: change them through a MachineDeployment or control-plane rollout, not in place.spec.providerIDcan only move from empty to non-empty, by the controller.spec.jobs,spec.driftandspec.remediationare operational policy and stay mutable at any time, so a stuck machine’s deadline or drift interval can be changed without rolling it. A direct delete is refused outside a few exceptions; see the reconcile lifecycle for deletion and finalizers.TerraformMachinePoolis owned by aMachinePool. Unlike aTerraformMachine, every field is mutable: a change tospec.source,spec.variables,spec.variablesFromor any other field re-applies the module on the next reconcile. A pool has nospec.applyPolicyand no delete guard. With autoscaling enabled, it also writes the group’s observed replica count back toMachinePool.spec.replicas; see machine pools.
Templates
TerraformClusterTemplate, TerraformMachineTemplate and
TerraformMachinePoolTemplate each hold spec.template, the metadata and
spec a TerraformCluster, TerraformMachine or TerraformMachinePool is
created with. A TerraformClusterTemplate and a TerraformMachinePoolTemplate
are each expanded by a ClusterClass topology into one TerraformCluster or
TerraformMachinePool, which the Cluster or MachinePool then
references directly. A TerraformMachineTemplate is expanded, with or
without a ClusterClass, by a MachineDeployment, a MachineSet, a
KubeadmControlPlane or another control-plane provider, into one
TerraformMachine per Machine. See
ClusterClass for how templates and
flavors fit together.
spec.template.spec is immutable once a template exists, like every
Cluster API template: create a new template and point the owner at it
instead of editing one in place. spec.template.metadata stays mutable.
The one exception is a ClusterClass topology dry-run, which the webhook
recognizes and exempts from the immutability check, so a topology patch
that only touches spec.template.spec can still be validated without
tripping it.
A TerraformMachineTemplate additionally reports status.capacity and
status.nodeInfo, resolved from its image’s labels, for Cluster
Autoscaler scale-from-zero. TerraformMachinePoolTemplate has no status:
a pool has no scale-from-zero, so there is no capacity to resolve.
TerraformClusterIdentity
TerraformClusterIdentity is cluster-scoped and has no Cluster API
counterpart: it exists only to hold spec.secretRef, the credentials
Secret a TerraformCluster’s, TerraformMachine’s or
TerraformMachinePool’s module run is allowed to use, and
spec.allowedNamespaces, which namespaces may reference it. Every field
is mutable, including secretRef; changing it re-checks that the
requester may read the new Secret. Deleting an identity is refused while
any object still uses it or still has its credentials mirrored into a
namespace. See identities for creating,
rotating and revoking credentials, and how they reach a Job.
What a cluster passes to its machines and pools
A TerraformCluster’s spec.defaults apply only to its
TerraformMachines and TerraformMachinePools, never to the cluster
itself, and are merged field by field: a field the machine or pool sets
wins, an unset one comes from spec.defaults, and a field neither sets
gets the built-in default. A machine or pool finds its Cluster by its
own cluster.x-k8s.io/cluster-name label, then the TerraformCluster
through the Cluster’s spec.infrastructureRef.
identityRef: a machine or pool uses its ownidentityRefwhen it sets one, else the cluster’sspec.defaults.identityRef, else the cluster’s ownspec.identityRef.jobs: merged field by field withspec.defaults.jobs, own over defaults; see Tuning Jobs for the full merge algorithm.drift: onlyintervalSecondsis inherited — the object’s own, elsespec.defaults.drift.intervalSeconds, else the manager’s built-in default. A pool’s owndrift.actionis never inherited; unset, it defaults toReportregardless of the cluster’s defaults, which carry noactionto inherit.
The pool exception: a TerraformMachine’s drift can be turned off
with intervalSeconds: 0, and an inherited spec.defaults.drift of 0
turns it off the same way, but a TerraformMachinePool’s drift can never
be turned off; see
Drift and health for why.
See also
- API Reference for every field, default and validation rule.
- The reconcile lifecycle for what runs when, retries, deletion and finalizers.
- Job inputs for what each role’s Job receives.
- ClusterClass for templates in use.
- Machine pools for autoscaling and replica write-back.
- Identities for
TerraformClusterIdentityin depth.
The Reconcile Lifecycle
TerraformCluster, TerraformMachine and TerraformMachinePool share one
reconcile flow. This page walks through a single pass of it: what runs
before any decision, how the next operation is chosen, how a failed
operation backs off, how two objects of the same cluster stay off each
other’s feet, and what each kind does once the shared flow is done. It is
for anyone who wants to know why CAPTF started (or did not start) a Job, or
why a delete is taking a while.
One pass, one status patch
Every reconcile writes the object’s status once, in a single patch at the end of the pass. Two things are written earlier, directly, because a crash between them and the deferred patch must not leave a gap:
- the clusterctl move block, set just before a Job is created;
- the preamble’s own writes, described below, which can stop the pass before there is anything else to decide.
A reconcile that reaches the end of the pass runs through, in order:
- Externally managed. An object carrying CAPI’s externally-managed annotation is left alone entirely; nothing below runs.
- Owner lookup. Without an owner reference of the expected kind yet,
the object waits (
DependenciesReady=Unknown) unless it is being deleted with no state and no running Job, in which case the finalizer is dropped at once: there is nothing to destroy. An owner reference whose target is gone does not stop deletion; a destroy still runs from the durable inputs. An owner gate such as the owningCluster’s infrastructure reference not naming aTerraformClustersetsDependenciesReadyand stops the pass, unless the object is being deleted, in which case deletion proceeds anyway. The lookup also requires the owner to reference this object back (its owninfrastructureRef, and, when set, a matching ownerRef UID and cluster-name label); an owner that fails this check is treated the same as one that is gone:DependenciesReady=False/OwnerMismatch, no Job, and deletion still proceeds. See Owner references are checked against the owner for the full rule and why it exists. - Finalizer. Added if missing; the pass stops for this reconcile so the write is visible before anything else happens.
- Pause. A paused object, or one whose
Clusteris paused, runs only the paused branch described below.
See Conditions for every condition this flow sets and what each reason means, and TerraformClusterIdentity for the identity and credential-mirror step that follows the preamble.
The paused branch
A paused object still processes its finished Jobs (bookkeeping, below) and
deletes a Job that can never start, but starts nothing new. Once no Job is
active, the clusterctl move block is cleared. This is what clusterctl move waits for after it pauses an object; see
the move runbook.
Bookkeeping
Unpaused, the reconcile reads the object’s Jobs and processes every one
that finished since the last pass it was seen: it records the run’s steps,
error and drift summary in status.lastRun, sets the ApplyJobSucceeded
and DriftJobSucceeded conditions, pins the image digest of a newly
succeeded apply, prunes history beyond the configured limits, and finds
whether a Job is active now. A stale state lock left by a Job whose pod is
gone is flagged for the next Job to force-unlock; a lock held by something
else (a workstation, another tool) is reported instead. A finished Job’s
per-run inputs Secret is deleted once, which is also the point at which its
run lease (and, for a TerraformCluster, its cluster write lease) is
given back.
What a finished Job’s pod reported is only durable once the status patch at the end of the pass has succeeded: the reconcile marks a Job “bookkept” only after that patch, so a failed patch leaves the Job to be read again next time instead of losing its result.
While a Job is active, state is not read and no new operation is decided: the reconcile deletes the Job if it looks stuck (its per-run Secret is missing and no pod ever started, checked only once the Job is at least a minute old) and otherwise waits, holding the clusterctl move block.
Choosing the next operation
Once no Job is active, the reconcile reads state, builds the current inputs (for a mutable kind, or an immutable one not yet provisioned), and looks for a requested state restore. It then picks the next operation in this order:
- A restore named by the
captf.io/restore-stateannotation, unless the object is being deleted: it is what was asked for, and checking or applying against the state about to be replaced would be wasted or worse. - Being deleted:
destroy, or drop the finalizer at once if there is no state to destroy. - No state yet:
apply(first apply). - State exists but carries no inputs hash:
apply. - A mutable kind whose current inputs hash differs from the one state
carries:
apply. - A mutable kind whose newest finished apply failed (and was not blocked,
was not stopped by a plan mismatch, and was not a drift remediation):
applyagain, even if the current inputs already match state, because the failed run may have changed reality without recording it. - A mutable kind with drift found and drift action
Remediate, while the object has not yet reached its remediation-failure cap since the last successful drift check:apply. Drift detection and remediation are Drift and health’s topic.
Whichever of those an apply reason names but the first (no state yet, with
nothing to break), an object under applyPolicy Manual does not apply
directly: it waits for a plan and its approval first, or backs off behind
the destructive-plan guard if neither policy applies. Both are
Plan approval’s topic; this page only
fixes where that wait sits in the order above.
When none of those apply reasons matched, the reconcile falls through to its refresh, membership and drift schedule, checked in this order and stopping at the first one due:
- A successful apply not yet followed by a
refresh, for kinds that ask for one (a machine or machine pool, so its addresses and health reach status without waiting for the next tick). - A
TerraformMachinePoolwhose membership is still converging:refreshevery 30 seconds, plus jitter, fixed rather than backed off. - Otherwise, while health reads pending:
refreshafter a delay that starts at 30 seconds and doubles (30s, 1m, 2m, 4m) up to a 5-minute ceiling, plus jitter. - A machine pool’s membership refresh interval elapsed:
refresh. - The health-check interval elapsed:
refresh. - The drift-check interval elapsed:
drift. - Otherwise, requeue at whichever of the above is soonest.
Every interval above adds a small deterministic jitter derived from the
object’s UID (up to a tenth of the interval), so objects created together,
such as a MachineDeployment’s machines or everything moved by one
clusterctl move, do not all check in lockstep. Configuring these
intervals is Drift and health’s and
Configuring drift’s topic; machine pool
membership convergence is Machine pools’s.
Retry backoff
A failed Job’s operation is not retried immediately: the delay after n
consecutive failures of that operation is one minute, doubling each time,
capped at ten minutes. Once n reaches the failed-Jobs history limit
(three by default), the delay jumps straight to the ten-minute cap instead
of continuing to double, so the sequence with the default limit is one
minute, two minutes, then ten minutes and no shorter. A Job stopped from
outside (its pod was interrupted, or an apply was blocked before a
destructive plan, or an approved apply’s plan changed since approval, or
the run never started because a lease was held) does not count toward this
backoff at all: those wait on a lease, an approval, or the object’s own
next reconcile instead. A failed state restore is a special case: it is
never retried for the same backup serial, only for a new one or after the
failed Job is deleted.
Run leases and the cluster operation gate
Before a Job is created, the reconcile takes the object’s own run lease, so
two managers (or two reconciles racing after a crash) never start two Jobs
for the same object at once. If another Job already holds it, the
operation waits instead: ApplyJobSucceeded, DriftJobSucceeded or
RestoreJobSucceeded (matching the operation) goes Unknown with a
message naming the Job to wait for.
With the manager’s --cluster-operation-gate flag (default enabled; see
Manager configuration), an apply,
destroy or restore Job of an object that names a cluster additionally
crosses a second gate meant to keep a cluster’s own apply or destroy from
running at the same time as its machines’ and machine pools’:
- A
TerraformClustertakes the cluster’s write lease, then checks whether any machine or machine pool of the cluster is currently applying or destroying. If so, it keeps both leases and waits for them to finish before it starts: new machine and machine pool operations wait behind it meanwhile, so a stream of machine creates cannot starve the cluster’s own operation. - A
TerraformMachineorTerraformMachinePoolchecks whether the cluster’s write lease is currently held. If it is, the machine or pool gives back the run lease it just took and waits for the cluster’s operation to finish first.
Deletion order
A TerraformCluster being deleted waits for every TerraformMachine and
TerraformMachinePool carrying its cluster name in the namespace to be
gone before its own destroy runs; machines and machine pools carry no such
wait of their own. An owner reference whose target has already been
removed does not block a destroy: deletion still runs from the durable
inputs. Once a destroy succeeds, or deletion finds no state to destroy at
all, the reconcile deletes the object’s state and durable inputs, releases
its leases, drops it from its credential mirror’s owners, and removes the
finalizer. See Secrets for what those
Secrets are and RBAC for the namespace RBAC
sweep that follows a finalizer’s removal.
The clusterctl move block
Clusterctl’s block-move annotation is set on the object just before a Job
is created and stays set for as long as a Job is active; it is cleared,
paused or not, once no Job is active. clusterctl move pauses every object
first and then waits for the annotation to clear, which is why the paused
branch above still runs bookkeeping and the stuck-Job check: a Job that can
never start would otherwise hold the block indefinitely.
Job names and history
A Job’s name is deterministic:
captf-<kindshort>-<name>-<operation>-a<attempt>-<hash>, where
<kindshort> is c, m or mp, <attempt> is one more than the
highest attempt number retained for that operation, and <hash> is six
hex characters derived from the inputs hash, the operation, the attempt
and, for refresh and drift, a tick that changes each time the operation
last succeeded (so a retry finds the same Job instead of starting a second
one); a drift-remediation apply and a restore carry a tick of their own
too, so neither collides with an unrelated Job of the same inputs hash and
attempt. A name that would exceed 57 characters replaces the object name
with a hash instead. History beyond the configured successful- and
failed-Jobs limits is pruned per operation, oldest first, while always
keeping each operation’s newest success and, while no success is newer,
its newest failure, since those are what retry backoff and drift
remediation’s failure cap read. Tuning those limits, and the rest of a
Job’s resources, deadlines and identity, is
Job tuning’s topic.
Jobs per bring-up
Bringing up one TerraformCluster with three control-plane and three
worker TerraformMachines runs eight Jobs when each machine’s apply
outputs give a definite health reading: one cluster apply, one apply per
machine (six), and one no-change re-apply once the cluster’s
control_plane_initialized input flips (the cluster’s inputs hash
changes, so the shared flow applies again even though nothing else about
the cluster changed). A machine whose apply outputs read pending instead
keeps the extra post-apply refresh described above; with all six machines
reading pending that adds six more Jobs, back up to fourteen. Count Jobs
over a bring-up with sum(captf_jobs_total) rather than listing Jobs: a
finished Job is pruned once it falls outside the history limits above.
After the shared flow
Each kind does a little more once the shared flow has read state and decided:
- A
TerraformClusterdecides itscontrol_plane_endpointinput while the shared flow builds inputs, and writes an apply’s endpoint output back tospec.controlPlaneEndpointwhile state is read: both are part of the shared flow, not a separate step. Each happens at most once. - A
TerraformMachinekeepscluster.x-k8s.io/remediate-machineon its owningMachinein step with its instance health once the shared flow finishes, skipped for a paused or externally-managed machine. See Machine health and remediation. - A
TerraformMachinePoolwrites its observed replica count back to its owningMachinePool’sspec.replicaswhen autoscaling is enabled, skipped for a paused, externally-managed or deleting pool. See Machine pools.
One pass, visually
flowchart TD
A["Preamble: externally managed?<br/>owner lookup, finalizer, pause"]
A -->|paused| P["Paused branch:<br/>bookkeeping, delete a stuck Job,<br/>clear the move block"]
A -->|not paused| B["Credentials: identity,<br/>credential mirror, runner RBAC"]
B --> C["Bookkeeping: finished Jobs,<br/>digest pin, lock check"]
C -->|a Job is active| D["Delete it if stuck,<br/>else wait for it"]
C -->|no Job is active| E["Clear the move block"]
E -->|deleting, destroy succeeded| F["Cleanup: remove the finalizer"]
E -->|deleting, machines or pools remain| G["Wait: deletion blocked"]
E -->|otherwise| H["Read state, build inputs,<br/>look for a requested restore"]
H --> K["Decide the next operation"]
K -->|start a Job| L["Take leases, start the Job"]
K -->|drop the finalizer| F
K -->|nothing to start| M["Requeue at the decided delay"]
D --> N["Patch status once"]
L --> N
M --> N
P --> N
F --> N
G --> N
See also
- Conditions — every condition this flow sets.
- Events — the events it emits.
- Drift and health — how drift checks and health work.
- The state backend — reading and adopting state, and backups.
- Plan approval — the Manual apply policy and the destructive-plan guard.
Job Inputs
This explains what goes into a CAPTF Job’s Terraform inputs —
terraform.tfvars.json and the generated root main.tf.json — and where
each value comes from on the Kubernetes side. For each input’s type and
whether it is required, see the module
contract; this page traces
the plumbing that fills the contract in, for anyone administering CAPTF or
writing a module against it.
A TerraformCluster runs the module’s cluster role, each TerraformMachine
runs its machine role, and each TerraformMachinePool runs its
machinepool role. All three roles are rendered the same way from the same
kind of input structure, so this page covers all three side by side.
1. The pipeline
flowchart TD
subgraph capi[CAPI objects]
Cluster["Cluster<br/>spec.topology.version<br/>spec.clusterNetwork<br/>spec.controlPlaneEndpoint<br/>status.initialization.controlPlaneInitialized"]
Machine["Machine<br/>spec.version<br/>spec.failureDomain<br/>spec.bootstrap.dataSecretName<br/>labels[cluster.x-k8s.io/control-plane]"]
MachinePool["MachinePool<br/>spec.replicas<br/>spec.failureDomains<br/>spec.template.spec.version<br/>spec.template.metadata.labels<br/>autoscaler min/max-size annotations"]
Bootstrap["Bootstrap Secret (KubeadmConfig/RKE2Config)<br/>keys: value, format"]
end
subgraph objs[Terraform* objects]
TC["TerraformCluster<br/>spec + status"]
TM["TerraformMachine<br/>spec + status"]
TMP["TerraformMachinePool<br/>spec + status"]
end
Exports["TerraformCluster state<br/>outputs: exports, failure_domains"]
Cluster --> ClusterInputs
TC --> ClusterInputs
ClusterInputs[Cluster inputs] --> RenderC[Render, cluster role]
RenderC --> DurableC["Durable Secret<br/>captf-inputs-c-name"]
RenderC --> RunC["Per-run Secret<br/>captf-run-job<br/>mounted at /captf/config"]
RunC --> RunnerC[Runner] --> ModuleC[Module, cluster role]
ModuleC --> Exports
Machine --> MachineInputs
Bootstrap --> MachineInputs
Cluster --> MachineInputs
TM --> MachineInputs
Exports -->|only while not yet provisioned| MachineInputs
MachineInputs[Machine inputs] --> RenderM[Render, machine role]
RenderM --> DurableM["Durable Secret<br/>captf-inputs-m-name"]
RenderM --> RunM["Per-run Secret<br/>captf-run-job<br/>mounted at /captf/config"]
RunM --> RunnerM[Runner] --> ModuleM[Module, machine role]
MachinePool --> PoolInputs
Bootstrap --> PoolInputs
Cluster --> PoolInputs
TMP --> PoolInputs
Exports -->|every reconcile| PoolInputs
PoolInputs[Machine pool inputs] --> RenderP[Render, machinepool role]
RenderP --> DurableP["Durable Secret<br/>captf-inputs-mp-name"]
RenderP --> RunP["Per-run Secret<br/>captf-run-job<br/>mounted at /captf/config"]
RunP --> RunnerP[Runner] --> ModuleP[Module, machinepool role]
DurableC -.->|destroy always; drift/refresh when gated or current inputs unavailable| RenderC
DurableM -.->|destroy and drift/refresh always, machine is immutable| RenderM
DurableP -.->|destroy always; drift/refresh when gated or current inputs unavailable| RenderP
Why two Secrets exist:
- Durable inputs Secret (
captf-inputs-<kindshort>-<name>, owned by the Terraform* object, moves with it acrossclusterctl move): the record of what was rendered when the controller last started an apply Job for this object (written at Job start, not gated on the Job’s success). Destroy always prefers it; a cluster or pool (both mutable) whose durable Secret is missing falls back to freshly built current inputs, but a machine (immutable) never does — with no durable Secret and no live spec to fall back to, destroy has nothing to run against until the Secret is restored. Drift and refresh follow the same asymmetry: a cluster or pool checks against its current inputs when they build cleanly, falling back to the durable Secret only when they don’t (a dependency gate, an owner gone); a machine, being immutable, always checks against the durable Secret. - Per-run Secret (
captf-run-<job>, owned by the Job, mounted at/captf/config, deleted when the Job finishes): a private copy the Job’s pod reads, so a rewrite of the durable Secret by a later reconcile can never change the files under a running Job.
2. Cluster role inputs
For the full type table, see
contract/v1alpha1/common.md
and
cluster.md. The
Kubernetes-side sources are:
| tfvars key | Type | Source |
|---|---|---|
captf_contract | string | The fixed contract version, "v1alpha1" |
captf_cluster | object{name,namespace} | The owning Cluster’s metadata.name/metadata.namespace |
captf_object | object{kind,name,namespace} | The TerraformCluster’s own kind/name/namespace |
captf_tags | map(string) | Fixed keys built from the Cluster name, the TerraformCluster namespace/kind/name, and its cluster.x-k8s.io/cloned-from-name annotation (empty string if absent) |
control_plane_endpoint | object{host,port} or null | Cluster.spec.controlPlaneEndpoint when valid; else TerraformCluster.spec.controlPlaneEndpoint when valid; else null. Always null once the captf.io/endpoint-source annotation is module |
kubernetes_version | string or null | Cluster.spec.topology.version; null without a ClusterClass |
control_plane_initialized | bool | Cluster.status.initialization.controlPlaneInitialized, latched: true once the durable Secret’s last rendered value was true, so it is never rendered false again after an apply saw true |
cluster_network | object{pods,services,service_domain,api_server_port} or null | Cluster.spec.clusterNetwork; null only when every attribute is unset. CIDR lists render as [], not null, when the network object is present but a list is empty |
A trimmed real example:
{
"captf_contract": "v1alpha1",
"captf_cluster": { "name": "prod", "namespace": "team-a" },
"captf_object": { "kind": "TerraformCluster", "name": "prod", "namespace": "team-a" },
"captf_tags": {
"captf.io/cluster": "prod",
"captf.io/kind": "TerraformCluster",
"captf.io/managed-by": "captf",
"captf.io/name": "prod",
"captf.io/namespace": "team-a",
"captf.io/template": ""
},
"control_plane_endpoint": { "host": "api.prod.example.com", "port": 6443 },
"kubernetes_version": "v1.36.2",
"control_plane_initialized": true,
"cluster_network": {
"pods": ["192.168.0.0/16"],
"services": ["10.96.0.0/12"],
"service_domain": "cluster.local",
"api_server_port": 6443
}
}
The matching root main.tf.json declares one Terraform variable per key
above (nullable ones get "default": null), a module "role" call
forwarding every variable by name to the module, and a sensitive = true
re-export of every required cluster output (control_plane_endpoint,
failure_domains, exports, health). The cluster role never declares
captf_cluster_outputs: it has no such input at all, not even null.
3. Machine role inputs
For the full type table, see
common.md and
machine.md. The
Kubernetes-side sources are:
| tfvars key | Type | Source |
|---|---|---|
captf_contract, captf_cluster, captf_object, captf_tags | as above | Same construction as the cluster role, but captf_object/captf_tags name the TerraformMachine |
captf_cluster_outputs | any | The owning TerraformCluster’s exports output, verbatim (section 5). {} for an externally managed cluster |
machine_name | string | The owning Machine’s metadata.name (may differ from the TerraformMachine’s own name) |
bootstrap_data | string, sensitive | Base64 (standard, padded) of the raw bytes of the bootstrap Secret’s value key, named by Machine.spec.bootstrap.dataSecretName. Encoded unconditionally for every bootstrap provider, since a gzip payload (CAPRKE2 gzipUserData: true) is not valid UTF-8 and cannot be a JSON/HCL string as-is. Contains the cluster CA and service-account keys for control-plane machines — treat it as a secret |
bootstrap_format | string | The bootstrap Secret’s format key (cloud-config or ignition); cloud-config when the key is absent |
failure_domain | string or null | Machine.spec.failureDomain; null when unset |
kubernetes_version | string or null | Machine.spec.version, passed verbatim (may carry a distro suffix such as +rke2r1) |
control_plane | bool | Whether the owning Machine carries the cluster.x-k8s.io/control-plane label |
A trimmed real example:
{
"captf_contract": "v1alpha1",
"captf_object": { "kind": "TerraformMachine", "name": "prod-md-0-abcde", "namespace": "team-a" },
"captf_cluster_outputs": { "network_id": "net-1", "note": "a<b & c>d" },
"machine_name": "prod-md-0-abcde-xyz12",
"bootstrap_data": "I2Nsb3VkLWNvbmZpZwpydW5jbWQ6...",
"bootstrap_format": "cloud-config",
"failure_domain": null,
"kubernetes_version": "v1.36.2",
"control_plane": false
}
The note: "a<b & c>d" above is deliberate: the machine role’s renderer
keeps a module-authored exports value’s exact bytes through the durable
Secret instead of turning </& into their \u... escapes (section
5 has the one exception
to this).
The root variables and re-exported outputs (provider_id, addresses,
failure_domain, interruptible, health) follow the same pattern as the
cluster role.
4. Machine pool role inputs
A TerraformMachinePool runs the module’s machinepool role. For the full
type table and lifecycle, see
contract/v1alpha1/machinepool.md
(this page does not duplicate it). Its inputs follow the same
captf_contract/captf_cluster/captf_object/captf_tags pattern as the
cluster and machine roles, plus captf_cluster_outputs (the owning
TerraformCluster’s exports, as for a machine), and these role-specific
inputs:
| tfvars key | Type | Source |
|---|---|---|
machinepool_name | string | The owning MachinePool’s metadata.name |
replicas | number | The group’s desired capacity. With autoscaling disabled: MachinePool.spec.replicas (1 when unset, CAPI’s own default). With autoscaling enabled: TerraformMachinePool.status.replicas (the observed count from the last refresh) once one exists, else still spec.replicas for the first apply — either way clamped into [autoscaling.min, autoscaling.max]. The write-back that syncs the raw observed count to spec.replicas is unclamped; see Machine pools |
bootstrap_data, bootstrap_format | as machine role | From the bootstrap Secret named by MachinePool.spec.template.spec.bootstrap.dataSecretName, encoded exactly as for a machine. Because the kubeadm bootstrap provider rewrites this Secret roughly every 7.5 minutes, a pool is re-applied at that cadence for its whole life; see the machinepool contract “Bootstrap rotation” |
failure_domains | list(string) | MachinePool.spec.failureDomains; may be [], meaning the module chooses |
cluster_failure_domains | list(string) | The names of the owning TerraformCluster’s own failure domains — from its state’s failure_domains output when it has state, else from TerraformCluster.status.failureDomains for an externally managed cluster. Lets a pool module spread across every domain when failure_domains is [] |
kubernetes_version | string or null | MachinePool.spec.template.spec.version; null when unset |
node_labels | map(string) | MachinePool.spec.template.metadata.labels verbatim; {} when absent. Pool instances have no Machine objects, so core CAPI never syncs labels onto their Nodes; the module renders these into the bootstrap or agent configuration it controls |
autoscaling | object{enabled,min,max} | Parsed from the owning MachinePool’s cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size/-max-size annotations, not from the TerraformMachinePool spec: a pool has no autoscaling field of its own. enabled is true only when both annotations are present and parse as 0 <= min <= max; otherwise {false, 0, 0} |
Unlike a TerraformMachine, a TerraformMachinePool is mutable: its inputs
are rebuilt, and its inputs hash recomputed, on every reconcile for the
whole life of the object, not only until it is provisioned; see
machinepool.md
for exactly which changes re-apply it.
The pool role’s renderer defaults failure_domains, cluster_failure_domains
and node_labels to []/{} when unset through a step that, as a side
effect, HTML-escapes <, > and & inside every string value it
touches — including a module-authored exports value carried into
captf_cluster_outputs. The cluster and machine roles do not escape these
characters (section 3’s example). A \uXXXX escape and its literal
character decode to the same JSON string, so a module never sees a
difference; only someone reading the raw Secret bytes does. Still, don’t
assume the two roles’ encodings match byte for byte when inspecting one.
5. What a cluster hands to its machines and pools
The cluster module’s exports output is the only handoff from a
TerraformCluster to the TerraformMachines and pools of its Cluster. It
is:
- Written by the cluster module as any Terraform value (an object is the convention).
- Stored, like every output, in the cluster’s own Terraform state — nothing else from that state is ever read by the controller for this purpose.
- Read by each machine’s or pool’s own reconcile whenever it builds its
inputs, and rendered into
captf_cluster_outputs(see section 4 for the pool role’s one encoding difference).
When it is read. A machine’s inputs (and so exports) are rebuilt on
every reconcile while it is not yet provisioned — before its status has
latched status.initialization.provisioned = true — whether or not that
reconcile ends up starting a Job. Once provisioned, the reconciler never
rebuilds that machine’s inputs again, so it never re-reads exports or the
bootstrap Secret. Concretely:
- Before a machine is provisioned, every reconcile re-reads the
cluster’s current
exportsand re-renderscaptf_cluster_outputsfrom it (along with the rest of the machine’s inputs). Only when an apply Job actually starts does the freshly rendered result get written to the durable Secret; a reconcile that only re-renders without starting a Job leaves the durable Secret untouched. - Once provisioned, a machine keeps exactly the
captf_cluster_outputsits last apply Job start wrote to its durable Secret, forever: a later change to the cluster’sexportsreaches only new machines created after that change, never existing ones. This is also whycaptf_cluster_outputschanges are never a trigger for a machine re-apply: machines are immutable, so their inputs hash, though recorded at apply time like any kind’s, is never compared against a fresh render to decide on a re-apply, and the only way a machine applies again is a fresh first apply (for example, retrying one that was interrupted before it recorded an inputs hash). - A pool re-reads
exportson every reconcile, for its whole life, unlike a machine: aTerraformMachinePoolis mutable and never latchescaptf_cluster_outputs, so a later change to the cluster’sexportsreaches every existing pool at its next reconcile, and (unlike a machine) a changedcaptf_cluster_outputsis in the pool’s inputs hash and so re-applies it. - Externally managed cluster (
cluster.x-k8s.io/managed-byannotation on theTerraformCluster): a machine’s or pool’scaptf_cluster_outputsis{}without reading any state, since there is no cluster module run and so no cluster state to read. - A malformed
exportsvalue — one that cannot be canonicalized for hashing, for example it contains a fractional or exponent number — is an output violation on the cluster side. Every unprovisioned machine, and every pool, waiting on it staysDependenciesReady=Unknownwith reasonWaitingForClusterExportsuntil the cluster module’sexportsoutput is fixed; this is the one operator-visible way a cluster module can break its machines and pools purely through inputs.
A concrete example is the libvirt modules: the cluster module’s exports
publishes
output "exports" {
value = {
network = local.network
pool = local.pool
vip = local.vip
lease = libvirt_volume.vip_lease.name
}
}
and the machine module’s captf_cluster_outputs variable carries a
validation block that rejects any value missing network, pool or
lease — refusing to run against an externally managed cluster ({}) or
a cluster module that doesn’t publish a lease, with a clear error instead of
an “unsupported attribute” failure deep inside the module.
6. User variables
Besides the contract inputs, a TerraformCluster, TerraformMachine or
TerraformMachinePool (and their templates) can pass the module its own
variables: instance sizes, CIDRs, SSH public keys, anything the module
declares. See Module Variables for how to set
them, the merge order between spec.variables and spec.variablesFrom, and
what the module sees.
7. What is (and isn’t) in the inputs hash
captf.io/inputs-hash covers exactly: the hash scheme, the contract
version, the role, spec.source.image as written (tag or digest —
never the resolved digest), the full rendered inputs of the role, and the
merged user variables when there are any (their canonical JSON and which of
them are sensitive; an object without variables keeps exactly the hash it
had before variables existed). Nothing else can change the hash: object
metadata (uid, labels, annotations, generation), the resolved image
digest, Job metadata, drift ticks, and spec.source.imagePullPolicy are
all absent from the hashed value and so cannot trigger anything through
it.
What a change does, per kind:
- TerraformCluster (mutable). A hash change re-applies: the controller
compares the freshly rendered inputs’ hash against the state’s recorded
hash on every reconcile and starts a new apply Job when they differ. A
cluster apply that would delete or replace a resource is blocked
(
ApplyJobSucceeded=False/DestructivePlanBlocked) until thecaptf.io/approve-destructive-planannotation names that exact inputs hash, and withspec.applyPolicy: Manualevery re-apply also waits for a plan approval; see Plan Approval. - TerraformMachine (immutable). Nothing, in the sense that matters: a
machine’s inputs hash is recorded at apply time like any kind’s, but
never compared against a fresh render to decide on a re-apply. Its
inputs are still rebuilt on every reconcile up to the
first successful apply (section 5), but once that apply has run and the
object is provisioned, nothing — not a spec field, not a bootstrap-data
change, not an
exportschange — is ever rendered or checked against it again. Roll out a new Machine instead. - TerraformMachinePool (mutable). A hash change re-applies, the same
way as a TerraformCluster, but with no destructive-plan guard and no
applyPolicy: an apply just runs.replicasis excluded from the hash while autoscaling is enabled, so the native autoscaler’s own scaling never triggers a re-apply on its own; seemachinepool.md“replicas”.
8. Inspecting the inputs of a live object
The durable Secret holds exactly two data keys, main.tf.json and
terraform.tfvars.json, plus three annotations recording the pinned
execution context: captf.io/image, captf.io/image-digest,
captf.io/identity. Its name follows the pattern
captf-inputs-c-<cluster-name> for a TerraformCluster,
captf-inputs-m-<machine-name> for a TerraformMachine, and
captf-inputs-mp-<pool-name> for a TerraformMachinePool.
To print the tfvars without the sensitive bootstrap_data key:
kubectl -n <namespace> get secret captf-inputs-m-<name> \
-o jsonpath='{.data.terraform\.tfvars\.json}' | base64 -d | jq 'del(.bootstrap_data)'
That still prints user variables that came from a Secret. To drop them too, delete every key the generated root declares sensitive:
s=$(kubectl -n <namespace> get secret captf-inputs-m-<name> -o json)
sensitive=$(jq -r '.data["main.tf.json"] | @base64d | fromjson | .variable // {}
| to_entries | map(select(.value.sensitive == true) | .key) | @json' <<<"$s")
jq -r '.data["terraform.tfvars.json"] | @base64d | fromjson' <<<"$s" \
| jq --argjson drop "$sensitive" 'del(.bootstrap_data) | delpaths([$drop[] | [.]])'
Warning: both the durable Secret (captf-inputs-<kindshort>-<name>) and
the per-run Secret (captf-run-<job>) carry bootstrap data in cleartext by
design — for control-plane machines that includes the cluster CA and
service-account private keys. Treat both Secrets as sensitive: avoid
kubectl get -o yaml or -o json on them, which print every key including
bootstrap_data unredacted (kubectl describe secret is safe: it prints
only key names and byte counts, never values).
9. Limits
The rendered root (main.tf.json plus terraform.tfvars.json combined) is
capped at 1,000,000 bytes, comfortably under the ~1 MiB a Kubernetes object
can hold, with headroom for the Secret’s own metadata (labels, annotations,
owner references). Exceeding it sets ApplyJobSucceeded=False/InputsTooLarge
and does not start a Job — retrying cannot help until the inputs themselves
shrink (fewer or smaller bootstrap_data, exports or user-variable bytes).
The size is reported on the captf_inputs_bytes gauge (per object, by
kind/namespace/name) even when the render was too large to run; see
metrics.md.
10. Not an input
- Cloud credentials. They come from the
TerraformClusterIdentity’s mirrored Secret, injected into the Job as environment variables and as files under a read-only mount. They are never rendered intoterraform.tfvars.json; a module reads them the way its provider expects (environment variables, a credentials file), not as a contract input. See Identities. - Backend configuration. The
kubernetesstate backend is a partial stub in the generatedmain.tf.json; the runner completes it atinittime with backend-config flags built from the object’s identity, never from tfvars. Seerunner-cli.md. - The image reference and pull policy.
spec.source.imageselects which Job runs, andspec.source.imagePullPolicyhow the kubelet pulls it; both are Kubernetes Pod fields, not Terraform inputs. Only the image’s resulting module code and its declared variables reach the tfvars. captf.io/inputs-hashand delete guards. Passed to the runner as command-line flags used to gate a destructive plan and, underapplyPolicy: Manual, an approved plan hash; seerunner-cli.md. Not part of the module’s own input surface. Nor isspec.applyPolicy: it is not hashed, so switching it re-applies nothing.- Object metadata and controller-written spec fields.
uid, labels, annotations,generation, and fields the controller writes back (spec.providerID,spec.providerIDList, an endpoint copied from the module) are never rendered or hashed, soclusterctl move, a label edit, or the controller’s own status-driven writes never trigger a re-apply. - Module-private variables. A module may declare extra variables of its own only if they carry a default; the generated root never sets them.
See also
contract/v1alpha1/README.md— the normative type and required-ness of every contract input.- Module Variables — passing your own variables to a module.
- Machine Pools — autoscaling, replicas write-back and membership refresh.
- Terraform State — where the durable Secrets sit relative to the state backend.
Terraform State
CAPTF stores every object’s Terraform or OpenTofu state in Kubernetes Secrets, alongside the object, and reads it back for outputs, drift and health. This page explains where that state lives, how CAPTF backs it up, and what happens to it when an object is deleted. For an inventory of every Secret CAPTF reads or writes, including state, see Secrets.
The backend
Every TerraformCluster, TerraformMachine and TerraformMachinePool
gets its own Terraform or OpenTofu run, backed by the kubernetes backend
configured to write into the object’s own namespace, in the default
workspace (the only one CAPTF ever uses). The backend’s secret_suffix
is derived from the object’s namespace, kind and name, which survive a
clusterctl move, and never from its UID, so state follows the object
across a move instead of being orphaned.
Secret names and the suffix
The suffix is the first 16 hex characters of
hex(sha256(<namespace>/<kind>/<name>)), a hyphen, and a short kind
code: c for TerraformCluster, m for TerraformMachine, mp for
TerraformMachinePool. It never ends in -<digits>, which the backend
would otherwise try to parse as a chunk index.
The base state Secret is named tfstate-default-<suffix>.
status.stateSecretSuffix records the suffix; it is informational only,
since the controller derives it deterministically and never reads it back.
The state lock is a coordination.k8s.io/v1 Lease named
lock-tfstate-default-<suffix> in the same namespace. See
Locks below.
Chunking and size caps
A Kubernetes Secret holds at most 1 MiB, and a large Terraform state can
exceed that once compressed. Terraform splits an oversized state across
additional Secrets named tfstate-default-<suffix>-part-1,
-part-2 and so on; OpenTofu does not chunk state at all and always
writes a single Secret. CAPTF reads whichever shape is present: it lists
every Secret carrying the backend’s own labels for the suffix, orders them
by chunk index, and concatenates their payloads before decompressing.
CAPTF caps what it is willing to read: at most 32 chunks and 64 MiB of
decompressed state. Real state compresses 10-20x, so these limits are far
beyond any plausible cluster or machine state; a state that exceeds them,
or whose chunk set is incomplete, duplicated or names an unexpected
Secret, is reported corrupt or inconsistent rather than partially read.
CAPTFStateNearSecretLimit
warns before a state’s compressed size approaches the 1 MiB Secret limit.
Locks
Both runtimes hold the lock for the duration of a run and release it on a
clean exit. A runner Job waits up to lockTimeoutSeconds (see
Job tuning for the field and its default)
for a held lock before failing. Before starting a Job, the controller
checks the lock Lease and reads its holder from the backend’s own lock
info: a lock whose holder is one of the object’s own runner pods, and
that pod either no longer exists or has already exited (a finished Job
keeps its pod object until it is pruned), is stale, and the controller
has the next Job force-unlock it automatically. A lock whose holder is
unknown, or is not one of the object’s own runner pods, is left alone
and reported as StateReadable=False/StateLocked. See the
stale state lock runbook to
force-unlock one by hand.
What CAPTF reads from state
CAPTF parses only the fields it needs from the Terraform state v4 file:
the serial, lineage, Terraform/OpenTofu version, the root module’s
outputs, and a count of managed resources (a resource with count or
for_each counts once). It also tracks the compressed size of the
concatenated chunks and, from an annotation on the base Secret, the inputs
hash of the last successful apply. Resource instance attributes are never
parsed. Outputs may be sensitive and are never logged.
After a successful apply or restore, the controller sets an owner
reference to the object on every chunk Secret (so state moves with a
clusterctl move and is garbage collected with the object) and records
the applied inputs hash on the base Secret, since Terraform’s own chunk
writes carry only the backend’s labels.
State backups
The manager keeps versioned copies of an object’s state so a lost,
corrupted or wrongly overwritten state Secret can be recovered. Whenever
the controller reads a state serial it has not observed before for the
object, it copies the state Secrets’ data verbatim into a backup set named
captf-state-backup-<suffix>-<serial> (with the same -part-N chunking
as the source), owned by the Terraform* object itself rather than by the
state, so a backup survives the state Secret being deleted by hand and is
garbage collected only when the object is. An unchanged serial is not
backed up again.
--state-backups (default 5; see manager
flags) sets how many backups per object
the manager keeps; it prunes older ones in the same pass. --state-backups=0
takes no new backups but leaves existing ones in place and restorable.
status.stateBackups lists the newest backups (see
reference/api.md for its fields).
A state that cannot be parsed with the reader’s own limits is never backed
up: encrypted, corrupt, inconsistent, or beyond the chunk and size caps
above; nor is one the manager failed to copy, for example on a transient
API error. Either way the manager logs why and counts it in
captf_state_backups_total with
result="skipped". A state that cannot be parsed is left that way and the
reconcile continues; a failed copy instead retries on the next reconcile,
since that serial is not recorded as observed. No backups are taken while
an object is being deleted.
Restore
An annotation asks the controller to push a listed backup’s content back
into the backend as a new state, through a restore Job that runs state push -force. See the state restore
runbook for the procedure,
what a restore does and does not undo, and how to verify one.
OpenTofu state encryption
OpenTofu’s client-side state encryption wraps the state file in an
envelope with no version field of its own. CAPTF detects that envelope
and reports the state as encrypted (StateReadable=False/StateEncrypted)
rather than misreading it as corrupt: reading an encrypted state’s outputs
is not supported.
State on deletion
Neither backend deletes its own state: a destroy only empties the
managed resources it recorded, and the default workspace cannot be
deleted. The controller removes the state Secrets and the lock Lease
itself, once it is safe to do so: after a destroy Job succeeds, or
immediately on deletion of an object that was never applied and so has no
state. That same cleanup also deletes the durable inputs Secret, the
object’s run lease and, for a TerraformCluster, its cluster write
lease. State backups are not deleted by
this cleanup; they are owned by the object and are garbage collected when
Kubernetes removes it after its finalizer is gone. If the finalizer is
removed by hand before the state is cleaned up, see the stuck destroy
runbook.
See also
- Secrets for the full Secret inventory, including state and its backups.
- State restore runbook.
- Stale state lock runbook.
- Stuck destroy runbook.
- Unreadable state runbook.
Drift and Health
This page explains how CAPTF checks a provisioned TerraformCluster,
TerraformMachine or TerraformMachinePool against reality: the drift
schedule, what the module’s health output means, and how both feed the
InfrastructureHealthy and Ready conditions and, for machines, Cluster
API remediation. It is for anyone who wants to understand the behavior
before changing it. To configure drift checking, see
Drift; to configure machine remediation, see
Machine Remediation; for the full reason
tables, see Conditions.
What runs, and when
Two kinds of Job read an object’s infrastructure after it is provisioned:
- A refresh Job runs
apply -refresh-only: it updates the state from reality and produces a fresh health reading, but plans nothing and finds no drift. - A drift Job does the same refresh, then plans with
-refresh=false: a plan with no changes means no drift, and any add, change or destroy in the plan is drift. A cluster or pool plans against its freshly rendered current inputs (falling back to the durable inputs Secret only when the current ones can’t be built); a machine, being immutable, always plans against its durable inputs. See Job inputs.
Both count as one health sample once the object is provisioned, and both
set the DriftJobSucceeded condition to report whether the Job
succeeded; only a drift Job sets DriftDetected. A failed refresh or
drift Job is retried under the reconciler’s normal backoff; see
Reconcile flow for how retries and requeues work in
general.
A drift check runs on a schedule: spec.drift.intervalSeconds, inherited
from the TerraformCluster’s spec.defaults.drift for a machine or pool
that sets none, else the manager’s --drift-default-interval (30
minutes by default; see Manager flags).
The first check after provisioning is due one interval after the last
successful apply, not immediately; each object’s schedule is jittered
deterministically by its UID so that objects created together don’t all
check at once. TerraformCluster and TerraformMachine both let
spec.drift.intervalSeconds: 0 turn drift checks off entirely, which
normally also stops routine health sampling after provisioning, since no
more refresh or drift Jobs run on a schedule (a reading of
InstancePending still gets occasional extra refreshes on its own,
below). A TerraformMachinePool’s own interval must be at least 1
second — the CRD rejects 0 outright — and even a cluster’s
spec.defaults.drift.intervalSeconds: 0, which does disable a machine’s
drift, leaves a pool’s on the manager’s default instead: a pool’s
periodic drift Job is what feeds its refreshed instance count into a
plan, so disabling it would stop that count from ever reaching a plan;
see
Machine pools.
Outside the drift schedule, a TerraformMachine or TerraformMachinePool
also takes a refresh right after each successful apply, to get an
immediate health reading (unless the apply’s own outputs already gave a
definite one); a TerraformMachinePool refreshes again on its own
membershipRefreshIntervalSeconds cadence to keep membership current
regardless of drift (see Machine pools),
and, with remediation.annotateMachine set, a machine refreshes on
remediation.healthCheckIntervalSeconds independent of drift too (see
Machine Remediation); none of these
apply to a TerraformCluster. For a TerraformCluster or
TerraformMachine, while a reading is InstancePending, the next
refresh backs off on its own doubling schedule (30 seconds up to 5
minutes) instead of waiting for the next drift interval, so a newly
launched instance is checked again quickly. A TerraformMachinePool’s
pending reading instead refreshes every 30 seconds flat, the same fixed
cadence its membership convergence uses, since the two converge
together; see Machine pools.
Report or remediate
Drift found by a drift Job is either just recorded or automatically
corrected, per the object’s drift action, Report or Remediate:
Report(the default for every kind) only setsDriftDetected; it records the finding but applies nothing.Remediateadditionally re-applies the object’s current inputs to remove the drift.
Only a TerraformCluster’s own spec.drift.action and a
TerraformMachinePool’s own spec.drift.action (never inherited from
the cluster’s spec.defaults.drift, which has no action field at all)
can select Remediate. A TerraformMachine’s drift policy has no action
field either, and is always Report: a machine’s instance is immutable
infrastructure, replaced by a Cluster API rollout, not reconciled in
place by a re-apply.
When a pool or cluster remediates, a finding sets DriftDetected
True/DriftPending rather than True/DriftReported, and the reconciler
starts an apply of the current inputs. While that apply runs,
DriftDetected reads True/DriftRemediating; if the apply succeeds
after the drift check that found the drift, DriftDetected clears to
False/NoDrift at once, without waiting for the next drift check. If the
apply instead fails, is blocked, or stops because its approved plan
changed, DriftDetected falls back to True/DriftPending with a note of
what happened, and the reconciler keeps retrying the remediation apply
(under the normal backoff) until as many attempts have failed since the
last successful drift check as the object’s failed-Job history keeps
(see Job tuning); past that, the drift
stays pending until the next successful check re-evaluates it.
A remediation apply is still an apply of that kind, with the same guard as any other:
- A
TerraformCluster’s remediation apply goes through the same destructive-plan guard as any other cluster apply — a plan that would delete or replace a resource blocks (ApplyJobSucceeded=False/DestructivePlanBlocked) until it is approved — or, underspec.applyPolicy: Manual, the plan-approval flow instead of the guard. See Plan preview and approval. - A
TerraformMachinePool’s remediation apply is not guarded: a pool has noapplyPolicyand no destructive-plan guard, so it just runs.
DriftDetected
DriftDetected is negative polarity: True means drift was found, and it
is never an input to Ready, so drift alone never makes an object
un-Ready. Before the first drift check it reads Unknown/DriftNotChecked.
Its full reason table, including the messages each reason carries, is in
Conditions.
From module health to InfrastructureHealthy
Every module role (cluster, machine, machinepool) declares a
health output with a state (pending, running, degraded,
stopped, terminated or unknown) and a healthy boolean, plus
optional message and reasons. Each refresh or drift Job’s outcome —
and, for a machine or pool, the reading an apply itself provides — maps
to InfrastructureHealthy:
Module health.state | healthy | InfrastructureHealthy |
|---|---|---|
| (no apply has started yet) | — | Unknown/WaitingForProvisioning |
| (before provisioned, since the first apply started) | — | False/Provisioning |
pending | any | False/InstancePending |
running | true | True/Healthy |
running | false | False/InstanceUnhealthy |
degraded | any | False/InstanceDegraded |
stopped | any | False/InstanceStopped |
terminated | any | False/InstanceTerminated |
unknown, or no health output at all | — | Unknown/HealthUnknown |
message and reasons, when the module sets them, are joined into the
condition’s message. The full reason table is in
Conditions.
For a TerraformMachine specifically, a provider_id output that turns
null after provisioning is read as terminated even though the module
reported no such health state: the instance is gone, whatever the health
output says, and spec.providerID is kept rather than cleared.
InfrastructureHealthy and Ready
InfrastructureHealthy is not itself mirrored to Cluster API, but once
an object is provisioned it becomes, with Deleting (and for a pool also
ApplyJobSucceeded), the entire set of conditions that Ready
summarizes — down from the full set of dependency, credential, RBAC,
apply, state and output conditions Ready watches beforehand. See
Reconcile flow for how status.initialization.provisioned
latches, and
Conditions for exactly
which conditions feed Ready, per kind and per phase.
Unhealthy samples
A TerraformMachine counts consecutive unhealthy samples in
status.unhealthySamples: each successful refresh or drift Job after
provisioning is one sample (an apply’s own reading counts too, when it
already gives a definite reading and stands in for the post-apply
refresh). An InfrastructureHealthy reading of Healthy resets the
count to zero; InstanceUnhealthy, InstanceDegraded or
InstanceStopped adds one; InstancePending, HealthUnknown or
InstanceTerminated leaves it as it is.
This count is what drives Cluster API machine remediation: see
Machine Remediation for the threshold,
the terminated-instance shortcut, the cluster.x-k8s.io/remediate-machine
annotation and its withdrawal.
See also
Security Model
This page states CAPTF’s trust boundary: what creating a TerraformCluster,
TerraformMachine or TerraformMachinePool grants, what the Job that runs
it can read, and what CAPTF keeps out of status, events and logs. Read it
before deciding who may create or update a Terraform* object, and before
deciding whether two tenants can share a namespace. For the mechanics behind
each control, see Identities and Credentials,
RBAC and Secrets.
A Terraform* object is a Pod
Once its Cluster API owner references it, a TerraformCluster,
TerraformMachine or TerraformMachinePool makes CAPTF run Jobs whose main
container runs spec.source.image, with jobs.env, as the namespace’s
captf-runner ServiceAccount (or an opted-in override named by
jobs.serviceAccountName), with the resolved identity’s credentials
injected as environment variables and mounted as read-only files
(Identities and Credentials). An object no
owner references runs nothing (see
Owner references are checked).
Running a Job is equivalent to granting whoever controls the image
everything the Pod runs with: the runner’s access to Secrets in the
namespace (below) and the identity’s cloud credentials. Module code, and so
any image a Terraform* object names, runs with that access; a
provider "kubernetes" {} block or a local-exec provisioner can use it
directly.
Consequences:
- The source image is the trust boundary. Whoever may set
spec.source.imageon an ownedTerraform*object in a namespace, or on the template it is cloned from, already has, in effect, the Secret access described below and the cloud credentials of every identity allowed in that namespace.TerraformClusterandTerraformMachinePoolare mutable, soupdateon them is enough. Restrictcreate/updateonterraform*kinds and their templates to principals who already hold that. The controller checks only thatspec.source.imageis a syntactically valid image reference; it does not police which registry or repository it names. An admission policy onspec.source.imageby registry prefix is the recommended control for restricting which images a namespace may run. - One identity and one workload cluster per tenant namespace. Every
Terraform*object in a namespace, and every image any of them names, can read the Secrets described below, including the credential mirrors of every other identity in use in that namespace and the state and inputs of every other object there. A namespace is the boundary between tenants; sharing one namespace between two tenants’ clusters or identities gives each tenant everything described in this page for the other’s cluster too. - Pod Security Admission applies to Jobs like any workload. A namespace that enforces it gets the defaults described in Pod security.
Owner references are checked against the owner
CAPTF treats a Cluster API object as the owner of a Terraform* object only
when the owner references it back. For a TerraformMachine, the Machine
named by its Machine-kind ownerReferences entry must pass three checks:
- The Machine’s
spec.infrastructureRefnames thisTerraformMachine: matchingapiGroup,kind: TerraformMachineandname. Cluster API sets it before it adds the owner reference, so everyTerraformMachineCluster API creates passes. - When the
ownerReferencesentry carries a UID, it equals the Machine’s UID. - When the
TerraformMachinehas acluster.x-k8s.io/cluster-namelabel, it equals the Machine’sspec.clusterName.
TerraformMachinePool applies the same checks to its MachinePool, using
spec.template.spec.infrastructureRef, and TerraformCluster to its
Cluster, using spec.infrastructureRef.
An owner that fails any check is not an owner:
DependenciesReady=False/OwnerMismatch names the failed check, no Job
runs, and CAPTF writes nothing to that object or its Cluster: no
remediation annotation on a Machine, no replica count on a MachinePool.
Deleting the Terraform* object works as it does when its owner is gone:
destroy runs from the durable inputs.
The TerraformMachine delete webhook uses the same checks. It refuses a
direct delete only while an owner that passes them exists and is not being
deleted, so an object whose ownerReferences name a Machine that does not
own it can always be deleted.
create on a terraform* kind alone therefore runs nothing: a Job starts
only once a Cluster API object that references the new object exists. Who
may create those, and who may set spec.source.image on an owned object,
is what the section above describes.
What the runner can read, and why
The runner ServiceAccount’s permissions come from a single, static
ClusterRole bound namespace-by-namespace; see RBAC
for the exact rules and how a custom ServiceAccount opts in. On Secrets, it
holds get, list, create, update and delete, and none of that is
scoped by name or label: the Terraform and OpenTofu Kubernetes state
backend needs list to enumerate its state chunks and workspaces on every
read and write, and list (like create) cannot be restricted to named
Secrets at all. get, update and delete could in principle be scoped
with resourceNames, but the backend names each state chunk itself, per
apply and per workspace, so no fixed rule can list them in advance. The
runner, and therefore any module image, can as a result read, replace or
delete every Secret in its namespace, which includes:
- the state and inputs Secrets of every
Terraform*object in the namespace, not only the one the running Job belongs to; - the credential mirrors (
captf-creds-*) of every identity in use in the namespace, not only the one the running Job was given (Secrets lists every Secret CAPTF reads or writes and its sensitivity); - any CAPI core Secret in the namespace, such as a cluster’s kubeconfig, certificate authority (CA) or a Machine’s bootstrap data. Write access here means a hostile module can substitute a cluster’s CA or kubeconfig, not only read them.
This is why one identity and one workload cluster per namespace, above, is the only real isolation CAPTF offers between tenants sharing a management cluster. The manager itself can also read every Secret in the cluster, as any CAPI infrastructure provider that runs the Kubernetes state backend effectively can.
Image pinning by digest
The first time an apply Job of a Terraform* object succeeds, the
controller records the image digest the kubelet actually ran (read from the
pod’s container status, not from spec.source.image) as captf.io/image-digest
on the object’s durable inputs Secret; a failed apply pins nothing. See
Annotations, Labels and Finalizers for
the annotation.
TerraformMachineis immutable (the kinds): once pinned, every later drift check and destroy Job for that machine runsrepo@sha256:…, never the tag inspec.source.image. A tag that moves after the successful apply can therefore never change the code that destroys an existing machine.TerraformClusterandTerraformMachinePoolare mutable: each spec-driven apply re-resolves the tag and re-pins the digest it ran.
Pinning guards against an accident — a tag moved out from under a running
cluster changing what a later destroy runs — not against a hostile image:
the pinned digest is whichever image spec.source.image named when the
apply succeeded, and that image’s own runtime computed the plan and ran the
providers. Approving a blocked destructive plan or a Manual-policy plan
preview is exactly as privileged as setting spec.source.image, since both
need update on the object; see
Plan Approval.
CAPTF’s own manager image is not signed, and nothing in CAPTF verifies a module image’s signature before running it: digest pinning fixes which image ran after the fact, it does not check who published it.
Pod security
The Job’s pod defaults satisfy the Pod Security baseline profile without
any configuration: seccomp defaults to RuntimeDefault, and fsGroup
defaults to 65532 so a non-root image user can read the credential
files, mounted at mode 0440, through that supplementary group
regardless of the image’s own user or group. The main
container additionally defaults to allowPrivilegeEscalation: false,
capabilities.drop: [ALL] and readOnlyRootFilesystem: true, and the
runner’s own init container is fixed non-root, fully locked down, and never
configurable. runAsNonRoot is not defaulted at the pod level, because
an image built FROM hashicorp/terraform runs as root unless it sets
USER; reaching the restricted profile needs an image that tolerates
runAsNonRoot: true, set through jobs.podSecurityContext or
jobs.securityContext.
Regardless of profile, the admission webhook rejects privileged: true,
allowPrivilegeEscalation: true and any capabilities.add in
jobs.securityContext, on every object and template that carries a jobs
policy (spec.jobs, spec.defaults.jobs on a TerraformCluster, and the
same field on a TerraformMachine, TerraformMachinePool and all three
*Template kinds), because that container holds the resolved identity’s
cloud credentials. Pod Security Admission remains the namespace-wide control
for everything else a jobs policy does not set, such as host namespaces and
volume types.
What CAPTF keeps out of status, events and logs
Terraform* object status is readable by anyone who can get the object —
far more people, in general, than can read Secrets in the namespace — so it
never carries raw process output. status.lastRun.error.summary is the
runner’s own short description of a failure, at most 512 bytes, never the
failing step’s stderr; the full output stays in the Job’s own logs, which
need pods/log access to read. The events the runner emits on the object
(RunStarted, StepStarted, …, RunFinished; see
Observability) carry step names, exit
codes, durations and resource-change counts, and, on failure, the same
curated summary as status — never tfvars, plan output or resource values.
A module variable sourced from a Secret
(Module Variables) is automatically declared
sensitive = true in the generated root — one sourced from a ConfigMap or
given inline never is — so Terraform redacts it from the Job’s own plan and
apply output and the controller redacts it from its own trace-level logs.
That redaction stops there: like every other input, the value is written
in clear into the object’s durable and per-run inputs Secrets and into
the Terraform state Secret, so anyone who can read
Secrets in the namespace can read it in either place
(Secrets,
Job Inputs). Cloud credentials never go through this path at
all: the resolved identity’s credentials are mounted into the Job and never
rendered into a variable, so they never reach the inputs Secrets or the
state.
Network exposure
CAPTF ships no NetworkPolicy by default; applying one is an opt-in step
covered in Installation. Without it,
nothing restricts which pods the manager or a Job can reach, or which pods
can reach them, beyond whatever the cluster otherwise enforces.
The manager needs only inbound traffic to its webhook, metrics and
health-probe ports, and outbound traffic to DNS, the API server and the
registries it reads module image metadata from. A Job’s pod is the more
sensitive workload: it holds the resolved identity’s cloud credentials and
a ServiceAccount token that can write every Secret in its namespace
(above), so its egress is worth restricting to DNS, the API server, and the
specific provider and registry endpoints the module it runs needs — nothing
reaches it inbound. A NetworkPolicy only has an effect on a CNI that
enforces one.
See also
Identities and Credentials
A TerraformClusterIdentity is a cluster-scoped object that names a Secret
of cloud credentials and the namespaces allowed to use it. Every
TerraformCluster, TerraformMachine and TerraformMachinePool needs one,
directly or by inheriting its cluster’s, before it can run a Job. This page
covers creating an identity, choosing which namespaces it allows,
referencing it from each kind, rotating and revoking its credentials, and
deleting it. For the trust an identity’s credentials carry once mounted
into a Job, see Security model.
Before you begin
- Cluster-scoped access to create
TerraformClusterIdentityobjects, and namespace access to create the Secret it names. getaccess to that Secret: creating an identity, or pointing an existing one at a different Secret, checks it.- A
TerraformCluster,TerraformMachineorTerraformMachinePoolto reference the identity from, or a namespace to create one in.
How credentials reach a Job
An identity’s credentials never reach a Job directly. Into each allowed
namespace where an object uses the identity, the controller mirrors the
source Secret as captf-creds-<identity> and keeps it in sync. Each Job
that resolves to this identity gets that mirror both ways:
envFrom, so every key becomes an environment variable — the convention most provider SDKs (AWS_*,ARM_*,GOOGLE_*, …) read. A key starting withTF_orKUBE_never reaches the module’s environment, exceptTF_IN_AUTOMATION,TF_INPUTandKUBE_NAMESPACE, which the Job sets itself and which a Secret key cannot override; see Runtime environment for why.- A read-only file mount at
/var/run/captf/credentials/<key>, one file per key, for file-based authentication (GOOGLE_APPLICATION_CREDENTIALS, a kubeconfig, a PEM key). TheTF_andKUBE_filtering above applies only to the environment: a matching key still appears as a file.
The runner ServiceAccount that a Job runs as can itself read every Secret
in its namespace; the identity mechanism controls which credentials a Job
carries, not what its ServiceAccount could otherwise reach — see
RBAC. The mirror Secret itself, its
lifecycle and its sensitivity are cataloged in
Secrets; its labels and annotations are
listed in Annotations and labels,
and the MirrorCreated/MirrorRemoved events in
Events.
Create the credentials Secret
Create a plain Opaque Secret with one key per credential your module’s
providers read:
apiVersion: v1
kind: Secret
metadata:
name: <identity-name>
namespace: <credentials-namespace>
type: Opaque
stringData:
AWS_ACCESS_KEY_ID: <value>
AWS_SECRET_ACCESS_KEY: <value>
<identity-name>— reused below as the identity’s own name; the two do not have to match, but it keeps the pair easy to find.<credentials-namespace>— any namespace the identity’s creator can read Secrets in. It is commonly a provider or platform namespace, separate from the tenant namespaces that use the identity.
The keys are exactly what reaches every Job that resolves to this identity: name them for what your module’s providers expect, not for CAPTF.
Create the identity
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformClusterIdentity
metadata:
name: <identity-name>
spec:
secretRef:
name: <identity-name>
namespace: <credentials-namespace>
allowedNamespaces:
list:
- <tenant-namespace>
allowedNamespaces decides which namespaces may use the identity:
| Value | Meaning |
|---|---|
| unset | No namespace may use it |
list | Exactly the named namespaces |
selector | Namespaces whose labels match |
selector: {} (empty selector) | Every namespace, including one that does not exist yet |
list and selector both set | The union of the two |
{} (empty object) | Rejected: write selector: {} for every namespace instead |
An empty selector reads as “every namespace” rather than “no namespace”,
so allowedNamespaces: {} is rejected outright rather than treated as one
or the other. list accepts up to 100 namespace names; selector follows
the usual label selector rules. Full field details, including size limits,
are in the
API reference.
Creating the identity, and any later change to spec.secretRef, runs a
SubjectAccessReview for the requesting user to get the named Secret,
and the request is rejected if they may not. This stops a role that may
manage identities but not read Secrets from pointing one at a Secret it
cannot read and having the controller mirror it somewhere that role can.
Other updates, including allowedNamespaces, are not re-checked.
Reference it
An identity is used through identityRef, which names it by name:
TerraformCluster:spec.identityRefis required and mutable — changing it re-resolves the cluster’s credentials on its next reconcile.TerraformClusterdefaults for machines and pools:spec.defaults.identityRef, mutable, is the identity a machine or pool without its ownidentityRefuses; unset, they fall back further to the cluster’s ownspec.identityRef. See Kinds for howspec.defaultsinheritance works, and Templates and ClusterClass for aTerraformClusterTemplatethat leavesidentityReffor a ClusterClass patch.TerraformMachine:spec.identityRefis optional, and immutable once the object is created — including from unset to a value. Change it by replacing the machine (aMachineDeploymentor control-plane rollout), not by editing it in place. A machine created with noidentityRefkeeps falling back to its cluster’s current default orspec.identityRef, so changing those still changes which credentials such a machine uses.TerraformMachinePool:spec.identityRefis optional and mutable, with the same fallback as a machine.
identityRef.name must resolve to an identity that exists and allows the
object’s namespace; the Conditions
reference has every reason IdentityAllowed and CredentialsMirrored can
carry.
Confirm it worked
kubectl get terraformclusteridentity <identity-name>
kubectl describe terraformcluster <cluster-name> -n <tenant-namespace>
The identity’s own Ready condition is True/SecretFound once its
Secret exists, False/SecretNotFound otherwise, and status.namespaces
lists every namespace currently holding a mirror of it. On the object
that references it, IdentityAllowed and CredentialsMirrored turn
True once the namespace is allowed and the mirror is in place; until
then no Job for that object starts.
Rotate credentials
-
Edit the credentials Secret’s data in place, keeping its name and namespace. The controller cannot watch that Secret, so this is not picked up at once: the mirror in each allowed namespace is rewritten the next time an object that uses the identity reconciles for any reason, and at the latest within one
--sync-period(10 minutes by default). A Job created after that reconcile gets the new values through itsenvFrom. A Job already running gets them too, but only in its file mount: the kubelet resyncs that volume from the mirror on its own schedule, even for a pod that started before the rotation; whether a long-running provider process re-reads a changed file is up to the provider. Only the environment is fixed for the life of the pod.A Job already running when you rotate keeps the old credentials in its environment for its whole run: the old credential must stay valid until every mirror has picked up the rotation and the longest
activeDeadlineSecondsany in-flight Job could still run for has passed (default 3600 seconds; see Deadlines and lock waits), not just until you edited the Secret.Every mirror carries a
captf.io/source-hashannotation of the source Secret’s data at the time it was last written; it changes whenever the mirror is rewritten, so watching it move off the value it held before the edit confirms that namespace’s mirror has picked up the rotation, without comparing credential values directly:kubectl get secret captf-creds-<identity-name> -n <tenant-namespace> \ -o jsonpath='{.metadata.annotations.captf\.io/source-hash}'Repeat for every namespace
status.namespaceslists, or diff the mirror’sdataagainst the source Secret’sdatadirectly if you want to confirm the values themselves rather than only that a rewrite happened. Only once every mirror has moved past its pre-rotation hash is the old credential safe to revoke. -
Point the identity at a different Secret, by editing
spec.secretRef. This is a change to the identity object itself, so it is watched: every object that uses the identity reconciles at once and refreshes its mirror. It also re-runs theSubjectAccessReview, so the user making the change needsgetaccess to the new Secret.
Revoke access
Editing spec.allowedNamespaces to drop a namespace is a change to the
identity object, so every object of that namespace using the identity
reconciles at once: the controller starts no new Job for it, leaves any
infrastructure it already created alone, and deletes the mirror Secret in
that namespace, whatever objects still reference it. This also blocks a
destroy: deleting such an object waits, reporting that the destroy is
waiting for the identity to allow the namespace again, until access is
restored.
Delete an identity
Deletion is refused while:
- a
TerraformCluster,TerraformMachineorTerraformMachinePool, in any namespace, currently resolves to this identity through its ownidentityRefor a cluster’s fallback — an object being deleted still counts, since its destroy still needs the credentials; or status.namespacesstill lists a namespace holding a mirror of it.
The rejection names one object still using it, or every namespace still
holding a mirror. Switching an object’s identityRef away from this
identity does not by itself release its mirror in that namespace: the
mirror keeps that object as an owner until the object is deleted, so
status.namespaces keeps listing the namespace, and the identity cannot
be deleted, until it is. If nothing resolves to the identity there
anymore, deleting the mirror Secret directly also releases the
namespace: it is only a copy, so this does not touch the source Secret,
and the identity’s status catches up as soon as the deletion is seen.
Deleting an identity never deletes the Secret it names: that Secret
belongs to you, and stays behind for you to remove or reuse.
Move considerations
TerraformClusterIdentity moves with clusterctl move, but its
credentials Secret deliberately does not: see the
move runbook for why, and the
procedure for copying it to the target cluster yourself.
See also
- Security model — the trust boundary a mounted identity’s credentials sit inside.
- Secrets — every Secret CAPTF reads or writes, including the credential mirror.
- RBAC — the runner ServiceAccount that Jobs run as.
- Kinds — how
spec.defaultsinheritance works. - Conditions — every
IdentityAllowedandCredentialsMirroredreason. - Module Variables — a
variablesFromSecret is a different mechanism: it supplies module-specific values (sizes, CIDRs, passwords), never cloud credentials, and CAPTF only ever reads it, in place, rather than mirroring it like an identity’s Secret.
Module Variables
This page shows you how to pass your own variables to a TerraformCluster,
TerraformMachine or TerraformMachinePool’s module: instance sizes,
CIDRs, SSH keys, database passwords, anything besides the contract
inputs CAPTF sets itself.
Before you begin
- The module declares each variable you set, with a default so it still
plans when nobody sets the variable
(
contract/v1alpha1/common.md). A variable the module does not declare fails the apply. - To pass a variable from a ConfigMap or Secret, you need permission to create and label one in the object’s namespace.
Set an inline variable
Add the variable to spec.variables (or spec.template.spec.variables on a
*Template kind), a JSON object:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachineTemplate
metadata:
name: workers-large
spec:
template:
spec:
source:
image: ghcr.io/example/aws-machine:v1.4.0
variables:
instance_type: m6i.xlarge
root_volume_gib: 100
subnet_ids: [subnet-0a1, subnet-0b2]
Each key becomes a named argument of the module: the generated root
declares it without a type and passes it through, so the module’s own
declared type converts the value (a ConfigMap string "100" becomes the
number 100 for a number variable).
Set a variable from a ConfigMap or Secret
-
Create the ConfigMap or Secret in the object’s namespace, and label it so CAPTF is allowed to read it:
kubectl create configmap aws-sizing -n <namespace> \ --from-literal=instance_type=m6i.xlarge kubectl label configmap aws-sizing -n <namespace> captf.io/variables=true<namespace>is the object’s own namespace:variablesFromnever reaches across namespaces. This Secret is unrelated to the credentials Secret aTerraformClusterIdentitynames: the two are separate mechanisms with separate labels and separate readers (see Identities and Credentials).The manager watches only labeled ConfigMaps and Secrets, and reads their data straight from the API server when it renders a Job; it never caches it. The label is an explicit opt-in, part of the trust boundary: an object cannot read an arbitrary Secret of its namespace just by naming it.
-
Reference it from
spec.variablesFrom:spec: variablesFrom: - configMapRef: name: aws-sizing - secretRef: name: db-credentials format: JSON - configMapRef: name: team-overrides optional: trueEvery data key of a labeled source becomes a variable of that name. Set
optional: trueon a source that may not exist yet: a missing or unlabeled required source blocks the object atDependenciesReady=False/VariablesSourceNotFoundand no Job starts; an optional one just contributes nothing. Every value, including a Secret’s, must be UTF-8; a ConfigMap’sbinaryDatais read the same asdata. -
Confirm:
DependenciesReadyturnsTrueonce every required source resolves, and the object’s next Job carries the variable. Inspect the renderedterraform.tfvars.jsonof a live object as described in Job inputs.
Merge order and formats
spec.variablesFrom, in list order: every source’s keys become variables, and a later source wins on a key it shares with an earlier one.spec.variables(inline) wins over every source.
A source’s format decides how its values are read:
String(the default): each value is passed as a string. Use it for scalars; the module’s declared type converts"3","true"and"0.5".JSON: each value is parsed as JSON, for lists, maps and objects (["a","b"],{"team":"a"}). A value that is not valid JSON holds the object atDependenciesReady=False/VariablesInvalid, naming the source and the key, never the value.
Names and limits
A variable name is a Terraform identifier: ^[a-zA-Z_][a-zA-Z0-9_-]*$.
Rejected, whatever the format:
- a name starting with
captf_; - a contract input of the object’s role — a machine cannot set
machine_name, but a cluster can, sincemachine_nameis not one of its inputs (contract reference); - a module meta-argument:
source,version,providers,count,for_each,depends_on,lifecycle,locals.
spec.variables holds at most 256 keys, and when set, at least one; the
admission webhook rejects a bad inline key or too many of them at once, and
also rejects a variablesFrom entry that names both a ConfigMap and a
Secret, or neither. A bad key in a referenced source is caught only when
the controller reads it, as VariablesInvalid naming the key.
spec.variablesFrom holds at most 16 sources.
Variables count toward the rendered-inputs limit shared with every other
input (1,000,000 bytes): past it the apply is refused with
ApplyJobSucceeded=False/InputsTooLarge
(Size limits). They are also
part of the inputs hash that decides whether a
mutable object’s next reconcile re-applies.
Sensitivity
A variable’s winning value (after the merge above) is declared
sensitive = true in the generated root when it came from a Secret, so
Terraform and OpenTofu redact it in plan and apply output. A value that
ends up inline or from a ConfigMap is not sensitive, even if an earlier,
losing source was a Secret, and even if the module’s own declaration marks
its variable sensitive. Sensitive or not, every variable is still stored
like any other input: in the object’s inputs Secrets and in the state
(Secrets, Terraform
state).
What a change does per kind
TerraformCluster: the variables are part of the inputs hash. Editingspec.variables,spec.variablesFromor the data of a referenced source re-applies the module, guarded like any other input change (Plan approval). The controller re-reads every referenced source on each reconcile, and a change to a labeled source it references wakes it.TerraformMachine:spec.variablesandspec.variablesFromare immutable. The controller reads the sources until the machine is provisioned; after that it never reads them again, and the machine’s destroy, refresh and drift runs use the variables already pinned in its durable inputs Secret. Editing a referenced ConfigMap or Secret therefore affects only machines created afterward: roll the MachineDeployment, or bump the template, to replace existing ones.TerraformMachinePool:spec.variablesandspec.variablesFromare mutable, and the sources are re-read on every reconcile for the pool’s whole life, the same as a TerraformCluster; a change to a labeled source it references wakes it immediately, the same way. The variables are part of the inputs hash, but the resulting re-apply is never guarded: a TerraformMachinePool has noapplyPolicy(Machine pools).- Templates:
spec.template.specof a*Templatekind, so itsvariablesandvariablesFromtoo, is immutable like the rest of the template (The Kinds); a ClusterClass topology patch onspec.template.spec.variablesrolls out through a new template (Templates and ClusterClass). spec.defaultson aTerraformClusterdoes not cover variables: a machine or pool without its ownidentityRef,jobsordriftinherits the cluster’s, but variables are never inherited, so set them on the machine, pool or their templates directly.clusterctl movenever carries avariablesFromsource: CAPTF puts no owner reference on a ConfigMap or Secret it reads. ATerraformClusterorTerraformMachinePool, or a machine not yet provisioned, waits atVariablesSourceNotFoundon the target until you recreate the source there (clusterctl move).
Troubleshooting
| Symptom | Reason | Fix |
|---|---|---|
DependenciesReady=False/VariablesSourceNotFound | A required variablesFrom source is missing, or not labeled captf.io/variables=true | Create or label the ConfigMap or Secret, or set optional: true |
DependenciesReady=False/VariablesInvalid | A source’s key is not a Terraform identifier, is reserved, or (format JSON) its value is not valid JSON | Rename or remove the key, or fix the source’s format |
Admission rejects spec.variables | Not a JSON object, an inline key is reserved or not a Terraform identifier, or there are more than 256 keys | Fix or trim the key named in the error |
ApplyJobSucceeded=False/ApplyFailed, status.lastRun.error.summary has Unsupported argument or Extraneous JSON object property | The module does not declare a variable you set | Add the variable to the module, or stop setting it |
See conditions.md for
every DependenciesReady reason.
See also
- Identities and Credentials — the separate Secret mechanism for cloud credentials.
common.md— how a module declares a user variable.- Job inputs — the full render pipeline and the inputs hash.
- Templates and ClusterClass — patching a variable per Cluster.
- Plan approval — the destructive-plan guard a TerraformCluster variable change is subject to.
reference/api.md— the full field list forvariablesandvariablesFrom.
Templates and ClusterClass
This page covers the clusterctl templates CAPTF ships, how to generate a
cluster from them, and the noop ClusterClass: what it creates, its
topology variables, and how to patch a module variable onto it. It is for
anyone creating clusters with CAPTF, especially with ClusterClass.
Before you begin
- The provider is installed and registered with
clusterctl(Installation). - A
TerraformClusterIdentityis applied for the namespace you are generating into (Identities and credentials). - Your cluster and machine module images are built, linted and pushed (tfcapi-lint, image contract).
- You know the values for the
clusterctlvariables you need (clusterctl variables).
The shipped flavors
CAPTF ships three flavors as templates/cluster-template*.yaml files;
once the provider is registered, clusterctl generate cluster --infrastructure terraform renders one of them from the registered
repository (Installation):
| Flavor | Template file | What it creates |
|---|---|---|
Default (no --flavor) | cluster-template.yaml | A Cluster, TerraformCluster, KubeadmControlPlane, two TerraformMachineTemplates (control plane and workers), a MachineDeployment with its KubeadmConfigTemplate, and a MachineHealthCheck for each of the control plane and the workers |
clusterclass | cluster-template-clusterclass.yaml | A Cluster with spec.topology.classRef.name: noop, referencing the noop ClusterClass |
libvirt | cluster-template-libvirt.yaml | The same shape as the default flavor, sized for the development host’s libvirt modules: kube-vip on the control plane, and nodes that install containerd and Kubernetes at boot from the version in the machine’s cloud-init metadata rather than a version clusterctl baked in |
Every flavor’s KubeadmControlPlane sets spec.remediation.maxRetry: 3
and retryPeriodSeconds: 300, so a module that fails deterministically
stops being retried instead of churning cloud resources, and its
localAPIEndpoint.bindPort (init and join) equals
Cluster.spec.clusterNetwork.apiServerPort (6443); kubeadm never reads
the Cluster field, so keep the two equal if you change one.
The libvirt flavor is for the project’s development host; see
the libvirt host guide and
the libvirt module.
Its identity template is templates/identity-libvirt.yaml, applied instead
of the default templates/identity.yaml.
Generate a cluster
Set the identity, image and sizing variables, then generate and apply. Default flavor:
export TERRAFORM_IDENTITY_NAME=<identity-name> \
TERRAFORM_CLUSTER_IMAGE=<cluster-module-image> \
TERRAFORM_MACHINE_IMAGE=<machine-module-image>
clusterctl generate cluster <cluster-name> --infrastructure terraform \
--target-namespace <namespace> \
--kubernetes-version <version> \
--control-plane-machine-count <count> --worker-machine-count <count> \
| kubectl apply -f -
ClusterClass flavor: enable the ClusterTopology feature gate when you
register the provider (Installation);
it is alpha and off by default. Apply the class to the namespace once, then
generate with --flavor clusterclass:
kubectl apply -n <namespace> -f templates/clusterclass-noop.yaml
clusterctl generate cluster <cluster-name> --infrastructure terraform \
--flavor clusterclass --target-namespace <namespace> \
--kubernetes-version <version> \
--control-plane-machine-count <count> --worker-machine-count <count> \
| kubectl apply -f -
clusterclass-noop.yaml carries no clusterctl variables of its own and no
namespace, so apply it once to every namespace that generates clusters from
it. Neither flavor defaults CLUSTER_NAME, KUBERNETES_VERSION,
CONTROL_PLANE_MACHINE_COUNT, WORKER_MACHINE_COUNT, the two image
variables or the identity name: a missing one fails clusterctl generate
instead of deploying something unintended. See
clusterctl variables for the full
list, including the libvirt flavor’s defaults.
The noop ClusterClass
templates/clusterclass-noop.yaml defines ClusterClass noop and the
templates it references:
TerraformClusterTemplate/noop, the infrastructure template.KubeadmControlPlaneTemplate/noop-control-planeandTerraformMachineTemplate/noop-control-plane, the control plane and its infrastructure.TerraformMachineTemplate/noop-workerandKubeadmConfigTemplate/noop-worker, the infrastructure and bootstrap for thedefault-workermachine deployment class.
Each of the three Terraform templates carries the placeholder image
example.invalid/captf/set-by-clusterclass:unset: a Cluster that fails to
patch a real image gets a failed image pull instead of running an
unintended module. The class also carries default MachineHealthCheck
timeouts for the control plane and the workers; the clusterclass flavor
overrides them through Cluster.spec.topology with the same variables as
cluster-template.yaml (see
clusterctl variables).
The class declares three required string topology variables —
identityName, clusterImage and machineImage — described in
the ClusterClass topology variables table.
Three patches apply them to the templates above:
identitysetsspec.template.spec.identityRefonTerraformClusterTemplate/nooptoidentityName, and also setsspec.template.spec.defaults.identityRefto the same value, so machines without their ownidentityRefinherit it.clusterImagereplacesspec.template.spec.source.imageonTerraformClusterTemplate/noopwithclusterImage.machineImagereplacesspec.template.spec.source.imageon bothTerraformMachineTemplate/noop-control-planeandTerraformMachineTemplate/noop-workerwithmachineImage, matched bymatchResources.controlPlaneandmatchResources.machineDeploymentClass.names: [default-worker].
cluster-template-clusterclass.yaml sets the three variables from
TERRAFORM_IDENTITY_NAME, TERRAFORM_CLUSTER_IMAGE and
TERRAFORM_MACHINE_IMAGE.
Patch a module variable through ClusterClass
A ClusterClass patch can also turn a topology variable into a module
variable (spec.template.spec.variables;
module variables). Declare the topology variable, then add a
patch whose jsonPatches write /spec/template/spec/variables (or one key
under it, if the template already sets others):
spec:
variables:
- name: workerInstanceType
required: true
schema:
openAPIV3Schema:
type: string
patches:
- name: workerInstanceType
definitions:
- selector:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachineTemplate
matchResources:
machineDeploymentClass:
names: [default-worker]
jsonPatches:
- op: add
path: /spec/template/spec/variables
valueFrom:
template: |
instance_type: {{ .workerInstanceType }}
The same patch on TerraformClusterTemplate reaches the TerraformCluster
instead, whose spec.variables is mutable: changing the variable
re-applies the cluster module, and a destructive plan still waits for
approval (plan approval). A ConfigMap or Secret
named in variablesFrom is not part of the ClusterClass: create it,
labeled captf.io/variables=true, in each namespace that uses the class
(module variables).
Template immutability and rolling out a change
A *Template kind’s spec.template.spec cannot be edited in place once
created; only its metadata can change (see
mutability per kind). ClusterClass topology
reconciliation is exempt from that rule for its own dry-runs, which is how
patching works at all, but a direct edit to a *Template object is
rejected.
Two situations follow from this:
- A field the class exposes as a topology variable. Changing the
variable’s value on a Cluster’s
spec.topology.variablesis enough: since the resultingTerraformMachineTemplatewould differ from the one already referenced, and templates are immutable, the topology controller creates a newTerraformMachineTemplatewith the patched spec and rolls the affectedMachineDeploymentor control plane onto it. ChangingmachineImageon aclusterclass-flavor Cluster works this way. - A field the class does not expose as a variable, such as a
KubeadmControlPlaneTemplatefield or a new patch. Create a new template object under a new name with the desiredspec.template.spec, then update theClusterClass’s reference to it —spec.infrastructure.templateRef,spec.controlPlane.templateRef,spec.controlPlane.machineInfrastructure.templateRef, or the matchingworkers.machineDeployments[].infrastructure.templateRef/bootstrap.templateRef— and apply the class. Topology reconciliation then rolls every Cluster on the class onto the new template, subject to each machine deployment’s or control plane’s own rollout strategy (thenoopclass’s control plane setsmaxSurge: 0; itsdefault-workermachine deployment class leaves the rollout strategy unset, so the defaultmaxSurge: 1applies).
A TerraformCluster itself is mostly mutable — only
spec.controlPlaneEndpoint is immutable once it has a host — so a
clusterImage or identityName change re-applies the cluster module in
place rather than creating a new object.
See also
Machine Pools
This page shows you how to create a MachinePool backed by a
TerraformMachinePool: a group of nodes CAPTF provisions and scales as one
native cloud scaling group (an autoscaling group, a scale set, an instance
group, or similar), rather than as individual TerraformMachines. It
covers fixed and autoscaled replica counts, the group’s reported members,
and deletion. See The Kinds for how a
TerraformMachinePool compares to a TerraformMachine, and the
machinepool contract
for everything the module role must implement.
Before you begin
- The provider installed, and a
TerraformClusterprovisioned or being provisioned (see Installation). - A
TerraformClusterIdentityallowed in your namespace (see Identities and Credentials). - A machinepool-role module image: one that implements the machinepool contract, managing one scaling group and reporting its provider IDs, desired capacity and members.
Create a MachinePool
A MachinePool has a single infrastructure object for its whole group,
not one per member: MachinePool.spec.template.spec.infrastructureRef
names one TerraformMachinePool directly, by kind and name, the same way
Cluster.spec.infrastructureRef names one TerraformCluster. There is no
per-replica cloning, so you create the TerraformMachinePool yourself
rather than pointing at a TerraformMachinePoolTemplate; a
TerraformMachinePoolTemplate exists only for a MachinePool a
ClusterClass topology manages, covered in
Templates and ClusterClass.
The TerraformMachinePool must carry the cluster’s
cluster.x-k8s.io/cluster-name label: CAPTF looks up the owning Cluster
by that label, not by the MachinePool’s spec.clusterName. Cluster
API’s MachinePool controller patches this label on from
spec.clusterName once the TerraformMachinePool exists, but only on
its own next reconcile; setting the label yourself avoids that wait.
For example:
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachinePool
metadata:
name: my-cluster-workers
namespace: team-a
spec:
clusterName: my-cluster
replicas: 3
template:
spec:
clusterName: my-cluster
version: v1.31.4
bootstrap:
configRef:
apiGroup: bootstrap.cluster.x-k8s.io
kind: KubeadmConfig
name: my-cluster-workers
infrastructureRef:
apiGroup: infrastructure.cluster.x-k8s.io
kind: TerraformMachinePool
name: my-cluster-workers
---
apiVersion: bootstrap.cluster.x-k8s.io/v1beta2
kind: KubeadmConfig
metadata:
name: my-cluster-workers
namespace: team-a
spec:
joinConfiguration: {}
---
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
metadata:
name: my-cluster-workers
namespace: team-a
labels:
cluster.x-k8s.io/cluster-name: my-cluster
spec:
source:
image: ghcr.io/example/machinepool-module:v0.1.0
identityRef:
name: aws-prod
The KubeadmConfig is the bootstrap provider’s own object (one per pool,
for the same reason: a pool has no per-member Machine to hold one each);
its content is the bootstrap provider’s concern, not CAPTF’s. With
spec.replicas unset, Cluster API defaults it to 1; set it to the fixed
size you want, as above. Every field of TerraformMachinePoolSpec is
mutable: changing spec.source, spec.identityRef, spec.variables or
any other field re-applies the module on the next reconcile, and there is
no spec.applyPolicy and no delete guard (see The
Kinds). Pass module-specific configuration through
spec.variables or spec.variablesFrom as for any other kind (see
Module Variables); spec.jobs tunes the Job the same way
it does for a TerraformMachine (see Tuning Jobs). A
pool with none of its own inherits spec.identityRef, spec.jobs and
spec.drift.intervalSeconds from the owning TerraformCluster’s
spec.defaults.
node_labels (rendered from spec.template.metadata.labels, not the
TerraformMachinePool’s own metadata) and the other inputs a module sees
are listed in the machinepool contract
and the environment reference; this page
does not restate them.
Choose fixed replicas or autoscaling
With no autoscaler annotations, MachinePool.spec.replicas is the sole
source of desired capacity, exactly as in the example above: change it to
resize the group.
To let the module’s own native autoscaling policy own the desired count
instead, set both annotations Cluster API’s autoscaler contract defines,
on the MachinePool:
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachinePool
metadata:
name: my-cluster-workers
namespace: team-a
annotations:
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "3"
cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "10"
spec:
clusterName: my-cluster
template:
spec:
clusterName: my-cluster
version: v1.31.4
bootstrap:
configRef:
apiGroup: bootstrap.cluster.x-k8s.io
kind: KubeadmConfig
name: my-cluster-workers
infrastructureRef:
apiGroup: infrastructure.cluster.x-k8s.io
kind: TerraformMachinePool
name: my-cluster-workers
Both annotations must be present, parse as non-negative integers, and
satisfy min ≤ max, or the pool reports
AutoscalingActive=False/AutoscalingAnnotationsInvalid
and applies without autoscaling. Valid, the pool reports
AutoscalingActive=True/ReplicasManagedByModule: the module owns the
group’s desired count and its own scaling policy (target tracking,
scheduled, or whatever it implements), and the controller claims the
cluster.x-k8s.io/replicas-managed-by annotation on the MachinePool so
Cluster API stops treating spec.replicas as authoritative. On every
reconcile the controller then writes the group’s observed desired
capacity back to MachinePool.spec.replicas, emitting a
ReplicasWrittenBack event when it changes; see
Annotations, Labels and
Finalizers
for both keys. Leave spec.replicas unset in this mode: Cluster API
defaults and clamps it from the annotations for the first apply, and the
write-back takes over from there. Removing both annotations returns
spec.replicas to being authoritative and releases
replicas-managed-by.
The Kubernetes Cluster Autoscaler itself does not drive these pools: its
clusterapi cloud provider requires MachinePool Machines, which CAPTF
does not implement. Running it against a CAPTF pool is unsupported, since
its spec.replicas patches would be overwritten by the write-back above.
Add an autoscaled pool to a generated cluster
This walks through adding an autoscaled MachinePool to a cluster
generated from the default flavor (Templates and
ClusterClass), since none of the shipped flavors creates
one on their own.
Generate the default flavor with no workers; a MachineDeployment is
still created, with spec.replicas: 0, rather than omitted:
clusterctl generate cluster my-cluster --infrastructure terraform \
--target-namespace team-a \
--kubernetes-version v1.31.4 \
--control-plane-machine-count 1 --worker-machine-count 0 \
| kubectl apply -f -
Add the pool alongside it: a MachinePool with the autoscaler
annotations, its KubeadmConfig (the flavor’s bootstrap provider is
kubeadm), and the TerraformMachinePool:
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachinePool
metadata:
name: my-cluster-workers
namespace: team-a
annotations:
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "2"
cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "5"
spec:
clusterName: my-cluster
template:
spec:
clusterName: my-cluster
version: v1.31.4
bootstrap:
configRef:
apiGroup: bootstrap.cluster.x-k8s.io
kind: KubeadmConfig
name: my-cluster-workers
infrastructureRef:
apiGroup: infrastructure.cluster.x-k8s.io
kind: TerraformMachinePool
name: my-cluster-workers
---
apiVersion: bootstrap.cluster.x-k8s.io/v1beta2
kind: KubeadmConfig
metadata:
name: my-cluster-workers
namespace: team-a
spec:
joinConfiguration: {}
---
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
metadata:
name: my-cluster-workers
namespace: team-a
labels:
cluster.x-k8s.io/cluster-name: my-cluster
spec:
source:
image: ghcr.io/example/machinepool-module:v0.1.0
identityRef:
name: aws-prod
spec.replicas is left unset on the MachinePool: with both autoscaler
annotations present and valid, Cluster API defaults and clamps it from
min-size/max-size for the first apply, and the pool’s own write-back
takes over from there (see Choose fixed replicas or
autoscaling above). Apply the
three objects, then confirm as in Confirm it worked
below.
Set the membership refresh interval
Between applies, the controller runs a refresh to pick up members joining
or leaving the group, on spec.membershipRefreshIntervalSeconds (15–86400
seconds; unset or 0 means 60):
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
spec:
membershipRefreshIntervalSeconds: 30
A new member is unschedulable until its provider ID reaches
spec.providerIDList, so a shorter interval gets new nodes ready for
workloads sooner, at the cost of more frequent Jobs. The controller also
refreshes right after every apply and, while the group has not converged
(spec.providerIDList’s length differs from status.replicas), every 30
seconds, or the configured interval instead when it is shorter.
Check the group’s members
kubectl get terraformmachinepool <name> -n <namespace>
shows the owning cluster, MachinePool, desired replicas and the Ready
condition. spec.providerIDList is every non-terminated member, sorted
and deduplicated; status.replicas is the desired capacity as of the
last refresh; status.instances is the module’s own per-member detail
(provider ID, an optional instance ID, addresses, failure domain and
health state), capped at 1000 entries:
kubectl get terraformmachinepool <name> -n <namespace> \
-o jsonpath='{.status.instances}'
See the API reference
for every status field. A pool reports provisioned, and the Ready
condition (mirrored onto the MachinePool’s InfrastructureReady)
true, once its state carries a successful apply and its module reports a
health state other than pending; unlike a fixed-replica machine, that
latch does not depend on spec.providerIDList being non-empty, since a
pool may legitimately scale to zero.
Drift on a pool
A pool’s drift check cannot be disabled: spec.drift.intervalSeconds: 0
falls back to the manager’s default interval rather than turning checks
off, because the drift Job’s own refresh is what feeds a plan; without
it, a cloud-side scaling change would never register as drift. With
spec.drift.action: Remediate, a detected difference re-applies the
pool’s current inputs; with autoscaling enabled the module is responsible
for excluding its own desired-count attribute from that plan, or every
cloud-side scale reports as drift. See Drift for setting the
interval and action, and Drift and Health
for how a check runs and feeds health.
Delete a MachinePool
Deleting the MachinePool deletes its KubeadmConfig and
TerraformMachinePool with it. Nothing blocks a TerraformMachinePool’s
deletion: it destroys the group from its durable inputs and removes its
finalizer once the destroy Job succeeds, whether or not the owning
MachinePool or Cluster still exist. See the reconcile
lifecycle for how deletion and finalizers work
across every kind.
Confirm it worked
kubectl get terraformmachinepool <name> -n <namespace>
Ready reads True once the group is provisioned, and Replicas shows
the desired capacity. kubectl get machinepool <name> -n <namespace>
shows the same replica count and InfrastructureReady=True once Cluster
API has copied spec.providerIDList and status.replicas across, which
happens only once the workload cluster is reachable.
See also
- The Kinds for how a
TerraformMachinePool’s mutability differs from aTerraformMachine’s. - The machinepool contract for every input and output a module role must implement.
- Drift and Drift and Health.
- Templates and ClusterClass for a
TerraformMachinePoolTemplateused through a ClusterClass topology. - Module Variables and Tuning Jobs.
- API Reference, Conditions, Events and Annotations, Labels and Finalizers.
Plan Approval
TerraformCluster guards two things before it changes infrastructure: a
plan that deletes or replaces a resource always waits for an approval, and
with spec.applyPolicy: Manual every plan waits for one. This page covers
both guards, how they interact, and what they do not cover.
Before you begin
- A
TerraformClusterwhose apply you want to guard or preview. kubectlaccess to annotate it and to read the source container’s logs of its Jobs.
The destructive-plan guard
A TerraformCluster is mutable: a new image tag, an edited field, or drift
Remediate re-applies its module. A module release that renames a
resource, or a different module pointed at the same state, can plan to
destroy infrastructure nobody asked to destroy. So every TerraformCluster
apply, including a drift remediation, runs its plan first and stops before
applying if any resource’s planned actions include a delete: a removal, or
a replacement in either order. Nothing changes and the state is untouched.
status.lastRun.error.kind is blocked, ApplyJobSucceeded turns
False/DestructivePlanBlocked naming the affected addresses and
actions, and a DestructivePlanBlocked event is emitted once per blocked
Job. See Conditions and
Events for the exact reasons.
While the newest apply of the current inputs hash is blocked, no apply of that hash starts; the controller re-checks at least every ten minutes. A blocked drift remediation keeps drift and health checks running; a blocked input change pauses them, since they would render the unapplied inputs. Changing the inputs, which produces a new hash, starts a new guarded apply at once, and so does approving the blocked hash. A blocked Job counts toward neither the retry backoff nor the remediation failure cap.
Approving a destructive plan
Read the plan first: kubectl logs job/<name> -c source has its
human-readable output, and the condition message lists what it deletes or
replaces. If the change is intended, approve the inputs hash named in the
condition:
kubectl annotate terraformcluster <name> -n <namespace> \
captf.io/approve-destructive-plan=<inputs-hash> --overwrite
<name> and <namespace> are the TerraformCluster’s; <inputs-hash> is
the hash from the condition message.
The approval is by inputs hash: the hash of everything the apply renders, the image reference and every module input (see Job Inputs). It approves exactly those inputs, not the object. Any later change produces a new hash, which the annotation does not name, so the next destructive plan is blocked again without anyone removing the approval.
The approval is consumed once: when an apply of the approved hash succeeds, the controller removes the annotation. A drift remediation re-applies the inputs the state already records, so without this, an approval of an input change would also cover a later destructive remediation of the same inputs; instead, that remediation is blocked again and needs its own approval.
The approval covers the plan computed when the approved apply runs, not
the plan you read: something that changed since is covered too. lifecycle { prevent_destroy = true } in the module remains the stronger control for
a resource that must never be replaced: it fails even an approved plan.
Anyone who can annotate the TerraformCluster can approve; that is the
same set of people who can change spec.source.image, which already
decides what the module does (see Security
Model). See
captf.io/approve-destructive-plan
for the annotation’s full definition.
A guarded apply whose plan has no changes is never blocked: it stops after the plan step, and the Job succeeds without an apply step, adopting the rendered inputs hash and pinning the image digest as a full apply would; see Reconcile Lifecycle.
Plan preview: applyPolicy Manual
The destructive-plan guard only stops deletes and replacements; an
in-place change to a load balancer, a network or a security group still
applies on the next edit. Teams that review every infrastructure change
set spec.applyPolicy: Manual on the TerraformCluster or its
TerraformClusterTemplate. The default, Automatic, keeps the behavior
above unchanged. applyPolicy is mutable: switching back to Automatic
applies whatever was waiting, and clears status.plan.
-
Plan. When the controller would apply new inputs, it starts a plan Job instead (
status.activeJob.operation: plan): init, validate, plan and read the plan back, with no apply step. -
Review. The runner reports the plan’s counts, the address and action of each changed resource (at most 50, never a value), and a plan hash over the sorted, non-no-op changes. The controller writes this to
status.plan(inputsHash,job,planHash,add,change,destroy,resources,truncated,createdAt), setsApplyJobSucceededtoUnknown/PlanAwaitingApprovalwith the exact approve command, and emits onePlanReadyevent with the counts. The Job’s log has the human-readable plan. -
Wait. Nothing applies and no further Job starts. The controller re-checks at least every ten minutes; a changed annotation or new inputs re-trigger it at once. The condition is
Unknown, notFalse: waiting never turnsReadyfalse. -
Approve. Set the annotation from the condition message to the plan hash in
status.plan.planHash:kubectl annotate terraformcluster <name> -n <namespace> \ captf.io/approve-plan=<plan-hash> --overwrite<name>and<namespace>are theTerraformCluster’s;<plan-hash>isstatus.plan.planHash. -
Apply. The apply Job starts (
PlanApprovedevent), plans again, and applies only if the new plan’s hash is the approved one. If anything changed since the plan was reviewed, it stops before applying:status.lastRun.error.kindisplan-changed, the new plan replacesstatus.plan,ApplyJobSucceededturnsUnknown/PlanChangedwith a new approve command, and aPlanChangedevent is emitted. The apply then waits for the approval of the new hash. A changed plan counts toward neither the retry backoff nor the remediation failure cap. -
Done. When the approved apply succeeds, the controller removes
captf.io/approve-plan, clearsstatus.plan, and emitsPlanApplied.
captf_plan_approvals_total{result="approved"|"changed"} counts applied
approvals and changed plans; see Metrics.
Gated: every apply but the first, which covers an input change (image, spec field, module variable, a Cluster input), a drift remediation, and the retry of a failed apply. Not gated: the first apply of a new cluster (there is no state yet, so nothing to break), deletion, refresh, and drift checks. A plan without changes needs no approval: its hash is that of the empty change list, and the apply proceeds at once, though it plans again first and stops if the plan is no longer empty.
With the destructive-plan guard
Approving a plan also approves its deletes and replacements: they are
listed in status.plan.resources, which the reviewer saw, and the apply
runs only that exact plan. So under Manual, the
captf.io/approve-destructive-plan annotation is never needed; the
destructive-plan guard is still armed for an apply that has no plan
approval, for example right after switching back to Automatic.
lifecycle { prevent_destroy = true } still fails an approved plan.
Caveats
- An approval names a plan hash, not an inputs hash: it approves a set of changed addresses and actions, not the values that change. Two plans that touch the same resources hash the same, whatever the new values are. Review the plan’s log, not only its hash.
- An approval is consumed only when the approved apply succeeds. One left
behind, because the plan changed or the inputs changed before the apply
ran, approves any later plan with the same hash, for example a
recurring drift on the same resource. Remove a stale approval with
kubectl annotate terraformcluster <name> -n <namespace> captf.io/approve-plan-. status.planis status, not durable state: afterclusterctl move, which does not restore status, the controller plans again, and an approval still on the annotation applies if the new plan hashes the same.- In a GitOps setup, the annotation is set by a person, or by a pipeline
after its own review of
status.planand the plan Job’s log. It is not part of the desired state in Git: a controller that syncs annotations from Git would re-add a consumed approval.
What is not guarded
Neither guard applies to a TerraformMachine or a TerraformMachinePool.
A machine is immutable: Cluster API replaces it with a new object instead
of changing it in place, so there is nothing for the destructive-plan
guard to catch, and applyPolicy does not exist outside
TerraformCluster. A machine pool is mutable and its spec.variables and
spec.variablesFrom change over its whole life, like a TerraformCluster
does, but its applies are never guarded: it has no applyPolicy field,
and the destructive-plan guard only watches TerraformCluster applies.
See Machine Pools and the
reconcile lifecycle.
Confirm it worked
- After approving a destructive plan,
kubectl describe terraformcluster <name> -n <namespace>showsApplyJobSucceededback toTrueand thecaptf.io/approve-destructive-planannotation gone. - After approving a plan under
Manual,status.planis empty andApplyJobSucceededisTrue/ApplySucceeded.
See also
- Drift for
drift.action: Remediate, which the destructive-plan guard also covers. - Reconcile Lifecycle for where these guards sit in the reconcile flow.
- Annotations, Labels and Finalizers for every annotation this page uses.
- Conditions and Events for the exact reasons.
Drift
CAPTF periodically refreshes and plans a TerraformCluster,
TerraformMachine or TerraformMachinePool to check whether the
infrastructure still matches its module’s inputs. This page covers how to
set the check interval and action for each kind, disable checks where
that is possible, read the results, and remediate what a check finds. See
Drift and Health for how the checks
work and how they feed the object’s health.
Before you begin
- A provisioned
TerraformCluster,TerraformMachineorTerraformMachinePool. kubectlaccess to edit the object and read its status.
Set the drift interval
Every kind checks for drift on a timer: spec.drift.intervalSeconds. Left
unset, it uses the manager’s --drift-default-interval flag (30 minutes
by default; see Manager Flags).
On a TerraformCluster:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformCluster
spec:
drift:
intervalSeconds: 900
A TerraformMachine and a TerraformMachinePool inherit an interval from
the owning TerraformCluster’s spec.defaults.drift.intervalSeconds
when they set none of their own:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformCluster
spec:
defaults:
drift:
intervalSeconds: 1800
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
drift:
intervalSeconds: 900
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachinePool
spec:
drift:
intervalSeconds: 900
A machine’s or pool’s own intervalSeconds always wins over the
inherited default. spec.drift is operational policy and stays mutable
for the object’s whole life, unlike the fields that define the machine or
pool itself (see The Kinds for what is mutable
per kind).
Set the drift action
spec.drift.action decides what happens when a check finds changes:
Report (the default) only records the finding, Remediate applies the
object’s current inputs to remove it. Both a TerraformCluster and a
TerraformMachinePool accept action:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformCluster
spec:
drift:
action: Remediate
A TerraformMachine’s drift is always reported, never remediated: the
underlying instance is immutable infrastructure, replaced by a rollout
rather than patched in place, so spec.drift has no action field.
spec.defaults.drift on the TerraformCluster has no action field
either: only a machine’s or pool’s own spec.drift.action sets it.
Disable drift checks
Setting intervalSeconds: 0 disables drift checks on a TerraformCluster
or a TerraformMachine. Disabling them also stops every health sample
taken after provisioning, since health is otherwise read only from a
completed refresh or drift Job: the InfrastructureHealthy condition
stops updating (a TerraformMachine still refreshes independently while
its spec.remediation.annotateMachine is true; see Machine
Remediation).
A TerraformMachinePool’s drift cannot be disabled; see
Drift and health
for why.
Read the results
DriftDetected reports the last check’s outcome:
Unknown/DriftNotCheckedbefore the first check completes.False/NoDriftwhen the last check found no changes.True/DriftReportedwhen it found changes and the action isReport.True/DriftPendingwhen it found changes, the action isRemediate, and the remediation has not started or failed.True/DriftRemediatingwhile a remediation apply runs.
kubectl get terraformcluster <name> -n <namespace> \
-o jsonpath='{.status.conditions[?(@.type=="DriftDetected")]}'
Its message names the drift Job and, when changes were found, the counts
of resources to add, change and destroy. status.lastRun.drift carries
the same counts as separate fields plus up to 20 affected resource
addresses, and status.lastDriftCheck is when the last successful check
completed. DriftJobSucceeded separately reports whether the check (or
refresh) Job itself succeeded, independent of what it found. See
Conditions and API
Reference for every reason and field.
Remediate drift
With action: Remediate set as above, a TerraformCluster or
TerraformMachinePool re-applies its current inputs whenever a check
finds changes, retried with the same backoff as any other apply (see
The reconcile lifecycle) and capped: after
failedJobsHistoryLimit (job tuning) failed remediation
applies since the last successful check, DriftDetected stays
DriftPending until the next check succeeds. A TerraformCluster
remediation apply also goes through the destructive-plan guard (or plan
approval under applyPolicy: Manual); the guard does not apply to a
TerraformMachinePool (see Plan Approval). See
Drift and Health
for exactly how a remediation runs and what each DriftDetected reason
during it means.
A TerraformMachine’s drift is never remediated this way: an unhealthy
instance is instead signaled to Cluster API for replacement. See Machine
Remediation.
Confirm it worked
After a check completes, status.lastDriftCheck advances and
DriftDetected reads False/NoDrift or True with a reason above.
After a Remediate apply succeeds, DriftDetected returns to
False/NoDrift and status.lastRun.drift is empty again.
See also
- Drift and Health for how checks and health sampling work.
- Machine Remediation for handling an unhealthy
TerraformMachine. - Plan Approval for the destructive-plan guard a
TerraformClusterremediation apply goes through. - Machine Pools for the pool’s separate membership refresh schedule.
- Conditions and Manager Flags.
Machine Remediation
A TerraformMachine reports its instance’s health on the
InfrastructureHealthy condition. By itself that only ever turns the
owning Machine’s InfrastructureReady condition False; nothing acts on
it unless something is watching. spec.remediation optionally asks
Cluster API to replace an unhealthy instance, through a
MachineHealthCheck and the Machine’s owner. This page covers
configuring it, how it reaches a replacement, and what it does to the
Machine along the way. See Drift and
Health for how health itself is
computed.
Before you begin
- A provisioned
TerraformMachinewhose owning Machine is selected by aMachineHealthCheckwith anunhealthyMachineConditionscheck onInfrastructureReady. kubectlaccess to edit theTerraformMachineand read its owning Machine.
Enable remediation
spec.remediation.annotateMachine defaults to false, meaning CAPTF
never touches the Machine. Set it to true to have CAPTF request
remediation once an instance has been unhealthy for long enough:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
remediation:
annotateMachine: true
spec.remediation is operational policy: it stays mutable for the
machine’s whole life, unlike spec.source and spec.identityRef, which
are fixed at creation.
Tune the threshold and sampling interval
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformMachine
spec:
remediation:
annotateMachine: true
unhealthyThreshold: 5
healthCheckIntervalSeconds: 120
unhealthyThreshold(1-100, default 3) is how many consecutive unhealthy health samples are needed before CAPTF signals remediation. A sample is one completed refresh or drift Job; a single transient reading is never enough on its own, since once aMachineHealthCheckacts on the signal, the Machine is replaced. A terminated instance skips the threshold; see Terminated instances below.healthCheckIntervalSeconds(60-86400, default 300) is how often a provisioned instance is refreshed to resample its health whileannotateMachineistrue, independent ofspec.drift.intervalSeconds. WithannotateMachineset tofalseit has no effect, and health is then re-read only at the drift (or refresh) cadence set byspec.drift.intervalSeconds; see Drift for that setting, including howintervalSeconds: 0stops health sampling entirely onceannotateMachineisfalse.
How it works with a MachineHealthCheck
- Health outside
running/healthy, read after provisioning, turnsInfrastructureHealthyFalse, which turnsReadyFalse, which Cluster API mirrors into the owning Machine’sInfrastructureReadycondition. - A
MachineHealthCheckwithspec.checks.unhealthyMachineConditions: [{type: InfrastructureReady, status: "False", timeoutSeconds: <N>}]selecting the Machine starts its own timeout onceInfrastructureReadyturnsFalse. - When the timeout (or, with
annotateMachine, the annotation below) marks the Machine unhealthy,MachineHealthChecksets the Machine’sOwnerRemediatedcondition. Only the Machine’s owner acts on that condition: aMachineDeployment’sMachineSet, or a control-plane provider that implements remediation, such asKubeadmControlPlaneorRKE2ControlPlane. A single-replica control plane refuses to remediate itself. - The owner that acts replaces the Machine (and, through its deletion,
the
TerraformMachine) with a new one rather than repairing it in place:spec.sourceis immutable, so there is nothing for CAPTF to patch.
annotateMachine does not replace this flow; it feeds it a faster, more
specific signal than the timeout alone (see the next section).
The remediate-machine annotation
With annotateMachine: true, once unhealthyThreshold consecutive
samples are unhealthy, degraded or stopped (or the instance is
terminated), CAPTF sets cluster.x-k8s.io/remediate-machine on the
owning Machine. This asks the MachineHealthCheck reconciler to treat
the Machine as unhealthy at once, bypassing unhealthyMachineConditions’
timeout; remediation still goes through the owner as described above.
CAPTF marks the annotation as its own with a second annotation,
captf.io/remediation-requested, whose value is why it was set, for
example InstanceUnhealthy for 5 consecutive samples or the instance is terminated.
CAPTF removes both annotations once the instance reads Healthy again
and the Machine is not being deleted, withdrawing a request that has not
yet been acted on. It never removes cluster.x-k8s.io/remediate-machine
when captf.io/remediation-requested is absent: an annotation set by
someone else is left alone.
See Annotations, Labels and Finalizers for both annotations’ full definitions.
Terminated instances
A terminated instance skips unhealthyThreshold entirely: CAPTF sets the
annotations on the very first sample that reads the instance terminated,
since there is nothing to wait for and no risk of a transient reading.
Confirm it worked
kubectl get machine <name> -n <namespace> -o jsonpath='{.metadata.annotations}'showscluster.x-k8s.io/remediate-machineandcaptf.io/remediation-requestedonce a request is made, and neither once the instance recovers or is replaced.- A
RemediationRequestedevent on theTerraformMachinemarks a new request, andRemediationWithdrawnmarks a withdrawal; see Events. captf_remediation_requests_total{action="requested"|"withdrawn"}counts both; see Metrics.
See also
- Drift and Health for how health samples are produced.
- Drift for
spec.drift.intervalSeconds, which paces health sampling whenannotateMachineisfalse. - API Reference for
MachineRemediation’s full field list. - Conditions for
InfrastructureHealthy’s reasons.
Tuning Jobs
Every operation CAPTF runs — apply, destroy, drift, refresh, restore and
plan — is one Kubernetes Job. spec.jobs on a TerraformCluster,
TerraformMachine or TerraformMachinePool tunes that Job: its resources,
deadlines, history, environment, pull secrets, ServiceAccount and security
contexts. Every field is optional and has a built-in default, so you only
set the fields you need to change. For the full field list, types and
validation see
JobPolicy in the API reference; for the
Job’s exact shape (containers, mounts, args) see
Job Environment.
Before you begin
- Know which object’s
spec.jobsyou are changing. ATerraformClusteralso hasspec.defaults.jobs, which itsTerraformMachines andTerraformMachinePools inherit; see Inheriting from a cluster’s defaults below. - A
TerraformCluster’s ownspec.jobsdoes not inheritspec.defaults.jobs: defaults are for its machines and pools, never for the cluster itself.
Resources
spec.jobs.resources sets the main container’s resources as a whole;
when unset the controller applies its own default (250m CPU / 512Mi memory
requested, 2Gi memory limit, deliberately no CPU limit — throttling a slow
apply is worse than a slow apply). Raise it when a module pulls large
provider plugins or holds a large plan in memory; the init container that
copies the runner binary is not configurable, since it never varies with
the module. See
Default resources for the
exact values.
Deadlines and lock waits
activeDeadlineSecondsbounds the whole Job, at most one day. Defaults to 3600 (one hour). Raise it for a module with a slow apply or many resources; a Job that hits its deadline is interrupted with SIGTERM, the same as a pod eviction or deletion. For an apply, destroy or plan Job,ApplyJobSucceededis set False with reasonJobDeadlineExceeded; for a drift or refresh Job,DriftJobSucceededis set False with reasonDriftJobDeadlineExceeded; for a restore Job,RestoreJobSucceededis set False with reasonRestoreFailed. See Conditions.lockTimeoutSecondsis passed to the runtime as-lock-timeout. Defaults to 300 (five minutes). Raise it when Jobs commonly queue behind each other on the same state lock.- When a policy sets both,
lockTimeoutSecondsmust be less thanactiveDeadlineSeconds: the webhook rejects a policy where a lock wait alone could fill the whole deadline. This check looks at one policy at a time, so it does not see a machine’s or pool’s field-wise merge with a cluster’sspec.defaults.jobs— settinglockTimeoutSecondson a machine while onlyspec.defaults.jobs.activeDeadlineSecondsbounds it is not caught at admission.
History limits
successfulJobsHistoryLimit and failedJobsHistoryLimit cap how many
finished Jobs of each kind are kept, both defaulting to 3. Each is scoped
per object and per operation: apply, destroy, drift, refresh, restore and
plan are pruned independently, so lowering one does not shrink another’s
history. The newest Job of an operation is always kept, even at a limit of
0 — for a failed operation, until a newer Job of the same operation
succeeds. failedJobsHistoryLimit also bounds retry backoff, since backoff
looks at the same retained failures; see
The reconcile lifecycle for how backoff is
computed.
Environment variables
spec.jobs.env adds environment variables to the main container. Entries
named TF_* or KUBE_* are reserved for the runner and the Job’s own
environment (TF_IN_AUTOMATION, KUBE_NAMESPACE, and so on): an entry
using one of those names is silently dropped instead of applied. See
spec.jobs.env rejected names
for the full list of names the Job already sets.
Image pull secrets
spec.jobs.imagePullSecrets covers both the role image (spec.source.image)
and the runner’s own init image, so one list is enough even when they come
from different registries.
ServiceAccount
spec.jobs.serviceAccountName overrides the runner ServiceAccount. Left
unset, the controller creates and uses captf-runner, bound to the static
captf-runner ClusterRole. An override must already exist and carry the
label captf.io/runner=true; without it, no Job is created and
RunnerRBACReady reports False with reason ServiceAccountNotOptedIn
(see Conditions). Setting up
and opting in your own ServiceAccount, and the RBAC CAPTF manages around
it, is covered in RBAC.
Security contexts
spec.jobs.securityContext sets the main container’s securityContext.
The container holds cloud credentials, so the webhook rejects
privileged: true, allowPrivilegeEscalation: true and any
capabilities.add on it, whatever else the policy sets. Every other field
you set is applied on top of the controller’s defaults (no privilege
escalation, every capability dropped, a read-only root filesystem), except
capabilities: setting your own replaces the default drop: [ALL]
wholesale, so include ALL in your own drop list to keep every
capability dropped. runAsNonRoot is left to you, since some role images
built from a hashicorp/terraform base still run as root.
spec.jobs.podSecurityContext sets the Job pod’s securityContext. The
controller defaults its seccompProfile to RuntimeDefault and its
fsGroup to the runner’s UID, so a non-root image user can read the
identity credential files (mode 0440) through that group; set your own
fsGroup to override it, or your own runAsNonRoot/runAsUser to run the
main container as a specific non-root user. For the trust boundary these
defaults protect — why the Job is treated like any other pod that holds
cloud credentials — see The security model.
Inheriting from a cluster’s defaults
A TerraformMachine’s or TerraformMachinePool’s spec.jobs is merged
field by field with its TerraformCluster’s spec.defaults.jobs: a field
the machine or pool sets wins, an unset one falls back to the cluster’s
default, and a field neither sets gets the built-in default above. Two
fields merge instead of falling back as a whole:
envis merged by name — the machine’s or pool’s entries first, then any of the cluster default’s entries whose name they do not already use.imagePullSecretsis the union of both lists, without duplicates, the machine’s or pool’s first.
resources, securityContext and podSecurityContext are each replaced
as a whole: setting any one of them on the machine or pool drops the
cluster default’s value for that field entirely, rather than merging
individual keys inside it. See
What a cluster passes to its machines and pools
for how this fits the rest of spec.defaults.
Defaults are resolved at reconcile time and never persisted, so raising a
cluster’s spec.defaults.jobs reaches every existing machine and pool on
their next reconcile.
Confirm it worked
kubectl get job -n <namespace> -l captf.infrastructure.cluster.x-k8s.io/owner-name=<object-name>
kubectl get job -n <namespace> <job-name> -o jsonpath='{.spec.activeDeadlineSeconds}{"\n"}{.spec.template.spec.serviceAccountName}{"\n"}'
<namespace> and <object-name> are the namespace and name of the
TerraformCluster, TerraformMachine or TerraformMachinePool you
changed; <job-name> is one Job name from the first command’s output.
Compare the Job’s spec.template.spec.containers[0].resources,
securityContext and env against what you set; a field you expected to
change but that still shows the built-in default usually means it was set
on the wrong object, or dropped by the merge rules above.
See also
- API Reference for every
spec.jobsfield, its default and its validation. - Job Environment for the Job’s fixed fields, mounts, security contexts and runner args.
- RBAC for the runner ServiceAccount, its RoleBinding and the opt-in sweep.
- The security model for the trust boundary a Job’s security context protects.
- The reconcile lifecycle for retry backoff and how a Job’s operation is chosen.
Module Contract
The interface between the CAPTF controller and the Terraform/OpenTofu modules it runs. A module targets exactly one contract version; the image contract is in Image Contract.
| Version | Status | Documents |
|---|---|---|
v1alpha1 | Provisional: frozen for implementation, and may still change until the first real module has provisioned a cluster | v1alpha1 (changelog) |
v1alpha1
cluster-api-provider-terraform is a Cluster API infrastructure provider that runs Terraform/OpenTofu child modules written to its contract, packaged as OCI images together with the runtime. Existing root modules do not drop in: the provider owns the root module, the state backend and the provider installation.
Read this first. A module written to this contract is a child module. It must not declare a
backend: the controller generates the root module, configures the state backend and supplies the provider installation (the image’s mirror or the registry), passes exactly the contract inputs and reads the contract outputs back from state.
Status: Provisional: frozen for implementation, and may still change
until the first real module has provisioned a cluster (see
CHANGELOG.md). This contract’s text was checked against
the OpenTofu v1.11.5 docs; that is a documentation-verification snapshot,
not a runtime requirement — the reference images
(image-contract.md) ship a newer OpenTofu.
A module is a Terraform/OpenTofu module that implements one role
(cluster, machine, machinepool). It is delivered as an OCI image
that bundles the module code and the runtime binary
(image-contract.md): the image is the artifact,
the version and the runnable object; there is no separate module source
and no separate runtime image. The controller never calls a module
directly: it generates a root module inside the Job that (a) configures
the kubernetes backend, (b) calls the module at /captf/module with
exactly the contract inputs, and (c) re-exports the contract outputs at
the root so they land in state. Everything the controller learns about
infrastructure comes from those root outputs.
Files:
| File | Contents |
|---|---|
| Image Contract | Fixed paths, optional provider mirror, labels, execution environment, pinning, reference Containerfiles |
| Common | Inputs injected into every role; the health output every role must produce |
| Cluster Role | Cluster role (InfraCluster) |
| Machine Role | Machine role (InfraMachine) |
| MachinePool Role | MachinePool role (InfraMachinePool) |
schemas/ | JSON Schemas (draft 2020-12) for the cluster, machine and machinepool inputs and outputs, shared definitions, and valid/invalid examples; checked by hack/verify-schemas.sh |
| Job Inputs | Where every terraform.tfvars.json value comes from in the running controller: field-by-field sources, the durable and per-run Secrets, the inputs hash, and how to inspect a live object’s inputs safely |
Versioning
- The contract version is a string:
v1alpha1. It is independent of the CRD API version and of the provider release. - The image tag is the module version.
spec.source.image(registry/repo:tagor@sha256:…) is the only version a CR carries; a new module version is a new image, referenced by a new Terraform*Template. The controller pins the digest after the first run (seeimage-contract.md“Versioning and pinning”). - A module does not declare its contract version or role in a way the
controller acts on; there is no manifest file. The role is implied by
which CRD references the image, and the contract version is the one the
controller release generates against (injected as
captf_contract, recorded in object status). The OCI labelsio.captf.contractandio.captf.roleare informational (seeimage-contract.md).tfcapi-linttakes--roleand--contracton the command line and checks the labels against them. - Compat policy:
- Within one version: only additive, optional changes (new optional inputs with defaults; new optional outputs). Never rename or retype.
- Making an input required, removing an output, or changing a type means a new contract version.
- The controller supports N and N-1 contract versions concurrently; the generated root differs per version.
v1alpha*versions carry no compat guarantee between alphas; this policy starts atv1beta1.
Naming rules
- Reserved prefix
captf_. Every controller-injected input and every contract-defined local starts withcaptf_. A module MUST NOT declare a variable or output with this prefix except those defined by the contract. - Contract outputs are not prefixed (
provider_id,addresses, and so on) because they are part of the module’s public interface. Their names are reserved per role; a module MUST declare every required output of its role. - Module inputs are the contract inputs plus user variables. Anything
module-specific (instance type, machine image, region) is either fixed
inside the module or a user variable: an ordinary module variable the
object sets through
spec.variables/spec.variablesFrom(seecommon.md“User variables”). A module SHOULD give every non-contract variable a default, so it plans when no user variable is set. - Extra, non-contract outputs are allowed on any role but are never read
by the controller. The only cluster-to-machine/pool handoff is the
cluster role’s
exportsoutput (seecluster.mdandcaptf_cluster_outputsincommon.md). - Tagging is mandatory. Every role receives
captf_tags(seecommon.md) and MUST apply it to the cloud resources it creates; the CAPI provider best practices require a mechanism for identifying a provider’s cloud objects.
Trust boundary
spec.source.image is executed, unpoliced, as a Job under the runner
ServiceAccount, which can read every Secret in the object’s namespace and
runs with the identity’s cloud credentials mounted: referencing an image is
equivalent to granting its publisher that access. See Security
Model for the full threat model; this
contract only requires that module authors and operators treat an image
reference accordingly.
Null semantics
- Required outputs must be declared; their value may be
nulluntil known.nullmeans “not yet known”, never “empty”. - The controller treats
nullfor a required output as “keep waiting” (the object stays unprovisioned, conditionOutputsValid=Unknown/OutputsPending), unless the role spec says otherwise (for examplecontrol_plane_endpointon cluster may stay null forever when CAPI supplies the endpoint). - An absent output (not declared in the module) is a hard error:
OutputsValid=False/OutputsMissing, with no retry until the module changes. - Empty collections are meaningful and distinct from null (
addresses = []means provisioned with no addresses).
Type conventions
- Types are Terraform type constraints. Object attributes listed as
optional use
optional(type)/optional(type, default)(Terraform ≥1.3 / OpenTofu; both in scope; see the Terraform v1.3.0 changelog and OpenTofu docs: type constraints).outputblocks take no type constraint: their arguments arevalue,description,sensitive,ephemeral,depends_on,deprecated, pluspreconditionblocks (see OpenTofu docs: outputs). Output “types” in this contract are the shape the controller decodes, and anyoptional(..., default)on an output is applied by the controller, not by Terraform.
- All outputs are read from state JSON; the controller decodes by the
contract’s expected shape, not by the
typerecorded in state. Extra attributes on objects are ignored; missing required attributes are an error. - Strings that map to Kubernetes fields must satisfy the target field’s
validation:
provider_id1–512 chars (seeinfra-machine.md“InfraMachine: provider ID”); failure-domain names 1–256 chars, no pattern (seecluster_types.goFailureDomain.Name); address types from the CAPIMachineAddressTypeenum (seecommon_types.go). The controller rejects invalid values withOutputsInvalidrather than writing them. sensitive = trueon a module output marks the value sensitive, and sensitivity propagates to anything derived from a sensitive value (for examplebootstrap_data). A rootoutput "x" { value = module.role.x }referencing a sensitive value failsplanwith “Output refers to sensitive values” unless the root output itself setssensitive = true(reproduced with OpenTofu v1.11.5; see OpenTofu docs: outputs “sensitive”). Sensitive values are still stored in state in cleartext (same page). The generated root marks every re-exported output"sensitive": true, so a module may mark any contract output sensitive without failingplan. This has no effect on the controller: the state Secret still stores the value in cleartext, andsensitiveonly changes CLI display.
Injection mechanics (for spec readers)
- The generated root lives at
/captf/work/rootinside the Job and calls the module asmodule "role" { source = "../../module" }, that is,/captf/moduleinside the image. Terraform/OpenTofu local paths must begin with./or../, so the absolute form/captf/moduleis not used. It is a local path, soinitfetches no module from anywhere (seeimage-contract.md“Fixed paths”). Providers come from/captf/providerswhen the image ships a mirror, else from the registry. - Contract inputs are supplied via
terraform.tfvars.jsonin the generated root (auto-loaded; see OpenTofu docs: variables) and forwarded to the module call by name. Bootstrap data is markedsensitive = trueon the root variable. - Every contract input is a root variable named identically to the module
input; the root forwards
x = var.x. Nothing else is passed to the module call. - Contract outputs are re-exported at root as
output "x" { value = module.role.x, sensitive = true }: every re-export is marked sensitive, so a module marking any contract output sensitive never failsplan(see Type conventions above). - The rendered
terraform.tfvars.jsonof the newest apply is kept in a Secret, written when that apply Job starts (a failed apply may already have created resources from those inputs, so destroy must use them), owned by the Terraform* object (captf-inputs-<kindshort>-<name>), together with the resolved image digest (captf.io/image-digest) and the identity. Drift and destroy render from that Secret, never from the live CAPI objects: destroy must work after the bootstrap Secret, the owning Machine/MachinePool, or the Cluster are gone, and drift on an immutable machine must see exactly the inputs it was applied with. The per-run Secret handed to a Job is deleted when the Job completes. The state Secret stores every value (includingbootstrap_data, sensitivity notwithstanding) in cleartext for as long as the object exists; modules that can avoid keeping bootstrap data in state (for example by hashing it into a trigger) SHOULD. - State lives in the
kubernetesbackend, in thedefaultworkspace (the only one CAPTF ever uses), locked with a Lease (see OpenTofu docs: kubernetes backend and Terraform State for Secret naming, locks and backups). - What triggers a re-apply (mutable roles): a change in the hash of a
canonical struct of the spec-derived inputs. Rendered inputs are
hashed inputs: object metadata (
uid, labels, annotations, generation) and fields the controller itself writes back (spec.providerID,spec.providerIDList,spec.controlPlaneEndpoint) are neither rendered nor hashed, so aclusterctl move, a label edit or the controller’s own status-driven writes never re-apply anything, and a driftRemediate(which renders the current hashed set) cannot apply them either (seecommon.md“What is hashed”).
Roles and CAPI mapping (summary)
| Role | CAPI kind | Immutable spec | Drift | Refresh | Destroy on delete |
|---|---|---|---|---|---|
| cluster | TerraformCluster | no (re-apply on change) | Report (default) / Remediate | with drift; while pending 30s, backing off to 5m | yes, after all machines/pools of the Cluster are gone (DeletionBlocked/DependentsExist otherwise) |
| machine | TerraformMachine | yes (provider choice; the InfraMachine contract only recommends classifying fields as immutable/mutable, infra-machine.md “support for in-place changes”) | Report only; unhealthy → InfrastructureHealthy=False → Ready=False → MHC | with drift; while pending 30s, backing off to 5m | yes; direct deletes refused while the owner Machine is not deleting |
| machinepool | TerraformMachinePool | no (re-apply on change) | Report (default) / Remediate; apply -refresh-only then plan, same as every kind — relies on the module’s ignore_changes when autoscaling.enabled, see machinepool.md | with drift; membership refresh every spec.membershipRefreshIntervalSeconds (default 60), and every 30s while pending or not yet converged | yes; destroy from the durable inputs Secret, not the bootstrap Secret |
Delivery pipeline for a module author: tfcapi-lint module <dir> --role <role>, then build the image (see
image-contract.md for the reference
Containerfiles), then push, then tfcapi-lint image <ref> --role <role>,
then reference the tag from a Terraform*Template. Lint runs on the source
directory before the build; the image check confirms the layout after
it.
Conditions: health feeds InfrastructureHealthy, which is an explicit
input to the Ready summary. Deleting is also in Ready’s
ForConditionTypes for every kind, declared in
NegativePolarityConditionTypes{Deleting}, so Ready=False while a
destroy runs (the CAPI v1beta2 convention). Drift results
(DriftDetected), drift-Job outcomes, Paused and DeletionBlocked are
not, so Ready — mirrored by CAPI into the Cluster’s/Machine’s
InfrastructureReady — reflects infrastructure only. Ready is Unknown
while an object waits on dependencies, False/Provisioning from apply
start, and True when provisioned (see
machine.md “Ready timeline”).
A destroy that fails permanently is resolved by cleaning up out of band
and removing the finalizer by hand, after backing up or un-owning the
tfstate-* Secrets (they are garbage-collected with the object). There is
no skip-destroy annotation.
Common
Applies to every role. Inputs here are injected by the controller; modules MUST declare them (they may ignore them). Outputs here MUST be declared by every role.
Inputs
| Name | Type | Required | Value source |
|---|---|---|---|
captf_contract | string | yes | Literal "v1alpha1"; lets a module assert the contract it was generated for |
captf_cluster | object (below) | yes | The owning CAPI Cluster |
captf_object | object (below) | yes | The Terraform* object being reconciled |
captf_cluster_outputs | any | machine, machinepool only | The cluster role’s exports output, read from the cluster state (below) |
captf_tags | map(string) | yes | Fixed tag keys the controller always sets (below) |
variable "captf_contract" { type = string }
variable "captf_cluster" {
type = object({
name = string
namespace = string
})
}
variable "captf_object" {
type = object({
kind = string # TerraformCluster | TerraformMachine | TerraformMachinePool
name = string
namespace = string
})
}
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" { type = map(string) }
captf_cluster_outputs
The value of the cluster role’s exports output
(cluster.md), read from the cluster state. Nothing
else from the cluster state is injected.
- Not passed to the cluster role. Its generated root declares no such
variable and its tfvars carry no key, so a cluster module either leaves it
undeclared or declares it with
default = null, as the skeleton below does. A declaration without a default failsvalidate. {}when the TerraformCluster is externally managed (thecluster.x-k8s.io/managed-byannotation; the provider then never runs the cluster module, so there is no cluster state to read).- A machine/pool whose owning Cluster’s
spec.infrastructureRefis not kindTerraformClusteris never rendered at all: it setsDependenciesReady=False/ClusterNotTerraformand does nothing, since there is no cluster state to readexportsfrom. - This is how a machine/pool module gets network ids, security groups,
subnet ids and similar cluster-wide values. The cluster module puts them
in its
exportsoutput; the controller reads the cluster state and injects exactly that value.exportsmay be markedsensitive: the generated root re-exports every contract output with"sensitive": trueregardless, so a sensitiveexportsnever failsplan, and sensitivity has no effect on the controller — the state Secret still stores the value in cleartext; only CLI display is affected (seeREADME.md“Type conventions”). Machine/pool apply is gated on the cluster being provisioned andexportsreadable, so the value is complete by then. - Because
captf_cluster_outputsisany, a change in its content is not treated as a spec change for machines (immutable) but is for pools and clusters (a value change triggers a re-apply). See each role’s Lifecycle section.
captf_tags
Fixed keys the controller always sets: captf.io/cluster (Cluster name),
captf.io/namespace, captf.io/kind (the Terraform* kind),
captf.io/name (the Terraform* object name), captf.io/managed-by =
captf, and captf.io/template.
captf.io/templateis the value of the object’scluster.x-k8s.io/cloned-from-nameannotation — the template it was cloned from byexternal.GenerateTemplate(util.go,common_types.go) — when present, else"". The key is always present: it is read from the Terraform* object’s own metadata, so an object created directly gets"". It is hashed like the other keys even though annotations in general are not:GenerateTemplatewrites it once at clone time, but for a TerraformCluster or TerraformMachinePool managed by topology, a ClusterClass rebase onto a differently named template rewrites the annotation through SSA, which re-applies that object once; for immutable machines it is pinned at first apply.- Modules MUST apply these tags to every cloud resource they create that
supports tags or labels, so resources can be attributed and
garbage-collected out of band. This follows the CAPI provider best
practices: a tagging/labeling mechanism for identifying cloud objects
“MUST always be provided”
(
best-practices.md;security-guidelines.md, housekeeping).
What is hashed
The controller re-applies mutable kinds when the hash of the rendered
inputs changes. Rendered inputs are hashed inputs: the hash covers a
canonical struct of exactly the spec-derived inputs the module receives,
and nothing is rendered outside it (the one special case is pool replicas
while autoscaling.enabled, rendered as the module’s own observed desired
capacity; see machinepool.md).
captf_cluster and captf_object therefore carry only kind, name and
namespace:
- no
uid— aclusterctl movere-creates objects with new UIDs, and hashing them would re-apply every cluster and pool after a move; - no
generation— it changes whenever the controller itself writesspec.providerID,providerIDListorcontrolPlaneEndpoint; - no
annotations— they toggle around every Job (clusterctl.cluster.x-k8s.io/block-move) or are edited by GitOps tools; - no
labels— a metadata edit must never reconcile live infrastructure.
Cloud tags come only from captf_tags, whose keys are fixed. Drift for
mutable kinds renders the current hashed set; immutable machines render
drift and destroy from the durable inputs Secret below.
Durable inputs
The rendered inputs (including bootstrap_data) are stored in a Secret
owned by the Terraform* object (captf-inputs-<kindshort>-<name>), so
drift and destroy of immutable machines re-feed exactly the values used at
apply, and the Secret moves with the object. The resolved image digest
(captf.io/image-digest, see
image-contract.md
“Versioning and pinning”) and the identity are pinned there too, so an
immutable machine’s drift and destroy always run the exact image that
created it.
User variables
Everything a module is parameterized by beyond the contract inputs
(instance type, disk size, subnets, a database password) is an ordinary
Terraform variable of the module. The object’s owner sets it through
spec.variables (inline JSON) or spec.variablesFrom (ConfigMaps and
Secrets labeled captf.io/variables=true) on the TerraformCluster,
TerraformMachine or their templates; see
Module Variables for the operator side.
For module authors:
- Declare each one as an ordinary variable, with a type and a
default. The generated root passes a user variable as a named argument ofmodule "role", declared in the root without a type, so your declaration’s type converts the value: anumbervariable accepts40inline and"40"from a ConfigMap. A variable nobody sets gets itsdefault; without one, every object that does not set it failsplanon the missing argument.tfcapi-lintreports such a variable asinput/user-variable-default(a warning: a module may deliberately require it). - Mark secrets
sensitive = truein the module. The root declares a variablesensitiveonly when its value came from a Secret; your declaration makes the value sensitive inside the module whatever its source. A value that arrives sensitive stays sensitive in everything derived from it: an output built from it must be declaredsensitive = true(orplanfails with “Output refers to sensitive values”), and it cannot drivefor_each. Either way the value is stored in cleartext in the inputs Secrets and in the state, like every input (see Module Variables); credentials the module runs with come from the identity, never from a variable. - Names. User variables are Terraform identifiers
(
^[a-zA-Z_][a-zA-Z0-9_-]*$). Reserved, and rejected before anything is rendered: everycaptf_name, every contract input of the module’s role (machine_nameon a machine,control_plane_initializedon a cluster, and so on) and the module meta-argumentssource,version,providers,count,for_each,depends_on,lifecycle,locals. Do not rely on a user variable to carry a contract input’s value. - Unknown names fail the apply. A variable the object sets and the
module does not declare is rejected by Terraform/OpenTofu (
Unsupported argument; for the JSON root,Extraneous JSON object property). The controller does not check names against the image. - Hashing. User variables are part of the inputs hash only when the
object sets some, so modules and objects without them hash exactly as
before. For a TerraformCluster a changed variable re-applies (guarded by
the destructive-plan approval, see
cluster.md); a TerraformMachine reads its variables until it is provisioned, and its later runs use the ones pinned in its durable inputs Secret, like every other input. - The role input schemas (
schemas/*-inputs.json) describe the contract inputs only; the renderedterraform.tfvars.jsoncarries the user variables next to them.
Outputs
| Name | Type | Required | Maps to |
|---|---|---|---|
health | object (below) | yes | Ready condition; machine remediation; drift reporting |
output "health" {
value = {
state = "running" # closed enum: pending | running | degraded | stopped | terminated | unknown
healthy = true
message = null # optional string, surfaced in condition message
reasons = [] # optional list(string), machine-readable
}
}
state is a closed enum. Any other string is
OutputsValid=False/OutputsInvalid.
Controller-side semantics. health is written to the
InfrastructureHealthy condition. InfrastructureHealthy is one of the
explicitly listed inputs to the Ready summary; so is Deleting (negative
polarity, so Ready=False while a destroy runs). Drift results, drift-Job
outcomes, Paused and DeletionBlocked are not, so Ready — and therefore
the Cluster/Machine InfrastructureReady mirror that MachineHealthCheck
acts on — reflects infrastructure health, not controller housekeeping.
state | healthy | InfrastructureHealthy | Effect |
|---|---|---|---|
pending | any | False/InstancePending | Provisioning in progress; object not marked provisioned even if other outputs are set. While pending, the controller refreshes (apply -refresh-only) regardless of drift.intervalSeconds, so provisioning never waits for a drift tick: for cluster and machine, 30s after the last refresh, then 1m, 2m, 4m and at most every 5m while the readings stay pending (plus the object’s jitter); for a machinepool, every 30s flat, since a pending pool is also not converged (see machinepool.md “Membership refresh”) |
running | true | True/Healthy | Ready=True once role outputs are satisfied |
running | false | False/InstanceUnhealthy | machine: remediation candidate |
degraded | any | False/InstanceDegraded | machine: remediation candidate |
stopped | any | False/InstanceStopped | machine: remediation candidate |
terminated | any | False/InstanceTerminated | machine: remediation candidate; object stays provisioned and spec.providerID is kept |
unknown | any | Unknown/HealthUnknown | Ready=Unknown; no remediation |
Once provisioned, the Machine controller treats an empty InfraMachine
spec.providerID as “waiting” and logs it; it is not a hard error, but
clearing it would stall the Machine
(machine_controller_phases.go
reconcileInfrastructure), so a terminated reading keeps providerID
rather than clearing it.
healthis re-read on every refresh (apply -refresh-only, which updates state and root output values to match remote objects; see OpenTofu docs:cli/commands/plan/“Planning Modes”): the pending refresh loop above, every drift tick, and for pools the membership refresh (seemachinepool.md). For immutable machines that is the only way it changes after apply. Modules that cannot observe health return{state="running", healthy=true}(the no-op module does this) orunknown.- Out-of-band termination. If, after provisioning, a refresh makes
provider_idcome backnull(the resource vanished from state) or the refresh itself fails because the instance is gone, the controller treats that asterminated:InfrastructureHealthy=False/InstanceTerminated,spec.providerIDkept. Modules SHOULD still reportterminatedexplicitly where the cloud API exposes it, because detection through a vanished resource is slower (drift interval plus a Job) and CAPI’s Node-based checks will usually fire first. - A group with zero members (pool role,
replicas == 0) reports{state="running", healthy=true}. - Only the machine role triggers remediation. Cluster/pool
healthfeeds conditions only.
Role-specific inputs
kubernetes_version is a role input, not a common one: for machines/pools
it comes from Machine.spec.version / MachinePool.spec.template.spec.version
(authoritative for nodes); for the cluster role it comes from
Cluster.spec.topology.version (null without ClusterClass; see
cluster.md). cluster_network,
control_plane_initialized and the non-module control_plane_endpoint
(including one set later by a control-plane provider) are cluster-role
inputs (see cluster.md); machine/pool modules get
endpoint-derived values through exports.
Cluster Role
Implements the CAPI InfraCluster contract
(infra-cluster.md)
for a TerraformCluster. One workspace/state per TerraformCluster.
Typically owns the cluster-wide substrate: network, subnets, security
groups, load balancer / control-plane endpoint, shared images/keys. Values
that machine/pool modules need are handed over through the exports
output, injected into them as captf_cluster_outputs.
Inputs
Common inputs apply (common.md), except
captf_cluster_outputs: it is not passed to this role. The skeleton below
declares it with default = null, which is safe; a declaration without a
default fails validate.
| Name | Type | Required | Value source |
|---|---|---|---|
control_plane_endpoint | object({host=string, port=number}) or null | yes (may be null) | Any endpoint the module does not own (below) |
kubernetes_version | string or null | yes (may be null) | Cluster.spec.topology.version; null without ClusterClass (below) |
control_plane_initialized | bool | yes | Cluster.status.initialization.controlPlaneInitialized, latched (below) |
cluster_network | object({pods=list(string), services=list(string), service_domain=string, api_server_port=number}) or null | yes (may be null) | Cluster.spec.clusterNetwork (below) |
control_plane_endpoint (input)
Value = Cluster.spec.controlPlaneEndpoint whenever the annotation
captf.io/endpoint-source is not module — whether the user set it before
the first apply or a control-plane provider (hosted control planes) sets it
later
(cluster_controller_phases.go).
While the Cluster field is not yet valid (APIEndpoint.IsValid, host and
port both set), the input falls back to a valid
TerraformCluster.spec.controlPlaneEndpoint (the two are equal once CAPI
copies ours,
cluster_controller_phases.go,
so the copy never changes the hash); else null. It is always null when
endpoint-source=module.
Non-null means “use this endpoint, don’t create one”: the contract allows
the endpoint to come from the user, the control plane provider, or the
infra provider
(infra-cluster.md
“InfraCluster: control plane endpoint”). Shape =
APIEndpoint{host 1–512 chars, port 1–65535}
(cluster_types.go).
Provenance (captf.io/endpoint-source): user is written before the
first apply when either field holds a valid endpoint; module is written
when the controller first copies a non-null module output (from an apply
whose rendered input was null) into TerraformCluster.spec.controlPlaneEndpoint;
otherwise the annotation is absent. Once written it never changes. With
module the input stays null for the life of the object even after CAPI
copies the module’s endpoint to Cluster.spec: feeding it back would tell
the module to drop the load balancer it created, and would change the
inputs hash. In the inputs hash: an endpoint set later by a control-plane
provider therefore re-applies the cluster once, which is how the module
learns it.
kubernetes_version (input)
Cluster.spec.topology.version
(cluster_types.go);
null when the Cluster has no topology (no ClusterClass). For modules
that provision managed/hosted control planes or version-dependent cluster
resources. Passed verbatim (it may carry a distro suffix such as
+rke2r1; strip it as the machine role does). In the inputs hash, so a
topology upgrade re-applies the cluster.
control_plane_initialized (input)
Cluster.status.initialization.controlPlaneInitialized
(cluster_types.go;
nil maps to false). In the inputs hash, so the cluster re-applies exactly
once when the control plane comes up (CAPI latches the field).
The controller latches it too: the rendered value is true when the
Cluster field is true or the last rendered value in the durable inputs
Secret (captf-inputs-c-<name>, which moves) is true; it is never rendered
false after an apply with true. Cluster status is not moved by
clusterctl move, so without this a first reconcile on the target could
render false and destroy every gated resource.
Modules that need a live workload API server (in-cluster add-ons,
cloud-controller secrets, DNS records pointing at registered nodes) gate
those resources on it, for example
count = var.control_plane_initialized ? 1 : 0.
cluster_network (input)
Cluster.spec.clusterNetwork (pods.cidrBlocks, services.cidrBlocks,
serviceDomain, apiServerPort;
cluster_types.go
ClusterNetwork, NetworkRanges); null when the Cluster sets none of
them; missing attributes are null or [].
api_server_port is not the kube-apiserver bind port with KCP —
CABPK/KCP never read Cluster.spec.clusterNetwork.apiServerPort; the real
bind port is KubeadmConfig localAPIEndpoint.bindPort (default 6443,
workload_cluster.go).
By convention the module’s LB backend port is
cluster_network.api_server_port ?? 6443, and templates must keep
apiServerPort and bindPort equal — the controller has no way to
enforce this.
Outputs
Common outputs apply (health).
| Name | Type | Required | Maps to |
|---|---|---|---|
control_plane_endpoint | object({host=string, port=number}) or null | yes (may be null) | TerraformCluster.spec.controlPlaneEndpoint → Cluster.spec.controlPlaneEndpoint (below) |
failure_domains | list(object({name=string, control_plane=optional(bool, true), attributes=optional(map(string), {})})) or null | yes (may be null or []) | TerraformCluster.status.failureDomains (below) |
exports | any (object recommended) | yes (may be null, treated as {}) | Injected verbatim into machine/pool modules as captf_cluster_outputs (below) |
control_plane_endpoint (output)
Maps to TerraformCluster.spec.controlPlaneEndpoint →
Cluster.spec.controlPlaneEndpoint, surfaced once InfraCluster
initialization completes
(infra-cluster.md).
It is written only when all of the following hold: the output is non-null;
the field is still empty; and the apply’s rendered control_plane_endpoint
input was null (the first such write records
captf.io/endpoint-source=module).
It is written only when the output is valid: CAPI’s copy-to-Cluster.spec
path only fires once APIEndpoint.IsValid() is true, which requires both
host and port to be set
(cluster_types.go),
so a half-set output (only one of the two present) is treated as
OutputsValid=False/OutputsInvalid, not partially written. CAPI copies the
InfraCluster endpoint only while Cluster.spec.controlPlaneEndpoint is not
yet valid, so later changes are never propagated
(cluster_controller_phases.go
reconcileInfrastructure); CAPI’s Cluster admission webhook has no
immutability rule for it
(cluster.go).
Once emitted, the value MUST be stable for the life of the object. CAPI
never updates Cluster.spec.controlPlaneEndpoint after the first valid
copy
(cluster_controller_phases.go),
so a module that later replaces its load balancer (a new DNS name or IP)
breaks every existing kubeconfig and the apiserver certificate SANs, with
no way for CAPI to notice or recover. Keeping it stable is the module
author’s responsibility first: the controller blocks, but cannot fix, a
planned replacement.
Every TerraformCluster apply (a re-apply, a new image tag, a drift
remediation) stops before a plan that deletes or replaces any resource,
sets ApplyJobSucceeded=False/DestructivePlanBlocked with the affected
addresses, and applies it only once the captf.io/approve-destructive-plan
annotation names that apply’s inputs hash; see
Plan Approval for the destructive-plan
guard. With spec.applyPolicy: Manual every re-apply instead waits for the
approval of its plan, which covers its deletes; see
Plan Approval for plan preview. An
operator can still approve a plan that replaces the endpoint’s resource;
lifecycle { prevent_destroy = true } on the resource behind the endpoint
turns such a plan into a failed Job even when approved, instead of a lost
cluster.
When no user endpoint exists, the module MUST emit this output no later
than the apply that makes the cluster provisioned (see EndpointAvailable
below).
failure_domains (output)
Maps to TerraformCluster.status.failureDomains (v1beta2 list shape:
listType=map, listMapKey=name, MaxItems=100;
FailureDomain{name 1–256 chars, controlPlane *bool, attributes map[string]string};
infra-cluster.md
“InfraCluster: failure domains”,
cluster_types.go).
CAPI’s controlPlane has no default; optional(bool, true) is our
default, applied by the controller. null and [] both clear the field.
exports (output)
Injected verbatim into machine/pool modules as captf_cluster_outputs. May
be marked sensitive: the generated root re-exports it "sensitive": true
regardless, so this has no effect on the controller (see
README.md “Type conventions”). Put secrets
in the identity, not here. Convention: publish the load balancer
target/backend-pool id(s) here for control-plane machine modules to
register against (see
machine.md “Control-plane
machines”).
Provisioned rule
status.initialization.provisioned is the v1beta2 contract field:
status.initialization.provisioned = true when the state Secret carries
the inputs-hash annotation of a successful apply, the required outputs are
valid, and health.state != "pending".
Latched. The formula applies only until it first holds; once
provisioned is true it stays true for the object’s life, whatever later
health or outputs say (CAPI never flips infrastructureProvisioned back;
machine_controller_phases.go
“should not flip back”).
There is no endpoint gate on provisioned: a null
control_plane_endpoint output with no endpoint on the Cluster is
allowed, because hosted-control-plane providers set the endpoint only
after infrastructure is provisioned. Provisioned is derived from state,
never from Job history, so it is rebuilt identically after clusterctl move (status is not moved; the target evaluates the formula once, then
latches).
Hairpin reachability
The control-plane endpoint MUST be reachable from the first control-plane
node itself, not only from the management cluster and other nodes.
kubeadm’s getKubeConfigSpecsBase (kubernetes/kubernetes release-1.33,
cmd/kubeadm/app/phases/kubeconfig/kubeconfig.go around lines 582-626)
sets the server URL of admin.conf, super-admin.conf and kubelet.conf
to GetControlPlaneEndpoint(cfg.ControlPlaneEndpoint, &cfg.LocalAPIEndpoint)
— that is, the endpoint, not the node’s local address — while only
controller-manager.conf/scheduler.conf use GetLocalAPIEndpoint.
Because kubeadm’s post-init phases and the first node’s own kubelet talk to
the workload cluster through admin.conf/kubelet.conf, the load
balancer or VIP that backs control_plane_endpoint must hairpin: the
control-plane instance that is itself a backend must be able to reach the
frontend it was just registered behind.
EndpointAvailable condition
Cluster only, not part of Ready: False/WaitingForEndpoint once
provisioned is true and neither the module’s control_plane_endpoint
output nor Cluster.spec.controlPlaneEndpoint is a valid APIEndpoint
(host and port both set); True/EndpointAvailable otherwise.
With KCP this case is not cosmetic: KCP creates no Machine at all until
Cluster.spec.controlPlaneEndpoint.IsValid()
(kubeadmcontrolplane_controller.go
— “Waiting for Cluster spec.controlPlaneEndpoint to be set”), so a module
that outputs null with no user-provided endpoint silently stalls the
cluster forever with no other signal. tfcapi-lint warns
(output/endpoint-never-set) when a module’s control_plane_endpoint
output is a literal null with no conditional path to a value.
Lifecycle
- Mutable. Any change to
spec.source.image(a pull-policy change alone only affects how the next Job pulls, and starts none), a non-modulecontrol_plane_endpoint(including one set later by a control-plane provider),cluster_network,kubernetes_version(Cluster.spec.topology.version), orcontrol_plane_initializedflipping true starts a new apply Job (detected by the hash of the rendered spec-derived inputs; labels, uid, and controller-written fields are neither rendered nor hashed, seecommon.md“What is hashed”). - Externally managed. A TerraformCluster carrying
cluster.x-k8s.io/managed-byis skipped entirely: no Jobs, no status writes (infra-cluster.md“Externally managed infrastructure”). Machines and pools of that Cluster receivecaptf_cluster_outputs = {}. - Drift. On
spec.drift.intervalSeconds:apply -refresh-onlyfirst (so outputs, health andexportsrefresh and the plan compares against current remote values), thenplan -detailed-exitcode(0 = no changes, 1 = error, 2 = changes; see OpenTofu docs:cli/commands/plan/). Exit 2 setsDriftDetected=True; withaction: Report(the default) that is all, and an operator decides; withaction: Remediatethe controller applies, re-rendering the current inputs. A changedexportsvalue re-applies pools (mutable) but not machines (immutable). Drift results and drift-Job outcomes never feedReady, and after first provisioning neither does a failed re-apply: it isApplyJobSucceeded=FalseplusDriftDetected, neverReady=False, because ClusterInfrastructureReady=Falsewould suspend every MachineHealthCheck of the cluster (machinehealthcheck_targets.go). - Delete ordering (enforced by the controller). Destroy is blocked with
DeletionBlocked=True/DependentsExist(requeue) while any TerraformMachine or TerraformMachinePool labeledcluster.x-k8s.io/cluster-name=<cluster>(ClusterNameLabel,common_types.go) exists in the namespace, since their modules depend on cluster resources viaexports. CAPI’s own Cluster deletion already removes Machines/MachinePools first — it requeues while descendants, including MachinePools, exist, then deletes the control plane object, then the InfraCluster (cluster_controller.goreconcileDelete,clusterDescendants) — so this guard only covers out-of-band deletions of the TerraformCluster. Then adestroyJob runs; success drops the state Secret(s) and Lease, and removes the finalizer. Module authors may rely on this ordering. - The module MUST be idempotent under repeated apply with unchanged inputs
(plan must be empty).
tfcapi-lintcan’t check this, and no automated test does yet; the module author must.
Control-plane provider requirements
What a control-plane provider needs from the cluster module, and how KubeadmControlPlane (KCP) and RKE2ControlPlane (RCP) differ. The KCP column draws on the KubeadmControlPlane integration guide; the RCP column draws on the RKE2ControlPlane integration guide.
| Requirement | KubeadmControlPlane | RKE2ControlPlane |
|---|---|---|
| Endpoint required before any Machine | Yes: KCP creates no Machine until Cluster.spec.controlPlaneEndpoint.IsValid() (kubeadmcontrolplane_controller.go; see the KubeadmControlPlane guide) | Yes: RCP returns before creating any Machine while !IsValid() (rke2controlplane_controller.go; see the RKE2ControlPlane guide) |
| LB listeners | endpoint.port → CP nodes :bindPort (default 6443); one frontend | endpoint.port → CP nodes :6443, and :9345 → CP nodes :9345 on the same host as the endpoint (the join URL is https://<endpoint.host>:9345, port ignored); two frontends |
| Health checks | TCP on the backend port, or HTTPS /readyz//healthz without certificate verification; must go green with a single backend | 9345: TLS GET /v1-rke2/readyz expecting 403 (unauthenticated), or plain TCP; 6443: TCP, or HTTPS /healthz only if anonymous auth is enabled |
| Backend membership | Module registers/deregisters CP instances in its own state; instance MUST be in the LB backend before kubeadm init/join finishes, or controlPlaneInitialized never latches (infra-machine.md) | Same registration contract; module MUST add a CP node before its Machine is Ready (joins through the LB during bring-up) and MUST tolerate a backend whose supervisor is not yet up |
| SG / firewall ports | Backend port (bindPort) from LB, nodes and management; 2379–2380 and 10250 between CP nodes; 10250 from CP to workers | CP↔CP 2379-2381; all→CP 6443 and 9345; all↔all 10250, NodePort range; CNI ports for the chosen serverConfig.cni (default canal: 8472/udp, 9099); egress to get.rke2.io/GitHub releases unless air-gapped |
cluster_network.api_server_port / service_domain | api_server_port is not wired to the kube-apiserver bind port by CABPK/KCP either — the real bind port is KubeadmConfig localAPIEndpoint.bindPort (default 6443); backend port is api_server_port ?? 6443 by convention only | No effect under RKE2: the domain comes from serverConfig.clusterDomain, and the API server is always 6443 on nodes; RKE2 never reads either field |
failure_domains[].control_plane default | KCP filters to entries with controlPlane == true; a nil controlPlane counts as false, so our optional(bool, true) default is required for domains to be visible | Same: RCP only uses entries with controlPlane: true (nil counts as false); our default true is compatible with both providers |
See the KubeadmControlPlane and RKE2ControlPlane guides for the reasoning and exact citations behind each row.
Minimal skeleton
A block written on one line may hold at most one argument
(OneLineBlock;
tofu validate reports “Invalid single-argument block definition” on a
violation), so a variable that needs both type and default (like
captf_cluster_outputs below) must use the multi-line block form:
variable "captf_contract" { type = string }
variable "captf_cluster" { type = object({ name = string, namespace = string }) }
variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) }
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" { type = map(string) }
variable "control_plane_endpoint" {
type = object({ host = string, port = number })
default = null
}
variable "kubernetes_version" {
type = string
default = null
}
variable "control_plane_initialized" { type = bool }
variable "cluster_network" {
# Every attribute is always present when cluster_network is non-null
# (an unset CIDR list renders as [], an unset scalar as null), so the
# type matches the Inputs table exactly: none of the four is optional().
type = object({
pods = list(string)
services = list(string)
service_domain = string
api_server_port = number
})
default = null
}
# Tag every cloud resource this module creates with captf_tags (common.md
# "captf_tags"); this stub only has to reference it, not create anything.
resource "terraform_data" "tags" { input = var.captf_tags }
output "control_plane_endpoint" { value = var.control_plane_endpoint != null ? var.control_plane_endpoint : { host = "...", port = 6443 } }
output "failure_domains" { value = [] }
output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } }
# Handed to machine/pool modules as captf_cluster_outputs:
output "exports" { value = { network_id = "...", subnet_ids = ["..."] } }
Machine Role
Implements the CAPI InfraMachine contract
(infra-machine.md)
for a TerraformMachine. One workspace/state per TerraformMachine. The
module creates exactly one node instance and reports its provider ID,
addresses, and health. The machine is immutable (spec.source and
spec.identityRef cannot change): apply once, refresh/drift for health,
destroy on delete.
Inputs
Common inputs apply (common.md), including
captf_cluster_outputs.
| Name | Type | Required | Value source |
|---|---|---|---|
machine_name | string | yes | Owning CAPI Machine.metadata.name (the TerraformMachine name may differ) |
bootstrap_data | string, sensitive | yes | Base64 of the bootstrap Secret’s value key (below) |
bootstrap_format | string | yes | cloud-config or ignition, from the bootstrap Secret’s format key (below) |
failure_domain | string or null | yes (may be null) | Machine.spec.failureDomain (below) |
kubernetes_version | string or null | yes (may be null) | Machine.spec.version (below) |
control_plane | bool | yes | true when the owning Machine carries the control-plane label (below) |
bootstrap_data
Base64 (standard alphabet, padded) of the raw bytes of the value key
of the Secret named by Machine.spec.bootstrap.dataSecretName. The
bootstrap contract requires a single key value
(bootstrap-config.md
“BootstrapConfig: data secret”; dataSecretName in
machine_types.go
Bootstrap).
The controller base64-encodes value when rendering tfvars, always and
for every bootstrap provider: CAPRKE2 with gzipUserData: true writes
value as raw gzip bytes, which are not valid UTF-8 and so cannot be a
JSON/HCL string (see the RKE2ControlPlane guide);
encoding unconditionally keeps one rule instead of sniffing content. The
input stays sensitive = true.
Modules use one of two patterns:
- pass it unchanged to a
user_data_base64-style argument (for example AWSaws_instance.user_data_base64oraws_launch_template.user_data, which both take base64); or base64decode(var.bootstrap_data)where the argument wants the plain payload — only valid when the payload is UTF-8, that is not gzipped:base64decodefails on non-UTF-8 bytes, so a binary payload must go to a base64-taking argument instead.
Apply waits until the Secret exists (the contract workflow also exits while
dataSecretName is nil;
infra-machine.md
“Typical InfraMachine reconciliation workflow”);
dataSecretName: "" (the field allows MinLength=0,
machine_types.go)
is treated exactly like nil (WaitingForBootstrapData).
For control-plane Machines the payload embeds the cluster CA and service
account private keys uncompressed
(controlplane_init.go,
controlplane_join.go),
so it is both larger and more sensitive than a worker’s. Size limits (for
example AWS’s 16 KiB user-data) and keeping keys out of readable instance
metadata are module business: gzip the payload, or stage it in a secret
store and pass only a small stub through bootstrap_data/user-data.
bootstrap_format
cloud-config or ignition, from the bootstrap Secret’s format key;
cloud-config when absent.
format is not part of the bootstrap contract, which specifies only
value
(bootstrap-config.md).
It is written by the kubeadm bootstrap provider
(kubeadmconfig_controller.go
writes value and format; enum cloud-config;ignition in
kubeadmconfig_types.go
Format). Other bootstrap providers need not write it. CAPRKE2 also
always writes both keys, with the same enum
(rke2config_controller.go;
see the RKE2ControlPlane guide). format
describes the decoded payload; with CAPRKE2 gzipUserData: true the
decoded bytes are gzip of that format.
failure_domain (input)
Machine.spec.failureDomain (1–256 chars;
machine_types.go).
The InfraMachine MUST be placed in this failure domain
(infra-machine.md
“InfraMachine: failure domain”).
kubernetes_version (input)
Machine.spec.version (optional, 1–256 chars;
machine_types.go).
The value may carry a control-plane-provider-specific distro suffix: with
RKE2 it is vX.Y.Z+rke2rN (for example v1.31.4+rke2r1), copied verbatim
from RKE2ControlPlane.spec.version to Machine.spec.version (see the
RKE2ControlPlane guide). The controller
passes this input to the module verbatim, unstripped; a module that
uses it for an image lookup or a semver comparison MUST strip the +…
build-metadata suffix itself (see the RKE2ControlPlane
guide).
control_plane (input)
true when the owning Machine carries the cluster.x-k8s.io/control-plane
label (MachineControlPlaneLabel,
machine_types.go;
CAPI’s own util.IsControlPlaneMachine checks only for this label’s
presence,
util.go).
There is no role string in MachineSpec.
Outputs
Common outputs apply (health).
| Name | Type | Required | Maps to |
|---|---|---|---|
provider_id | string or null | yes (may be null until known) | TerraformMachine.spec.providerID → Machine.spec.providerID (below) |
addresses | list(object({type=string, address=string})) | yes (may be []) | TerraformMachine.status.addresses → Machine.status.addresses (below) |
failure_domain | string or null | yes (may be null) | TerraformMachine.status.failureDomain → Machine.status.failureDomain (below) |
interruptible | bool or null | declared by every machine module; value optional (null = false) | TerraformMachine.status.interruptible (below) |
provider_id (output)
Maps to TerraformMachine.spec.providerID → Machine.spec.providerID. Must
equal the Node’s spec.providerID exactly (see “Node providerID matching”
below). 1–512 chars
(infra-machine.md
“InfraMachine: provider ID”; MinLength=1, MaxLength=512 markers in
machine_types.go).
"" is treated as null.
Written once; a later non-null change is
OutputsValid=False/ProviderIDChanged. If it turns null after
provisioning, the instance is treated as terminated (see
common.md “Out-of-band termination”) and
spec.providerID is kept. We latch it — never clear it once set — because
CAPI does not: the Machine controller recopies Machine.spec.providerID
from TerraformMachine.spec.providerID on every reconcile rather than
latching it once
(machine_controller_phases.go),
and clearing our side would blank the Machine’s field and stop
address/failure-domain copying from the InfraMachine on that same path.
addresses (output)
Maps to TerraformMachine.status.addresses → Machine.status.addresses
(infra-machine.md
“InfraMachine: addresses”). type is one of Hostname, ExternalIP,
InternalIP, ExternalDNS, InternalDNS (MachineAddressType enum,
common_types.go);
address is 1–256 chars; at most 256 entries (MachineAddress and
Machine.status.addresses markers,
common_types.go
and
machine_types.go).
The controller validates every mapped output against these CRD markers
before patching status; a violation is OutputsValid=False/OutputsInvalid,
and nothing is written.
Canonical order. Before writing status the controller sorts the
module’s list by type precedence InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname, then by address. CAPI itself copies the list
verbatim
(machine_controller_phases.go),
so an unsorted module output would otherwise reach Machine.status.addresses
in whatever order the module emitted it. Rationale: RKE2’s
registrationMethod: internal-first takes the first address in list
order whose type is InternalIP or ExternalIP (this is what the code
does, not “internal preferred” as its own docs claim);
internal-only-ips needs an InternalIP present, and external-only-ips
needs an ExternalIP present (see the RKE2ControlPlane
guide) — a fixed, sorted order makes that
first-match deterministic instead of depending on module authoring order.
failure_domain (output)
Maps to TerraformMachine.status.failureDomain → Machine.status.failureDomain
(the actual placement;
infra-machine.md
“InfraMachine: failure domain”). When the failure_domain input is
non-null the output MUST equal it: the contract says the InfraMachine MUST
be placed in the requested domain, and the controller enforces it — a
mismatch is OutputsValid=False/FailureDomainMismatch and Ready=False,
so KCP/MachineDeployment spreading can never be silently broken. The
output adds information only when no domain was requested.
interruptible (output)
Maps to TerraformMachine.status.interruptible (bool); true for
spot/preemptible instances. CAPI reads status.interruptible from the
InfraMachine and, when true, sets the Node label
cluster.x-k8s.io/interruptible (InterruptibleLabel;
machine_controller_noderef.go),
which termination handlers and schedulers key on. The output must be
declared because the generated root re-exports every contract output
unconditionally; a module without spot support declares
output "interruptible" { value = false }. Null is written as false.
Provisioned rule
status.initialization.provisioned = true when the state Secret carries
the inputs-hash annotation of a successful apply, provider_id is
non-null, and health.state != "pending"; derived from state only, never
from Job history, so it is rebuilt after clusterctl move (evaluated once
on the target, then latched).
Latched. The formula applies only until it first holds; once
provisioned is true it stays true for the object’s life, whatever later
health or outputs say (CAPI never flips infrastructureProvisioned back;
machine_controller_phases.go
“should not flip back”). A later null provider_id or non-running health
is reported on InfrastructureHealthy (for example InstanceTerminated),
never by un-provisioning.
While pending, the controller refreshes regardless of
drift.intervalSeconds, 30s after the last refresh and backing off to at
most 5m (not a field; see common.md). After a
successful apply it refreshes once more unless the apply’s own outputs are
valid with a health.state other than pending or unknown. Ready=True
additionally requires health.healthy and state == "running".
Ready timeline
What CAPI mirrors into Machine.status.conditions[InfrastructureReady],
which MachineHealthCheck can act on:
| Phase | Ready | Reason |
|---|---|---|
Waiting for owner Machine, cluster infrastructure, exports, bootstrap data | Unknown | WaitingFor* (on the DependenciesReady condition; Ready reason ReadyUnknown) |
| Apply Job started | False | Provisioning |
| Provisioned, healthy | True | Ready |
Provisioned, health not running/healthy | False | NotReady (detail on InfrastructureHealthy) |
| Deleting | False | Deleting |
Unknown rather than False while waiting matters for MHC:
unhealthyMachineConditions timeouts count from the condition’s
lastTransitionTime, and a reason change does not reset it. Workers
created alongside a slow control plane can wait a long time for bootstrap
data; starting the False window only at apply start means an MHC
InfrastructureReady=False timeout measures apply time, not control-plane
bring-up. After provisioning, Ready is computed from
InfrastructureHealthy and deletion only; drift results, drift-Job
failures, identity or RBAC problems never turn it False (they are
separate conditions), so a credentials outage cannot trigger fleet-wide
remediation.
Lifecycle
-
Immutable. Changes to
spec.sourceandspec.identityRefafter creation are rejected by the webhook, andspec.providerIDcan only be set once, by the controller. Those changes come via MachineDeployment/KCP rollout.spec.jobs,spec.driftandspec.remediationare operational policy and stay mutable: they change how the controller treats the machine (a Job deadline, drift checks, auto-remediation), never what was applied.captf_cluster_outputschanges do not trigger re-apply: the rendered inputs of the first apply (includingbootstrap_dataandcaptf_cluster_outputs) are stored in the object-ownedcaptf-inputs-<kindshort>-<name>Secret and re-fed unchanged on drift and destroy; the effective image and identity are pinned there too, so a later change toTerraformCluster.spec.defaults(for example a ClusterClass in-place update) never changes what runs against an existing machine. Bootstrap data is read once. The kubeadm bootstrap provider does not rewrite a Machine’s bootstrap Secret after creation — it only extends the join token’s TTL in the workload cluster until the Node joins (kubeadmconfig_controller.gorefreshBootstrapTokenIfNeeded) — so “read once” loses nothing. Pools are different (seemachinepool.md). -
Gating. Apply waits for: the owner
Machine(looked up through the Machine ownerRef,util.GetOwnerMachine, then the Cluster throughcluster.x-k8s.io/cluster-name; a TerraformMachine never has a Cluster ownerRef);Cluster.status.initialization.infrastructureProvisioned(cluster_types.go); the cluster’sexportsreadable (WaitingForClusterExports); and the bootstrap Secret (WaitingForBootstrapData; the contract workflow also exits whiledataSecretNameis nil, and""counts as nil). The owning Cluster’sspec.infrastructureRefmust be kindTerraformCluster; otherwiseDependenciesReady=False/ClusterNotTerraformand nothing is done. Once those pass, thespec.variablesFromsources must exist and be labeled (DependenciesReady=False/VariablesSourceNotFound) with valid keys and values (False/VariablesInvalid); seecommon.md“User variables”. While gated,Ready=Unknown(timeline above). -
Retry. A state Secret that exists without the inputs-hash annotation (a failed or interrupted first apply) means the apply is retried with backoff even though the spec is immutable; each attempt gets a distinct Job name (attempt counter).
-
Refresh/drift. On
spec.drift.intervalSeconds:apply -refresh-onlythenplan -detailed-exitcode. Drift is reported only (DriftDetected=True, a negative-polarity condition kept out ofReady); never auto-applied for machines (MachineDriftPolicyhas noactionfield: there is nothing to remediate against for an immutable machine). The refresh updatesaddressesandhealth. A failed drift Job setsDriftJobSucceeded=Falseand nothing else. -
Health → remediation.
healthoutsiderunning/healthyafter provisioning setsInfrastructureHealthy=False, thenReady=False/NotReady, mirrored toMachine.status.conditions[InfrastructureReady]=False, which aMachineHealthCheckcan act on. See Machine Remediation for configuringspec.remediation, the unhealthy threshold and sampling interval, and how a MachineHealthCheck reaches a replacement.healthis purely cloud-side: the controller does not cross-checkMachine.status.nodeRef/nodeInfo(Node existence,machine_controller_noderef.go); a missing or NotReady Node is MHC’snodeStartupTimeout/Node-condition checks’ job. -
Delete. Deletion of a TerraformMachine is normally initiated by the Machine controller after drain and pre-terminate hooks. Ordering is CAPI’s: drain (bounded by
Machine.spec.deletion.nodeDrainTimeoutSeconds) and volume detach (nodeVolumeDetachTimeoutSeconds) finish before the TerraformMachine is deleted, sodestroynever races a running drain; the Node object is deleted by core-machine only after the TerraformMachine is gone (bounded bynodeDeletionTimeoutSeconds;machine_types.go,machine_controller.go). No contract input carries these timeouts.The webhook allows a direct delete when the object has no Machine ownerRef at all (an orphan: KCP deletes the InfraMachine itself when KubeadmConfig or Machine creation failed,
helpers.go), when the owning Machine is itself deleting, or when the request carriesclusterctl.cluster.x-k8s.io/delete-for-move; it refuses deletion otherwise, so drain and the KCP etcd-membership hook are never bypassed by deleting the infra object directly.Then a
destroyJob runs, rendered from thecaptf-inputs-<kindshort>-<name>Secret (never from the bootstrap Secret or the Cluster, which may already be gone); on success it drops the state, the Lease and the inputs Secret, and removes the finalizer. If the owning Machine or Cluster is gone the same path runs. If there is no state Secret and no active Job (deleted while still gated), the finalizer is removed without a Job. If destroy fails, the object stays withApplyJobSucceeded=False/DestroyFailed(Ready=False) and retries with backoff; a Lease whose holder Job no longer exists is force-unlocked automatically. There is no skip-destroy annotation: a permanently failing destroy is resolved by cleaning up out of band and removing the finalizer by hand, after backing up or un-owning thetfstate-*Secrets, which are otherwise garbage-collected with the object (see the stuck-destroy runbook). A stuck control-plane machine destroy blocks KCP remediation, scale and upgrade until resolved (remediation.go). -
The module MUST accept
bootstrap_dataopaquely: it is base64 of the bootstrap payload (Inputs), passed to a base64-taking user-data argument orbase64decode()d into a plain one; the module MUST NOT need to parse the payload. -
Taints and labels.
Machine.spec.taintsand Machine labels are applied to the Node by core-machine (machine_types.go;machine_controller_noderef.go), not by the infrastructure; the machine role has nonode_labelsinput (pools do, seemachinepool.md); no role has a taints input. -
In-place updates (the CAPI
InPlaceUpdatesfeature gate, alpha) are unsupported: no field is mutable, and the webhook rejects the MachineSet’s SSA patch; the rollout path is a new Machine.
Node providerID matching (module authors)
CAPI links a Machine to its Node only when Node.spec.providerID equals
Machine.spec.providerID exactly
(machine_controller_noderef.go).
Our provider_id output is copied to the Machine verbatim, so the module
MUST emit the same string the Node will carry:
- With a cloud-controller-manager (CCM): the CCM sets
Node.spec.providerIDin its own format (for exampleaws:///<zone>/<instance-id>,azure:///subscriptions/...,openstack:///<uuid>,vsphere://<uuid>). Emit exactly that format. - Without a CCM: the kubelet must be started with
--provider-id=<value>(viaKubeadmConfignodeRegistration.kubeletExtraArgs), and the value must be derivable on the instance at boot (an instance-metadata lookup, or a per-machine value baked into the module’s user-data wrapper), becausebootstrap_datais shared per template and opaque to the module. - Not supported in v1: a controller-side patch of
Node.spec.providerIDonce the Node registers, as CAPD does (machine.go), would need a workload-cluster client and a manager flag; the controller does not implement it. In v1 the CCM or kubelet--provider-idmust set it. - RKE2: CAPRKE2 itself never sets kubelet
provider-id,node-ipornode-name(the fields exist in its config but nothing in the repo writes them; see the RKE2ControlPlane guide). Without a CCM, the image MUST setagentConfig.kubelet.extraArgs: [provider-id=<same format>]from instance metadata so the kubelet-registered value matches ourprovider_idoutput exactly. This matters more under RKE2 than KCP: every RKE2 join (control plane or worker) requiresRCP.status.availableServerIPsto be non-empty, which requires at least one Ready control-plane Machine, and a Machine only becomes Ready once its Node exists with a matchingproviderID— so a providerID mismatch on the first control-plane Machine blocks every subsequent join, not just that one Machine’s own readiness.
Autoscale-from-zero (InfraMachineTemplate.status.capacity / nodeInfo;
infra-machine.md
“InfraMachineTemplate: support cluster autoscaling from zero”) is supported
through the image: instance size is fixed inside the module, so the
image declares it with the OCI labels io.captf.capacity and
io.captf.node-info (see image-contract.md
“OCI labels”), and a TerraformMachineTemplate reconciler copies them into
status.capacity/status.nodeInfo. An image without the labels leaves
both fields unset; the Cluster Autoscaler’s
capacity.cluster-autoscaler.kubernetes.io/* annotations on the
MachineDeployment/MachineSet remain the fallback.
Control-plane machines
When control_plane = true, the module is responsible for registering the
instance in the control-plane load balancer’s backend or target group. The
ids it needs (load balancer, target group, backend pool — whatever the
cluster module’s cloud exposes) come from captf_cluster_outputs (see
cluster.md exports, which describes the
publishing side of this same convention). Registration is done in the
machine module’s own Terraform state, not the cluster module’s, so
destroy on the machine deregisters it as an ordinary part of tearing down
that state
(infra-machine.md).
Ordering matters: the instance MUST be in the load balancer backend before
kubeadm init/kubeadm join finish on it, because the contract requires
the endpoint to be reachable through the load balancer during
control-plane bring-up
(status.go;
see the KubeadmControlPlane guide on
reachability) — Cluster.status.initialization.controlPlaneInitialized
never latches if the first control-plane node can’t be reached at the
endpoint it just joined. In practice this means the module’s apply must
register-then-boot (or register-then-poll-healthy) rather than
boot-then-register as an afterthought.
For worker Machines (control_plane = false) no load balancer
registration applies; captf_cluster_outputs is still read for any other
cluster-level values the module needs.
Minimal skeleton
A block written on one line may hold at most one argument
(OneLineBlock;
tofu validate reports “Invalid single-argument block definition” on a
violation), so a variable that needs both type and sensitive, or
type and default, must use the multi-line block form below.
variable "captf_contract" { type = string }
variable "captf_cluster" { type = object({ name = string, namespace = string }) }
variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) }
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" { type = map(string) }
variable "machine_name" { type = string }
variable "bootstrap_data" { # base64 of the bootstrap Secret's `value`; e.g. user_data_base64 = var.bootstrap_data
type = string
sensitive = true
}
variable "bootstrap_format" { type = string }
variable "failure_domain" {
type = string
default = null
}
variable "kubernetes_version" {
type = string
default = null
}
variable "control_plane" { type = bool }
# Tag every cloud resource this module creates with captf_tags (common.md
# "captf_tags"); this stub only has to reference it, not create anything.
resource "terraform_data" "tags" { input = var.captf_tags }
output "provider_id" { value = "noop:///${var.captf_object.namespace}/${var.captf_object.name}" }
output "addresses" { value = [{ type = "InternalIP", address = "10.0.0.1" }] }
output "failure_domain" { value = var.failure_domain }
output "interruptible" { value = false }
output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } }
MachinePool Role
Implements the CAPI InfraMachinePool contract
(infra-machinepool.md;
MachinePool types in
machinepool_types.go)
for a TerraformMachinePool. One workspace/state per
TerraformMachinePool. The module manages a native scaling group (an
ASG, VMSS, MIG, instance pool, or similar) and reports the group id, the
list of provider IDs currently in it, and the desired capacity. Mutable:
re-applied on spec or replica changes.
The v1 controller implements fixed replicas and native autoscaling
(autoscaling.enabled). It does not watch the bootstrap Secret: a rotated
token reaches the pool at its next reconcile, at the latest one
membership-refresh interval later (see “Bootstrap rotation” below). Drift
relies entirely on the module’s own ignore_changes: the drift Job already
runs apply -refresh-only then plan -detailed-exitcode for every kind,
and the controller does no plan-JSON filtering of the result — an
autoscaled pool’s module MUST put its desired-count attribute under
lifecycle { ignore_changes = [...] } (see Lifecycle and the
autoscaling input) or a cloud-side scale reports as drift.
MachinePool Machines
(status.infrastructureMachineKind; optional,
infra-machinepool.md
“MachinePoolMachines support”) are out of scope. Two consequences follow:
there is no drain on scale-down, replacement or version roll — drain
exists only for Machines
(infra-machinepool.md
“MachinePoolMachines support … draining a node before scale down”), so
modules SHOULD use native lifecycle hooks or termination handlers instead —
and MachineHealthCheck never selects pool instances.
Inputs
Common inputs apply (common.md), including
captf_cluster_outputs.
| Name | Type | Required | Value source |
|---|---|---|---|
machinepool_name | string | yes | Owning CAPI MachinePool.metadata.name |
replicas | number | yes | The desired capacity of the group (below) |
bootstrap_data | string, sensitive | yes | Base64 of the bootstrap Secret’s value key (below) |
bootstrap_format | string | yes | as machine role; the format key is kubeadm-specific, not contract (see machine.md) |
failure_domains | list(string) | yes (may be []) | MachinePool.spec.failureDomains (below) |
cluster_failure_domains | list(string) | yes (may be []) | Names from the cluster’s own state (below) |
kubernetes_version | string or null | yes (may be null) | MachinePool.spec.template.spec.version (below) |
node_labels | map(string) | yes (may be {}) | MachinePool.spec.template.metadata.labels verbatim (below) |
autoscaling | object({enabled=bool, min=number, max=number}) | yes | Parsed from the MachinePool autoscaler annotations (below) |
replicas (input)
The desired capacity of the group.
- With
autoscaling.enabled = false:MachinePool.spec.replicas, authoritative and in the inputs hash (*int32, “Defaults to 1”; the pointer distinguishes an explicit 0 from unset, and 0 is allowed). - With
autoscaling.enabled = true:TerraformMachinePool.status.replicas(the observed desired capacity from the last refresh) once one exists, elseMachinePool.spec.replicasfor the first apply before any refresh has run — either way clamped into[autoscaling.min, autoscaling.max]so a render never asks the cloud for an out-of-range desired count. The clamp applies only to this rendered input; the write-back (Lifecycle) always writes the raw observed value, unclamped.
replicas is excluded from the inputs hash while autoscaling.enabled, so
a non-replica-triggered apply (bootstrap rotation, module change, exports
change) never fights the native autoscaler, and an observed-count change
alone never triggers InputsChanged; a fixed-replica pool hashes
replicas exactly as before.
bootstrap_data (input)
Base64 of the raw bytes of the value key of the bootstrap Secret
named by MachinePool.spec.template.spec.bootstrap.dataSecretName
(template is MachineTemplateSpec; see
bootstrap-config.md),
encoded by the controller for every bootstrap provider exactly as for the
machine role (see machine.md
bootstrap_data: CAPRKE2 gzipUserData: true makes value raw gzip
bytes that cannot be a string).
Modules pass it to a base64-taking launch-configuration argument (for
example aws_launch_template.user_data, which expects base64) or
base64decode() it for a plain-text argument (UTF-8 payloads only).
Hashed by content; dataSecretName: "" is treated as nil
(WaitingForBootstrapData).
Rotates. For MachinePools the kubeadm bootstrap provider re-creates the
join token when it is past half its TTL (default TTL 15m, checked every
TTL/3) and rewrites the same Secret in place, roughly every 7.5 minutes
(token.go
refreshBootstrapTokenIfNeeded/recreateBootstrapToken;
kubeadmconfig_controller.go storeBootstrapData). See Lifecycle for what
that requires from modules.
failure_domains (input)
MachinePool.spec.failureDomains (at most 100 items, each 1–256 chars).
[] means the module chooses. To make that choice informed, the cluster’s
failure-domain names are injected as cluster_failure_domains (below); the
full attributes maps are only available through exports.
cluster_failure_domains (input)
Names read from the cluster’s own state, off the failure_domains output
(see cluster.md) at render time — not from
TerraformCluster.status.failureDomains — so a pool module can spread
across all of them when failure_domains is [] without depending on
exports conventions. Sourcing it from state rather than status keeps it
stable across clusterctl move (status is not moved, and is empty until
the target’s first cluster reconcile); it stays in the inputs hash.
kubernetes_version (input)
MachinePool.spec.template.spec.version. A change MUST roll the instances
(see Lifecycle). As with the machine role, the value may carry a
control-plane-provider-specific distro suffix (RKE2: vX.Y.Z+rke2rN, for
example v1.31.4+rke2r1); the controller passes it to the module
verbatim, and a module that needs it for an image lookup or a semver
comparison MUST strip the +… suffix itself (see
machine.md and the RKE2ControlPlane
guide).
node_labels (input)
MachinePool.spec.template.metadata.labels verbatim (absent maps to
{}). Pool instances have no Machine objects, so core CAPI never syncs
labels onto their Nodes; the module MUST render these into kubelet
registration (--node-labels) itself. node_labels is in the inputs
hash, so an edit re-applies: new members pick it up, and existing members
keep their registration.
The precise requirement and why it is not simple. bootstrap_data is
opaque (see machine.md “the module MUST NOT
need to parse the payload”) and, depending on the bootstrap provider and
its settings, may be plain cloud-config, plain Ignition, or gzip of
either. The module cannot edit the kubelet flags inside that payload
without parsing it, which the contract forbids. The realistic options,
in order of how much of bootstrap_data they need to understand:
-
Cloud-config, uncompressed (
bootstrap_format == "cloud-config", CAPRKE2gzipUserDataunset orfalse): wrapbootstrap_dataand a second, module-generated part in amultipart/mixedMIME message (for example Terraform’scloudinit_configdata source, or an equivalent built by hand) so cloud-init runs both parts at boot. The second part writes a kubelet drop-in with--node-labels:data "cloudinit_config" "node" { gzip = false base64_encode = true part { content_type = "text/cloud-config" # bootstrap_data is base64 of the raw payload (machine.md); decode it # only because this part must stay unmodified UTF-8 cloud-config, not # to parse or edit it. content = base64decode(var.bootstrap_data) } part { content_type = "text/cloud-config" content = yamlencode({ write_files = [{ path = "/etc/systemd/system/kubelet.service.d/20-node-labels.conf" content = "[Service]\nEnvironment=\"KUBELET_EXTRA_ARGS=--node-labels=${join(",", [for k, v in var.node_labels : "${k}=${v}"])}\"\n" }] }) } } # data.cloudinit_config.node.rendered is base64 multipart/mixed; pass it # to the same user-data argument bootstrap_data alone would have gone to.RKE2’s agent config has the same shape through a second part that drops a file under
/etc/rancher/rke2/config.yaml.d/*.yamlwith anode-label:list, merged by the RKE2 agent at boot alongside the bootstrap part’s own/etc/rancher/rke2/config.yaml. -
Ignition (
bootstrap_format == "ignition") or a gzipped payload (CAPRKE2gzipUserData: true): neither is multipart-mixable the same way — Ignition has its own merge/append config mechanism, and a gzipped payload must be decompressed, edited or merged, and recompressed (or passed through a launch-configuration field that decompresses it itself). These need format-aware handling specific to the format; there is no one wrapper that covers them. -
A module MAY instead declare, in its own documentation, only the
bootstrap_format/compression combinations it supports, and fail loudly (a precondition, or an unsupported-value error) on the others, rather than silently droppingnode_labels.
Labels the kubelet may not self-assign (the node-role.kubernetes.io/*
and other restricted kubernetes.io/k8s.io prefixes under
NodeRestriction) are the module’s to filter.
There is no taints input: the v1.14.2 MachinePool webhook rejects any
spec.template.spec.taints (“taints feature for MachinePools is not yet
implemented”,
machinepool.go;
MachineTaintPropagation is off by default,
feature.go).
autoscaling (input)
Parsed from the MachinePool annotations
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size and
-max-size
(AutoscalerMinSizeAnnotation/AutoscalerMaxSizeAnnotation,
common_types.go).
This is the pool’s only autoscaling switch; there is no separate drift
flag.
enabled=true with both min/max parsed if and only if both annotations
are present and 0 ≤ min ≤ max; otherwise {enabled=false, min=0, max=0}
(no annotations sets AutoscalingActive=False/AutoscalingDisabled), plus,
when an annotation is present but the pair is incomplete, unparsable or
min > max, the condition AutoscalingActive=False/AutoscalingAnnotationsInvalid
(the pool still applies without autoscaling; AutoscalingActive never
feeds Ready).
How this differs from CAPI’s own reading of the same annotations. CAPI
already reads these annotations on MachinePools, but only in a narrower
path than ours: the MachinePool admission webhook defaults and clamps
spec.replicas from them only when spec.replicas is nil (an explicit
value, including 0, is left alone), requires both annotations to be
present, rejects the create/update outright if either annotation is
unparsable, and does not reject min > max — our controller does,
via AutoscalingAnnotationsInvalid
(machinepool.go,
calculateMachinePoolReplicas, called from Default) — which is what
makes them a sound source for the initial size. Our controller does not
rely on the webhook: it reads the annotations itself on every reconcile
for autoscaling.min/.max, and treats a missing or invalid annotation as
enabled=false plus AutoscalingActive=False/AutoscalingAnnotationsInvalid
rather than rejecting anything.
enabled means the module owns the desired count and its scaling
policy: it sets the group’s min/max to these values, configures whatever
native scaling policy it wants (target tracking, scheduled, and so on),
and MUST put the group’s desired-count attribute under
lifecycle { ignore_changes = [...] } so an apply never resets what the
cloud autoscaler decided.
On the CAPI side the controller then claims
cluster.x-k8s.io/replicas-managed-by: captf on the MachinePool when the
annotation is absent or "false" (ReplicasManagedByAnnotation; CAPI
counts the annotation as set for any value except the literal string
"false", hasTruthyAnnotationValue,
helpers.go,
called by ReplicasManagedByExternalAutoscaler — a foreign truthy value is
left alone, never overwritten), and, on every reconcile pass on which
status.replicas is known and differs from MachinePool.spec.replicas,
writes the raw observed desired capacity back to MachinePool.spec.replicas.
The InfraMachinePool docs make this the provider’s job
(machine-pool.md
“It is the provider’s responsibility to update Cluster API’s Spec.Replicas
property to the value observed”); see Lifecycle “Write-back” for the exact
rules. With the annotation set, the MachinePool controller reports phase
Scaling instead of ScalingUp/ScalingDown
(machinepool_controller_phases.go).
The Kubernetes Cluster Autoscaler does not act on these pools: its
clusterapi provider requires MachinePool Machines
(README.md),
which this role excludes. Running it against a CAPTF pool anyway is
unsupported, because its spec.replicas patches would be overwritten by
the write-back.
Outputs
Common outputs apply (health).
| Name | Type | Required | Maps to |
|---|---|---|---|
provider_id | string or null | yes (may be null) | TerraformMachinePool.spec.providerID, the scaling-group id (below) |
provider_id_list | list(string) | yes (may be []) | TerraformMachinePool.spec.providerIDList → MachinePool.spec.providerIDList (below) |
replicas | number | yes | The group’s observed desired capacity (below) |
instances | list(object({provider_id=string, instance_id=optional(string), addresses=optional(list(object({type=string, address=string})), []), failure_domain=optional(string), state=optional(string)})) | yes (may be []) | TerraformMachinePool.status.instances (below) |
provider_id (output)
Maps to TerraformMachinePool.spec.providerID — the scaling-group id;
optional in the contract and not used by core CAPI, 1–512 chars
(infra-machinepool.md
“InfraMachinePool: providerID”). May stay null for group-less
implementations.
provider_id_list (output)
Maps to TerraformMachinePool.spec.providerIDList →
MachinePool.spec.providerIDList (mandatory; at most 10000 items, each
1–512 chars;
infra-machinepool.md
“InfraMachinePool: providerIDList”). Rules:
-
It MUST contain every non-terminated member of the group regardless of health — pending, starting, standby, rebooting or unhealthy instances included — because CAPI deletes the Node of any providerID that leaves the list (
machinepool_controller_noderef.godeleteRetiredNodes; only an empty list withstatus.replicas != 0is guarded,machinepool_controller_phases.go). Listing only “running” or “healthy” instances kills live Nodes. -
Each entry MUST equal the Node’s
spec.providerIDexactly (machinepool_types.goproviderIDList“must match the provider IDs as seen on the node objects”; seemachine.md“Node providerID matching”), otherwise Nodes never get a nodeRef, keep thenode.cluster.x-k8s.io/uninitialized:NoScheduletaint that the kubeadm bootstrap provider applies, and are eventually deleted. -
Order is irrelevant to CAPI: the controller sorts and de-duplicates before writing, since CAPI compares the list with
reflect.DeepEqualand any reorder churns status (CAPD sorts too).tfcapi-lintstill warns (output/provider-id-list-shape) when the output’s own expression is not wrapped insort()ordistinct(), because a module is applied and refreshed independently of the controller’s write path: an unsorted expression makesprovider_id_listreorder between identicalapply -refresh-onlyruns, which is a plan diff (and, for a module whose desired-count is notignore_changesd, spurious drift) even though the controller’s own write is stable. Wrap the output’svalueexpression itself — the check inspects that expression directly, not a local it reads from:output "provider_id_list" { value = sort(local.raw_ids) # or distinct(...), or both }value = local.sorted_idsdoes not satisfy the check even whensorted_idsis itselfsort(...)-derived, and JSON-syntax modules are not scanned at all. Eithersort()ordistinct()alone satisfies it; the controller sorts and de-duplicates regardless of which the module uses.
replicas (output)
The group’s desired capacity as observed at refresh, not a count of
running instances. Maps to TerraformMachinePool.status.replicas →
MachinePool.status.replicas
(infra-machinepool.md
“InfraMachinePool: replicas”; read in
machinepool_controller_phases.go).
Without MachinePool Machines, CAPI sets the MachinePool’s
readyReplicas/availableReplicas equal to this value
(machinepool_controller_status.go);
only the deprecated v1beta1 readyReplicas is Node-based.
Outside a scaling transition it MUST equal length(provider_id_list); the
controller derives status.replicas from this output and serializes 0
explicitly (a missing value would make CAPI keep the old count and block
scale-to-zero forever,
machinepool_controller_phases.go).
instances (output)
Maps to TerraformMachinePool.status.instances — optional, provider-defined
shape, not used by core CAPI
(infra-machinepool.md
“InfraMachinePool: instances”; the contract’s own example shape is
{addresses, instanceName, providerID, version, ready}, ours differs,
which is allowed).
One Go type, MachinePoolInstance{ProviderID, InstanceID, Addresses []MachineAddress, FailureDomain, State}, maps field by field:
provider_id→providerID, instance_id→instanceID,
addresses→addresses (validated like the machine role’s),
failure_domain→failureDomain, state→state (the health.state
enum).
Size. providerIDList alone may reach 10000×512 bytes, and an object
must stay under the etcd request limit of roughly 1.5 MiB, so the
controller caps status.instances at 1000 entries (reported as
OutputsValid=True/InstancesTruncated when it does), and the documented
practical pool size is at most 2000 members.
Provisioned rule
status.initialization.provisioned = true — and, for as long as CAPI
v1.14 reads it, status.ready = true; the MachinePool controller decides
provisioning solely from status.ready
(machinepool_controller_phases.go
via external.IsReady, and CAPD still sets it) — when the state Secret
carries the inputs-hash annotation of a successful apply and
health.state != "pending". provider_id is not required, since the
contract makes it optional. provider_id_list may legitimately be empty
when replicas == 0, and a 0-member group reports running/healthy (see
common.md).
Latched. The formula applies only until it first holds; once true,
provisioned and status.ready stay true for the object’s life, whatever
later health or outputs say. It is derived from state only, so it is
rebuilt after clusterctl move (the formula is re-evaluated once on the
target, then latched again); CAPI does not latch its copy, so the
MachinePool shows infrastructureProvisioned=false briefly after a move
until the first reconcile on the target.
Separately, CAPI copies TerraformMachinePool.spec.providerIDList/status.replicas
to the MachinePool only after ClusterCache.GetClient succeeds — that is,
only once the workload cluster is reachable
(machinepool_controller_phases.go);
until then the TerraformMachinePool side is correct but the MachinePool
does not reflect it.
Lifecycle
-
Mutable. Re-apply on: a
spec.sourcechange (image, pull policy); areplicaschange on the MachinePool (unlessautoscaling.enabled, wherereplicasis excluded from the inputs hash and a user edit ofMachinePool.spec.replicasis overwritten by the write-back, surfaced as an event); afailure_domainschange; anode_labelschange (applied likebootstrap_data: a launch-configuration update, no instance replacement); akubernetes_versionchange; abootstrap_datachange; or acaptf_cluster_outputscontent change. Every apply, including a driftRemediate, re-renders the current inputs; nothing is replayed from an older run. -
Bootstrap rotation. Because the kubeadm bootstrap provider rewrites the pool’s bootstrap Secret roughly every 7.5 minutes (
bootstrap_dataabove), a pool is re-applied at that cadence for the life of the pool. The controller does not watch the bootstrap Secret — it is notcaptf.io/managed, so it is not in the cache — so a rotated token reaches the pool at its next reconcile, at the latest one membership-refresh interval later, and is then coalesced into the next apply. Module rules that make this cheap and safe:- a
bootstrap_datachange MUST be applied as an in-place update of the group’s launch configuration (a launch template version, instance template, or VMSS model) so that new members get the current token, and MUST NOT replace existing instances; - a
kubernetes_versionchange MUST roll (replace) the instances, because a ClusterClass upgrade completes only when the Nodes’ kubelet versions match (upgrade.goIsMachinePoolUpgrading; CAPD’s DevMachinePool rolls on version only and never on bootstrap data,dockermachinepool_backend.go); - modules SHOULD expose no other trigger that replaces instances on an
apply with unchanged
kubernetes_version.
Each rotation re-applies the pool, so a module that never converges (see below) is applied roughly every 7.5 minutes indefinitely, on top of its 30s refresh cadence.
- a
-
Membership refresh.
provider_id_list,replicas,instancesandhealthare only as fresh as the last apply or refresh, and a Node joining the group is unschedulable until its providerID appears inMachinePool.spec.providerIDList— CAPI removes the kubeadm-appliednode.cluster.x-k8s.io/uninitialized:NoScheduletaint only then (machinepool_controller_noderef.go). The controller therefore runsapply -refresh-onlyon its own membership-refresh interval,spec.membershipRefreshIntervalSeconds(default 60, 15–86400, may not be 0), and additionally right after every apply and repeatedly (every 30s) whilehealth.state == "pending"orlength(provider_id_list) != replicas— a module that never converges is therefore refreshed every 30s indefinitely.drift.intervalSeconds: 0is rejected by the CRD schema (minimum 1) for a pool’s own field, by contrast with a machine’s or the cluster’sdrift.intervalSeconds, where 0 is a valid setting that disables drift; an inheriteddefaults.drift.intervalSeconds: 0(unset) falls back to the controller’s default drift interval (--drift-default-interval, 30m) rather than to 0. Drift may be sparse but membership never is. -
Drift order. Drift for a pool is the same job every other kind runs —
apply -refresh-onlyfirst, thenplan -detailed-exitcode(internal/runner/plan.go) — and the controller does no plan-JSON filtering of the result: whatever the refreshed plan reports is drift. This is why the desired-countlifecycle { ignore_changes = [...] }in theautoscalinginput is load-bearing, not optional: withautoscaling.enabled, the module owns the desired count, and a module that does notignore_changesit reports every cloud-side scale as drift. WithRemediatea detected diff re-applies the pool; withautoscaling.enabled = falsethat scales the group back toMachinePool.spec.replicas— the only mode where the module doesn’t alreadyignore_changesthe desired count. -
Write-back (
autoscaling.enabled). On every reconcile pass, one patch (SyncReplicas): the controller claimscluster.x-k8s.io/replicas-managed-by: captfon the MachinePool when the annotation is absent or"false"(a foreign truthy value already means another controller manages it, and is left untouched), and whenstatus.replicasis known and differs fromMachinePool.spec.replicas, writes the raw observed value (unclamped by[min,max], unlike the renderedreplicasinput) tospec.replicasin the same patch. AReplicasWrittenBackNormal event is emitted on the TerraformMachinePool whenever a write happens. Withautoscaling.enabled = false, the controller removesreplicas-managed-byonly if it still carriescaptf(never a foreign value), and never writesspec.replicas. Without the write-back, CAPI would reportScalingUp/ScalingDownforever andupToDateReplicaswould follow the stale spec (machinepool_controller_status.go). A ClusterClass-managed MachinePool MUST leavereplicasunset in the topology for this to work — the topology controller would otherwise reassert it every reconcile, fighting the write-back.Module note. Because
lifecycle { ignore_changes }cannot be conditional on a variable, a module that supports bothautoscaling.enabled = trueand= falsefrom the same scaling-group resource should, when disabled, pin the group’smin_size/max_sizetovar.replicas— not tovar.autoscaling.min/.max, which are0when disabled — so the sameignore_changesblock still lets the group trackreplicasexactly. -
Health.
healthfeeds conditions only; there is no remediation for pools in v1 (MHC selects Machines, and pool replicas only have Machines under MachinePool Machines, which this role excludes;infra-machinepool.md“MachinePoolMachines support”). Instances terminated by the cloud simply leaveprovider_id_listat the next refresh, and CAPI deletes their Nodes. -
Delete. A
destroyJob renders from the object-ownedcaptf-inputs-<kindshort>-<name>Secret, not from the bootstrap Secret: CAPI deletes the bootstrap config and the InfraMachinePool in the same pass (machinepool_controller_phases.goreconcileDeleteExternal), so the bootstrap Secret is usually already gone when destroy runs. Then the controller drops state, the Lease and the inputs Secret, and removes the finalizer.Node cleanup is partial. On MachinePool deletion, CAPI deletes only the Nodes whose providerIDs are absent from the last-seen
providerIDList(reconcileDeleteNodes→deleteRetiredNodes,machinepool_controller_noderef.go); Nodes still listed survive and are removed by the cloud-controller-manager once their instances are gone, or remain as orphans without a CCM. A module MAY emptyprovider_id_liston the final refresh before destroy, but the controller does not depend on it.
Minimal skeleton
A block written on one line may hold at most one argument
(OneLineBlock;
tofu validate reports “Invalid single-argument block definition” on a
violation), so a variable that needs both type and sensitive, or
type and default, must use the multi-line block form below.
variable "captf_contract" { type = string }
variable "captf_cluster" { type = object({ name = string, namespace = string }) }
variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) }
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" { type = map(string) }
variable "machinepool_name" { type = string }
variable "replicas" { type = number }
variable "bootstrap_data" { # base64 of the bootstrap Secret's `value`; e.g. launch template user_data = var.bootstrap_data
type = string
sensitive = true
}
variable "bootstrap_format" { type = string }
variable "node_labels" { type = map(string) }
variable "failure_domains" { type = list(string) } # always set, may be []; no default (input/default)
variable "cluster_failure_domains" { type = list(string) } # always set, may be []; no default (input/default)
variable "kubernetes_version" {
type = string
default = null
}
variable "autoscaling" { type = object({ enabled = bool, min = number, max = number }) }
# Tag every cloud resource this module creates with captf_tags (common.md
# "captf_tags"); this stub only has to reference it, not create anything.
resource "terraform_data" "tags" { input = var.captf_tags }
locals {
# Unsorted; provider_id_list must wrap sort()/distinct() directly in its
# own output expression: see "provider_id_list order" below.
raw_ids = [for i in range(var.replicas) : "noop:///${var.captf_object.namespace}/${var.captf_object.name}/${i}"]
}
output "provider_id" { value = "noop-group:///${var.captf_object.namespace}/${var.captf_object.name}" }
output "provider_id_list" { value = sort(local.raw_ids) }
output "replicas" { value = var.replicas }
output "instances" { value = [for id in sort(local.raw_ids) : { provider_id = id, state = "running" }] }
output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } }
A real module with autoscaling.enabled sets the group’s
min_size/max_size from var.autoscaling, its desired_capacity from
var.replicas, and declares lifecycle { ignore_changes = [desired_capacity] }
on the group resource; output "replicas" then reads the group’s
current desired capacity from the resource (refreshed by
apply -refresh-only), not var.replicas.
Per-instance state
instances[*].state uses the health.state enum (see
common.md) so per-instance conditions can be
derived later.
Deriving group health from mixed instance states. The controller does
not aggregate instances[*].state into health itself: health is copied
to the InfrastructureHealthy condition, and instances is copied to
status.instances (capped, see “Size” above), with no cross-referencing
between the two — a module that reports a healthy health and a list full
of degraded instances gets a healthy InfrastructureHealthy and a
degraded-looking status.instances, since nothing else reconciles the
difference. A pool module SHOULD derive health from the instances it is
about to report:
healthy = true,state = "running"when the group has reached its desired capacity and every counted instance’s own state isrunning;healthy = false,statereflecting the worst instance (for exampledegradedif any instance isdegraded, elsestopped, and so on down the enum) withreasonsnaming the affected instances, when some instances are notrunning;state = "pending"only while the group itself does not yet exist or has no members to report at all — not merely while some members are still starting, sincelength(provider_id_list) != replicasalready forces the controller’s fast (30s) refresh loop on its own (see “Membership refresh” below), so a module does not needhealth.state = "pending"to get that cadence during a scale-up or scale-down.
This is a SHOULD, not a MUST: a module that cannot observe per-instance
health may still report {state="running", healthy=true} as long as the
group itself is healthy (see common.md).
Changelog
Changes to the module contract, newest first. Within v1alpha* there is no
compatibility guarantee (see README.md
“Versioning”); every change is still recorded here.
Unreleased
Added
- Machinepool role, including native autoscaling (additive
within v1alpha1;
machinepool.md,README.md“Roles and CAPI mapping”): the v1 controller now reconcilesTerraformMachinePool. New API kindsTerraformMachinePoolandTerraformMachinePoolTemplate, and new fieldTerraformMachinePoolSpec.membershipRefreshIntervalSeconds(default 60, 15–86400, rejected below 15 by the CRD schema, not the webhook). NewOutputsValidreasonInstancesTruncated, set whenstatus.instancesis capped at 1000 entries. Manager flag--terraformmachinepool-concurrency(default 10). The controller does not watch the bootstrap Secret (a rotated token reaches the pool at its next reconcile, at the latest one membership-refresh interval later) and does not support MachinePool Machines. - Machinepool autoscaling (additive within v1alpha1;
machinepool.md“Inputs”autoscaling/replicas, “Lifecycle” “Drift order” and “Write-back”): theautoscalinginput is parsed from the MachinePool’scluster-api-autoscaler-node-group-min-size/-max-sizeannotations; when enabled,replicasrendersstatus.replicas(falling back toMachinePool.spec.replicasbefore the first observation) clamped into[min,max], is excluded from the inputs hash, and the controller writes the raw observed value back toMachinePool.spec.replicasand claimscluster.x-k8s.io/replicas-managed-by: captf(never a foreign truthy value), removing its own claim when autoscaling is disabled again. New conditionAutoscalingActive(informational, pool-only, never inReady; reasonsReplicasManagedByModule,AutoscalingDisabled,AutoscalingAnnotationsInvalid). New eventReplicasWrittenBack. New RBAC:patchoncluster.x-k8s.iomachinepools(write-back needs it;get/list/watchalready covered reads). Drift for a pool already ranapply -refresh-onlythenplanlike every other kind; nothing changed there — it is the module’slifecycle { ignore_changes }on the desired-count attribute that keeps a cloud-side scale from reporting as drift, and that responsibility is now load-bearing rather than theoretical. Upgrade behavior: a MachinePool that already carries both autoscaler annotations before the upgrade starts autoscaled and re-applies once right after it — itsautoscalinginput flips from{enabled=false,min=0,max=0}to the parsed, enabled value, which is hashed, so the inputs hash changes and the pool re-applies. - Plan preview and approval (the module contract is unchanged; see
Plan Approval):
TerraformCluster and TerraformClusterTemplate gain
spec.applyPolicy(Automatic, the default applied at reconcile, orManual; mutable, not hashed). UnderManualevery apply but the first (no state yet), including a drift remediation and a retry, first runs a new operation,plan(added to theOperationenum): a Job that runsinit,validate,plan -outandshow -jsonunder the run lease only and applies nothing. Its plan (counts, up to 50"<address> (<action>)"entries, never values, and a plan hashp1:<sha256>of the sortedaddress|actionsof every change) goes to the newstatus.plan(inputsHash,job,planHash,add,change,destroy,resources,truncated,createdAt), and the apply waits until the newcaptf.io/approve-planannotation names that hash. The approved apply (runner flag--expect-plan) plans again and applies only an identical plan; otherwise it stops with the newRunErrorKindplan-changedand the new plan waits for its own approval. A plan without changes never waits. Approving a plan also approves its deletes; the destructive-plan annotation is not needed on top. The annotation is removed andstatus.plancleared once the approved apply succeeds. NewApplyJobSucceededUnknown reasonsPlanAwaitingApprovalandPlanChanged; eventsPlanReady,PlanApproved,PlanApplied,PlanChanged; metriccaptf_plan_approvals_total{kind,result}(approved,changed) andcaptf_jobs_totalresultplan_changed; Job annotationscaptf.io/approved-plan,captf.io/plan-changedandcaptf.io/plan-unreadable. - State backups and restore (the module contract is unchanged; see
Terraform State and
State Restore):
every new state serial the manager
reads is copied verbatim into
captf-state-backup-<suffix>-<serial>Secrets (plus-part-Nper chunk), owned by the Terraform* object and labeled forclusterctl move; the newest--state-backups(default 5, 0 disables) are kept. TerraformCluster and TerraformMachine gainstatus.stateBackups(up to 16{serial, takenAt, bytes}). Thecaptf.io/restore-state: "<serial>"annotation starts a new operation,restore(added to theOperationenum ofstatus.activeJobandstatus.lastRun): a Job that runsinit,state push -forceof the backup andstate list, under the run lease (and the cluster write lease for a TerraformCluster). It precedes apply, drift and refresh but not deletion; the annotation is removed on success, and a failure is not retried for the same serial. New conditionRestoreJobSucceeded(never inReady) with reasonsStateRestored,RestoreFailed,RestoreBackupNotFoundand the three lease waits; eventsStateBackedUp,StateRestored,StateRestoreFailed; metricscaptf_state_backups_total{kind,result}andcaptf_state_restores_total{kind,result}; runner stepsstate-pushandstate-list. A TerraformCluster’s restore takes part in the cluster operation gate like its apply. - User variables (see
common.md“User variables” and Module Variables): TerraformCluster and TerraformMachine (and both templates) gainspec.variables, an inline JSON object, andspec.variablesFrom, up to 16 ConfigMaps or Secrets labeledcaptf.io/variables=true, eachoptionaland read informatString(default) orJSON. Sources merge in list order, a later one winning; inline variables win over all. Each variable becomes a named argument ofmodule "role", declared in the generated root without a type (sensitive = truewhen its value came from a Secret) and valued interraform.tfvars.json; a name the module does not declare fails the apply. Reserved:captf_names, the role’s contract inputs and the module meta-arguments. The inputs hash covers the variables only when an object sets some, so existing objects keep theirh2hash and do not re-apply. On a TerraformCluster a changed variable or source re-applies; on a TerraformMachine both fields are immutable and read only until the machine is provisioned. NewDependenciesReadyFalse reasonsVariablesSourceNotFoundandVariablesInvalid. The manager’s ClusterRole gainsget,listandwatchonconfigmaps, and it runs a second, label-scoped cache (captf.io/variables=true, data stripped) for the source watches. - Run leases (the module contract is unchanged; see
The Reconcile Lifecycle): the manager
holds a
coordination.k8s.io/v1Leasecaptf-run-<suffix>per object before it creates a Job, so a stale Job cache, a leader-election handover or two manager instances with overlapping--watch-filtervalues cannot run two Jobs of one object. With the new--cluster-operation-gate(default true) a TerraformCluster’s apply or destroy and its machines’ applies and destroys never run at once (cluster write Leasecaptf-cluster-<hash>); refresh and drift are not gated. NewApplyJobSucceededUnknown reasonsWaitingForRunLease,WaitingForClusterOperationandWaitingForMachineOperations(DriftJobSucceededgainsWaitingForRunLease), the Normal events of the same names once per wait, andcaptf_lease_waits_total{kind,reason}. The manager’s ClusterRole gainscreateandupdateonleases. - Events for every stage (the module contract is unchanged; see
Observability “Events”): the
runner now reports its progress as
events.k8s.io/v1Events on the object that owns the Job (RunStarted,StepStarted,StepSucceeded,StepFailed,PlanSummary,ResourcesChanged,RunFinished; reporting controllercaptf.io/runner), best effort and never failing the run. Job pods get two new runner flags,--event-objectand--job-name; the manager’s new--runner-events(default true) turns them off. Thecaptf-runnerClusterRole gainsevents.k8s.ioeventscreate. The manager emits new reasons once per transition:JobSucceeded,JobInterrupted,JobDeadlineExceeded,StuckJobDeleted,DestructivePlanApprovalConsumed,DeletionStarted,FinalizerRemoved,Paused,Resumed,ProviderIDSet,ControlPlaneEndpointSet,FailureDomainsChanged,InputsChanged,StateAdopted,StateLost,StateLocked(both formerlyStateUnreadable),DriftResolved,DriftRemediationStarted,InstanceHealthy,InstanceUnhealthy,MirrorCreated,MirrorRemoved,IdentitySecretFound,IdentitySecretNotFound,CapacityResolved, andConditionChangedfor any other owned condition’s status or reason change.JobFailednow also covers failed drift and refresh Jobs;JobCreatednames the attempt, the image and why the Job started. - Destructive-plan guard: a TerraformCluster
apply, including a drift remediation, runs
plan -out,show -jsonand an apply of the saved plan, and stops before a plan whose actions includedelete(a removal or a replacement) unless thecaptf.io/approve-destructive-planannotation names the inputs hash it renders. Newstatus.lastRun.error.kind: blocked, reasonApplyJobSucceeded=False/DestructivePlanBlocked, eventDestructivePlanBlockedandcaptf_jobs_total{result="blocked"}. The image’s runtime must supportplan -out=<file>andshow -json <planfile>(see Image Contract). Machine applies are unchanged. - Controller remediation additions (the module contract is unchanged):
- New
remediation.healthCheckIntervalSeconds(60–86400, default 300): withannotateMachine, a provisioned machine is refreshed at that interval to sample health, independent of drift. A health sample is one completed refresh or drift Job, not a new state serial. - New condition reasons:
StateReadable=False/StateLost(a provisioned object’s state is gone or lost its inputs hash; no Job runs),StateReadable=False/StateLocked(the state lock is held by something other than the object’s runner) andApplyJobSucceeded=False/InputsTooLarge.DriftDetectedstarts asUnknown/DriftNotChecked. captf_job_attemptsrecords the retry number of a successful Job (failures of that op since its last success, plus one);status.activeJob.attemptis the Job’s sequence number.
- New
tfcapi-lint’sinput/extraerror is replaced by theinput/user-variable-defaultwarning (a non-contract variable without a default is legitimate when every object is meant to set it throughspec.variables/spec.variablesFrom).
Changed
- The contract pages (
README.md,common.md,cluster.md,machine.md,machinepool.mdand this changelog) were reorganized for readability: long paragraphs were split into sections, table cells were shortened with their detail moved into a section below each table, and source citations became links. The contract itself did not change: every MUST/SHOULD/MAY requirement, every schema and every skeleton is unchanged. - Fewer Jobs per bring-up and in steady state (the module contract is
unchanged; see Machine Remediation,
Plan Approval and
The Reconcile Lifecycle): a
TerraformMachine no
longer refreshes after a successful apply whose own outputs are valid
with a
health.stateother thanpendingorunknown; that reading is the post-apply health sample (status.lastRefreshis the apply’s finish, and it counts once towardstatus.unhealthySamples). While health ispendingthe refresh backs off 30s, 1m, 2m, 4m, then every 5m (plus the UID jitter) instead of a fixed 30s; TerraformCluster and TerraformMachine gainstatus.pendingRefreshes, the consecutive pending samples since the last other reading or apply (status only; it restarts afterclusterctl move). A guarded or approved (--expect-plan) apply whose plan has no changes skips the apply step and succeeds right after the plan, with zero changes and noResourcesChangedevent, so its leases go as soon as it is bookkept. The example bring-up (one cluster, six machines, one no-change re-apply) drops from 14 Jobs to 8. - Controller remediation behavior (the module contract is unchanged):
- The
cluster.x-k8s.io/remediate-machineannotation CAPTF set (markedcaptf.io/remediation-requested) is removed once the instance reads Healthy and the Machine is not being deleted. - The first drift check runs one interval after the last successful apply; drift and health deadlines get a per-object jitter of up to 10%.
- A failed apply is retried even when the inputs equal the state’s hash. An image change clears the pinned digest in the durable inputs Secret.
- The
TerraformMachine.spec.driftandspec.defaults.driftare aMachineDriftPolicywithintervalSecondsonly: machine drift is always reported. The cluster’sspec.drift.actiondefaults toReport, notRemediate.- Machine jobs and drift policies are merged field by field over the
cluster’s
spec.defaults(env by name,imagePullSecretsas a union), instead of replacing them as a whole.spec.defaultsno longer applies to the TerraformCluster itself. TerraformMachine.spec.jobs,driftandremediationare mutable;sourceandidentityRefstay immutable.jobs.lockTimeoutSecondsis optional as a pointer;jobs.activeDeadlineSeconds(at most 86400) andremediation.unhealthyThresholddefault when unset (0), not through a pointer. A policy that sets bothlockTimeoutSecondsandactiveDeadlineSecondsmust keep the lock timeout below the deadline.jobs.securityContextmay not setprivileged,allowPrivilegeEscalationorcapabilities.add.TerraformCluster.spec.identityRefis required; machines fall back tospec.defaults.identityRef, then tospec.identityRef.spec.controlPlaneEndpointis immutable once it has a host.TerraformClusterIdentity.spec.allowedNamespaces: {}is rejected; writeselector: {}to allow every namespace. The identity has a status (conditions,namespaces), filled by the manager.status.lastRun.error.tailis nowsummary(at most 512 bytes): the runner’s summary of the failure, not raw stderr.captf_cluster_outputsis not passed to the cluster role. Previously the text said it wasnull. The generated cluster root declares no such variable, and its tfvars carry no key. A cluster module may leave it undeclared or declare it withdefault = null, as the skeletons do. A declaration without a default failsvalidate. Updatedcommon.md,cluster.md, and thecluster-inputs.jsondescription; the schema still acceptsnull.- Consequence of the inputs hash: the hash accepts only integer numbers. The
cluster role’s
exportsvalue becomes every machine’scaptf_cluster_outputsinput, which is hashed. So a fractional or exponent number anywhere inexports(for example{"ratio": 0.5}) makes those machines’ inputs unhashable, and they cannot apply. Module authors should export numbers as integers or strings. The controller reports it asOutputsValid=False/OutputsInvalidonexports. Notfcapi-lintcheck catches this today. The alternative, a canonical float encoding in the hash, was not chosen, because the design calls for integers only and a float that does not round-trip exactly would change the hash silently. - API field names: the drift and job durations are integer seconds,
following the Cluster API v1beta2 convention and the kube-api-linter
nodurationsrule.spec.drift.intervalis nowspec.drift.intervalSeconds,spec.drift.refreshIntervalisspec.drift.refreshIntervalSeconds, andjobs.lockTimeoutisjobs.lockTimeoutSeconds. Defaults are unchanged (1800, pool-only 60 with a minimum of 15, and 300 seconds). References in these documents were renamed.
Removed
jobs.backoffLimit(Jobs always getbackoffLimit: 0) andjobs.ttlSecondsAfterFinished(Jobs never get a TTL): the controller owns retries, and backoff, digest pinning and conditions are derived from the Jobs it retains.spec.drift.refreshIntervalSeconds(it did nothing),spec.source.imagePullSecrets(usejobs.imagePullSecrets, which covers the source and runner images) andspec.source.command(the image contract fixes/captf/runtime; the image’s/captf/runtimeis now the only executable the runner starts). The durable inputs Secret no longer carriescaptf.io/command, and the inputs hash no longer covers a command: the scheme is nowh2, so every provisioned mutable object re-applies once after the upgrade.- The five no-op mutating webhooks; only validating webhooks remain.
Unused condition reasons
RBACPending,StateReadPendingandCapacityResolving.
Fixed
machinepool.md’s “Minimal skeleton” gavefailure_domainsandcluster_failure_domainsadefault = []; both are always set (non-null) by the controller, sotfcapi-lint’sinput/defaultcheck warned on either one, and--strictfailed. The defaults are removed; the inputs are documented as always set (may be[]), matchingmodules/noop/machinepool. Every role’s “Minimal skeleton” (cluster.md,machine.md,machinepool.md) now also referencescaptf_tags(an emptyterraform_datastub), since the skeletons are now actually run throughtfcapi-lint module --role <role> --strict, which otherwise warnsinput/tags-unusedon all three.cluster.md’s “Minimal skeleton” typedcluster_network’s attributes withoptional(...), while the Inputs table types them as required and the controller always renders all four (an unset CIDR list as[], an unset scalar asnull, never omitted). Both forms passtfcapi-lint’sinput/typecheck, but the skeleton now matches the table.- The
output/provider-id-list-shapecheck registers at botherror(the output is missing or not a list) andwarning(the expression is not sorted and deduplicated);reference/tfcapi-lint-cli.mdlisted it as two separate rows with the same description and no way to tell them apart. It is now one row whose Severity cell names both. machinepool.mdsaidprovider_id_list’s order was irrelevant (true for the controller’s own write, which sorts and de-duplicates) without mentioning thattfcapi-lintseparately checks the output’s own expression forsort()/distinct(), and the skeleton’s comment pointed atmodules/noop/machinepool/outputs.tfinstead of explaining why. Both are now explained inline, with the pattern shown directly.machinepool.md’snode_labelssaid only that the module “MUST render these into kubelet registration … through the bootstrap/agent config it controls,” without saying how, whilebootstrap_datais opaque and may be gzipped or Ignition. The realistic options (amultipart/mixedcloud-config extra part, with a worked example; Ignition and gzipped payloads needing format-aware handling; or a module declaring only the formats it supports) are now spelled out. No new MUST or SHOULD is added.runtime-environment.mddid not say that a provider’s own debug logging (TF_LOG) cannot be turned on for a Job: it is dropped like every otherTF_*name, andspec.jobs.envrejects it outright. The page now says so and points at running the pinned image locally against copies of the rendered inputs (the stuck-destroy runbook’s recipe) withTF_LOGset, as the supported way to get provider debug output.runtime-environment.mddid not say that a network provider mirror or a private registry cannot be configured at run time:TF_CLI_CONFIG_FILEis runner-owned and every otherTF_*name is dropped. The page now states this and points at a filesystem mirror baked into the image (/captf/providers) as the supported path.machinepool.mdandcommon.mddescribedhealth,instancesand their controller-side effects individually, without saying how a pool module should combine mixedinstances[*].statevalues into one grouphealth— the controller does not aggregate them itself.machinepool.md“Per-instance state” now states what the controller does (nothing: both outputs are copied independently) and what a module SHOULD do to derivehealthfrominstances(a new SHOULD, not a MUST).image-contract.mdrecommended theorg.opencontainers.image.*labels but the reference Containerfiles (docs/book/src/module-author/examples/Containerfile.terraformand.opentofu) never set them. Both now takeIMAGE_SOURCE,IMAGE_REVISIONandIMAGE_VERSIONbuild args (empty by default) and set the three labels from them.
v1alpha1 (provisional freeze)
Frozen for implementation: Go types and the controller are built against this version. It may still change until the first real module has provisioned a cluster (the contract is frozen for good only once one real cloud module has created a cluster); every such change is recorded here.
Scope
README.md,common.md,cluster.md,machine.md,machinepool.mdandschemas/in this directory.- Image Contract and the reference
Containerfiles (
docs/book/src/module-author/examples/,modules/noop/). - The control-plane guides in Control-Plane Integration.
Changed at the freeze
- Provider mirror wildcard. The runner writes
include = ["*/*/*"]andexclude = ["*/*/*"], not*/*. Verified with the pinned runtimes Terraform v1.16.4 and OpenTofu v1.12.6, and separately with OpenTofu v1.11.5:*/*matches providers on the default registry host only, so a provider from any other host fell through todirect;*/*/*matches every host. - Remediation wording (
machine.md“Health → remediation”). Remediation happens for Machines whose owner acts onMachineOwnerRemediated: a MachineSet or a control-plane provider that implements remediation (KubeadmControlPlane, RKE2ControlPlane), not only a MachineSet or KubeadmControlPlane.
Added at the freeze
- Reserved paths.
/captf/bin(the injected runner) and/captf/config(the per-run Secret) are mount points; an image MUST NOT ship anything there. Added to the image contract’s “Fixed paths” table and thetfcapi-lint imagechecklist. - Machine-checkable JSON Schemas (draft 2020-12) in
schemas/:definitions.json,cluster-inputs.json,cluster-outputs.json,machine-inputs.jsonandmachine-outputs.json, with a valid and an invalid example for each. The schemas change nothing in the contract; they encode it.- Inputs are closed at every object level (module inputs are exactly the contract inputs); outputs are open (extra, non-contract outputs are ignored).
- Outputs are validated after the controller’s normalization:
""for an ID such asprovider_idbecomesnullfirst, so the schema requires 1–512 characters for a non-nullprovider_id. - Limits come from the CAPI v1.14.2 API markers, including three the role
documents do not spell out:
service_domain1–253 characters, at most 100 CIDR blocks of 1–43 characters each,api_server_port1–65535. - Not expressible in the schemas and enforced by the controller: unique
failure-domain names, the canonical address order, and a machine’s
failure_domainoutput equal to its input when one was requested.
- The image contract’s reference Containerfiles create
/captf/providersbefore runningproviders mirror. Neitherterraform providers mirror(v1.16.4, in the reference build) nortofu providers mirror(checked with OpenTofu v1.11.5 on the build host) creates the target directory when the module requires no providers, so the previous recipe failed withlstat /captf/providers: no such file or directoryfor a provider-less module (found building the reference images).
Decisions
- Operator decisions: native pool autoscaling, runner Secret access accepted and documented (see Security Model), static credential Secrets only in v1.
- Every candidate CAPI mapping considered was decided; the adopted ones
are in the role documents (e.g. address order,
interruptible, clusterkubernetes_version,control_plane_initialized, endpoint provenance, emptydataSecretName,captf.io/template). - The role documents’ claims were checked against CAPI source and
resolved. Resolutions for the items raised in the RKE2ControlPlane
guide:
- binary bootstrap data: resolved by always-base64
bootstrap_data; - providerID under RKE2: no contract change; the machine module makes
the kubelet
provider-idmatch (CCM or kubelet extra args), as the guide and checklist say; - a second listener on 9345: a cluster-module convention, not a contract field, since CAPI has nowhere to put it;
- the
InternalIPguarantee: no contract change; notfcapi-lintcheck exists for it ininternal/lint; - the version string: stays verbatim; modules strip
+rke2rN(asmachine.mdsays), because normalizing would change the inputs hash; - remediation wording: fixed above;
- CAPRKE2 on CAPI v1.14: needs a compatibility smoke test against CAPI v1.14.2 before RKE2ControlPlane is relied on in production;
- a nondeterministic join target: upstream behavior, documented in the guide.
- binary bootstrap data: resolved by always-base64
- The contract is authoritative on
captf_contract: every role receives it (common.md), and the renderer renders it for every role. - Proposals considered and not adopted for
common.md,cluster.md,machine.mdandmachinepool.md(extra tags,captf_identity,health.observed_at, aresourcesinventory output, clustertags/dns_zone, machineimage/instance_type/disk/ssh_authorized_keys/machine_uid,failure_domain_attributes,bootstrap_data_hash,instance_id, poolrolling_updateandready_replicas): none loosens the closed input schemas (schemas/) or is needed by a shipped module; several are fixed in the module instead because the CRDs have no free-form per-object variables outsidespec.variables/spec.variablesFrom.
Image Contract
The OCI image is the deliverable a module author ships. This page is the
normative image contract for the v1alpha1 module contract: the fixed
paths CAPTF’s runner looks for, the labels it and tfcapi-lint read, the
user the image should run as, and multi-arch publishing. tfcapi-lint image checks an image against this contract. What the module actually
sees when the runner executes it — the generated root, the environment,
the commands run and their order — is on
Runtime Environment.
One image bundles one role module (cluster, machine or machinepool)
and the runtime that runs it (tofu or terraform). There is no
separate module source and no separate runtime image: the image tag is the
module version, the image is what spec.source.image references, and the
image is what gets pinned, moved, rolled out and audited. The controller
runs it as a Job with its own runner binary as the entrypoint; the image
itself never needs a shell.
Fixed paths
The runner discovers a module by fixed path, not by label or registry
metadata, and fails the Job at start (error.kind: image-layout) if a
required path is missing or unusable.
| Path | Required | Contents |
|---|---|---|
/captf/module/ | yes | The role module: at least one .tf, .tf.json, .tofu or .tofu.json file at its top level, plus any local module it references by a relative source. It cannot declare a terraform { backend … } or cloud block: those are only valid in a root module, and the generated root, not this one, is the root. |
/captf/runtime | yes | The tofu or terraform binary: a regular file, or a symlink to one inside the image, executable by the image’s USER. It must support the Terraform 1.x / OpenTofu 1.x CLI surface: version, init, validate, plan, apply, destroy, force-unlock, show, and state push/state list. There is no override for this path. |
/captf/providers/ | no | An optional provider filesystem mirror (below). Without it, init needs registry egress. |
/captf/work/, /captf/bin/, /captf/config/, /var/run/captf/credentials/ | must be empty | The Job mounts an emptyDir, the runner binary, the per-run Secret and the identity’s credential files at these paths respectively. Anything the image ships under them is shadowed (or, for /captf/work, never used, since the image’s own root filesystem is read-only by default). |
Everything else in the image is the author’s business: CA certificates,
git for provider blocks that shell out, or a helper binary a
local-exec provisioner calls.
Provider mirror layout
Build /captf/providers with terraform providers mirror <dir> or tofu providers mirror <dir>, for every platform the image publishes (for
example -platform=linux_amd64 -platform=linux_arm64). Both runtimes
accept either layout it can produce: the packed layout
(HOST/NAMESPACE/TYPE/terraform-provider-TYPE_VERSION_TARGET.zip plus
.json index files) or the unpacked layout
(HOST/NAMESPACE/TYPE/VERSION/TARGET/).
providers mirror creates its target directory only when it writes at
least one provider, so a module that requires none needs the mirror stage
to create /captf/providers itself (see the reference Containerfiles
below) or the later COPY --from=mirror fails. It also fails with
“Module not installed” when the module calls local modules that are not
yet installed, so run terraform get / tofu get first, in the same
stage; that installs modules only, never providers.
Providers are optional: an image without /captf/providers still works,
but needs registry egress at init and is slower and non-hermetic. The
reference images in this repository ship a mirror. How the runner uses the
mirror at run time is on
Runtime Environment.
OCI labels
Labels are metadata only: the runner never reads them for behavior.
tfcapi-lint image checks them for consistency with --role/--contract.
| Label | Value |
|---|---|
io.captf.contract | Contract version, for example v1alpha1. |
io.captf.role | cluster, machine or machinepool. |
io.captf.runtime | tofu or terraform: what /captf/runtime is. |
io.captf.runtime.version | For example 1.12.6. |
org.opencontainers.image.source, .revision, .version | Standard OCI annotations; .version should equal the tag. |
The reference Containerfiles below take these three as IMAGE_SOURCE,
IMAGE_REVISION and IMAGE_VERSION build args (empty by default), so a
build pipeline sets them with --build-arg from the source repository
URL, the commit, and the tag being built.
Capacity labels (machine role only, optional) let
TerraformMachineTemplate support Cluster Autoscaler scale-from-zero: the
module fixes the instance type, so the image is the only place that knows
the node’s size. Set both identically on every platform of a multi-arch
index; an image without them leaves status.capacity/status.nodeInfo
unset.
| Label | Value | Maps to |
|---|---|---|
io.captf.capacity | JSON object, resource name to Kubernetes quantity string, for example {"cpu":"4","memory":"16Gi","nvidia.com/gpu":"1"}. Each key is a valid Kubernetes resource name; each value parses as a quantity. | TerraformMachineTemplate.status.capacity |
io.captf.node-info | JSON object {"architecture":"amd64","operatingSystem":"linux"}; architecture is one of amd64, arm64, s390x, ppc64le; at least one key set. | TerraformMachineTemplate.status.nodeInfo |
A module whose instance type varies needs one image per instance type to use these labels; pool images may carry them, but they are ignored.
User
Any UID works for the runner, but recommend a non-root USER (for example
65532) so the Job can run under a namespace that enforces the Pod
Security restricted profile. Files and directories under /captf/module
and, when present, /captf/providers must be readable, and directories
traversable, by that UID. Credential files are mounted mode 0440, so a
non-root image user reads them through the pod’s fsGroup, not through
ownership. See Security Model
for how the pod’s own security context defaults interact with the image’s
USER.
Multi-arch
Publish a multi-arch manifest (linux/amd64, linux/arm64) or pin the
management cluster’s node architecture to the platform the image ships:
tfcapi-lint image checks the linux/amd64 platform by default, and
--platform/--all-platforms select others. A provider mirror must carry
a package for the target platform it is checked against.
Versioning and pinning
The image tag is the module version: spec.source.image is
registry/repo:tag or registry/repo@sha256:…, and a new module version
is a new tag referenced by a new Terraform*Template. The controller pins
the digest it actually ran after the first successful apply and, for
immutable machines, uses that digest for every later drift and destroy
Job; see Security Model
for the full mechanics and why it matters.
Contract version is not declared in the image: labels are informational.
The controller injects the contract it generates against as
captf_contract, and tfcapi-lint takes --contract on the command
line.
Trust boundary
Referencing an image is equivalent to granting its publisher the runner’s Secret access and the resolved identity’s cloud credentials in that namespace: see Security Model for the full trust boundary and what it means for review and tenancy.
Building an image
Lint the module first, then build, then lint the pushed image:
tfcapi-lint module ./cluster --role cluster --strict
podman build -t "$IMAGE" .
tfcapi-lint image "$IMAGE" --role cluster --strict
The reference images below ship as
examples/Containerfile.terraform and
examples/Containerfile.opentofu, each
taking ARG ROLE and ARG RUNTIME_VERSION.
Reference: Terraform base
hashicorp/terraform is Alpine with git, openssh and CA certificates,
its binary at /bin/terraform, and ENTRYPOINT ["/bin/terraform"]; the
runner replaces that entrypoint, so it has no effect. The image has no
USER (root); the Containerfile below adds one.
# Reference source image, Terraform base (docs/book/src/module-author/image-contract.md "Reference:
# Terraform base"). Run from your module's root directory:
#
# podman build -f Containerfile.terraform --build-arg ROLE=cluster \
# --build-arg IMAGE_SOURCE=https://github.com/<org>/<repo> \
# --build-arg IMAGE_REVISION="$(git rev-parse HEAD)" \
# --build-arg IMAGE_VERSION=<tag> -t <registry>/<repo>:<tag> .
#
# The image tag is the module version. Lint first:
# tfcapi-lint module . --role cluster --strict
ARG RUNTIME_VERSION=1.16.4
# The org.opencontainers.image.* labels below (image-contract.md "OCI
# labels"): leave these unset for a local/test build, or pass them from
# your CI pipeline (source repo URL, commit SHA, the image tag).
ARG IMAGE_SOURCE=""
ARG IMAGE_REVISION=""
ARG IMAGE_VERSION=""
# Optional but recommended: hermetic provider mirror for the platforms you
# publish. Needs registry egress at build time; drop this stage (and the
# COPY --from=mirror below) for a non-hermetic image.
FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION} AS mirror
WORKDIR /src
COPY . /src
# get: `providers mirror` refuses a module whose nested local modules are not
# installed; get installs them (no providers), in this stage only.
# mkdir: `providers mirror` does not create the target when the module
# requires no providers, and the COPY --from=mirror below needs it.
RUN terraform get \
&& mkdir -p /captf/providers \
&& terraform providers mirror -platform=linux_amd64 -platform=linux_arm64 /captf/providers
FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION}
ARG ROLE=cluster
ARG RUNTIME_VERSION
ARG IMAGE_SOURCE
ARG IMAGE_REVISION
ARG IMAGE_VERSION
COPY --from=mirror /captf/providers /captf/providers
COPY . /captf/module
RUN ln -s /bin/terraform /captf/runtime \
&& adduser -D -u 65532 captf \
&& chown -R 65532:65532 /captf
USER 65532
LABEL io.captf.contract="v1alpha1" \
io.captf.role="${ROLE}" \
io.captf.runtime="terraform" \
io.captf.runtime.version="${RUNTIME_VERSION}" \
org.opencontainers.image.source="${IMAGE_SOURCE}" \
org.opencontainers.image.revision="${IMAGE_REVISION}" \
org.opencontainers.image.version="${IMAGE_VERSION}"
Mirroring providers needs registry egress at build time; drop the
mirror stage (and its COPY --from=mirror) for a non-hermetic image.
Reference: OpenTofu base
opentofu:*-minimal is FROM scratch with only the tofu binary — no CA
certificates, no shell, no git — which is why the Containerfile below
copies it into an Alpine stage instead of using it directly as the final
base. The full (non-minimal) OpenTofu image refuses to be used as a
FROM base.
# Reference source image, OpenTofu base (docs/book/src/module-author/image-contract.md "Reference:
# OpenTofu base"). Run from your module's root directory:
#
# podman build -f Containerfile.opentofu --build-arg ROLE=machine \
# --build-arg IMAGE_SOURCE=https://github.com/<org>/<repo> \
# --build-arg IMAGE_REVISION="$(git rev-parse HEAD)" \
# --build-arg IMAGE_VERSION=<tag> -t <registry>/<repo>:<tag> .
#
# The image tag is the module version. Lint first:
# tfcapi-lint module . --role machine --strict
#
# The full ghcr.io/opentofu/opentofu image refuses to be a FROM base
# (ONBUILD RUN exit 1); use the -minimal tag via COPY --from as below.
ARG RUNTIME_VERSION=1.12.6
# The org.opencontainers.image.* labels below (image-contract.md "OCI
# labels"): leave these unset for a local/test build, or pass them from
# your CI pipeline (source repo URL, commit SHA, the image tag).
ARG IMAGE_SOURCE=""
ARG IMAGE_REVISION=""
ARG IMAGE_VERSION=""
FROM ghcr.io/opentofu/opentofu:${RUNTIME_VERSION}-minimal AS tofu
# Optional but recommended: hermetic provider mirror for the platforms you
# publish. Needs registry egress at build time; drop this stage (and the
# COPY --from=mirror below) for a non-hermetic image.
FROM docker.io/library/alpine:3.22 AS mirror
RUN apk add --no-cache ca-certificates
COPY --from=tofu /usr/local/bin/tofu /usr/local/bin/tofu
WORKDIR /src
COPY . /src
# get: `providers mirror` refuses a module whose nested local modules are not
# installed; get installs them (no providers), in this stage only.
# mkdir: `providers mirror` does not create the target when the module
# requires no providers, and the COPY --from=mirror below needs it.
RUN tofu get \
&& mkdir -p /captf/providers \
&& tofu providers mirror -platform=linux_amd64 -platform=linux_arm64 /captf/providers
FROM docker.io/library/alpine:3.22
ARG ROLE=machine
ARG RUNTIME_VERSION
ARG IMAGE_SOURCE
ARG IMAGE_REVISION
ARG IMAGE_VERSION
RUN apk add --no-cache ca-certificates \
&& adduser -D -u 65532 captf
COPY --from=tofu /usr/local/bin/tofu /captf/runtime
COPY --from=mirror /captf/providers /captf/providers
COPY . /captf/module
RUN chown -R 65532:65532 /captf
USER 65532
LABEL io.captf.contract="v1alpha1" \
io.captf.role="${ROLE}" \
io.captf.runtime="tofu" \
io.captf.runtime.version="${RUNTIME_VERSION}" \
org.opencontainers.image.source="${IMAGE_SOURCE}" \
org.opencontainers.image.revision="${IMAGE_REVISION}" \
org.opencontainers.image.version="${IMAGE_VERSION}"
A distroless/static:nonroot final stage also works, and needs no shell,
if the module needs no other tool: tofu is statically linked.
Checklist for tfcapi-lint image
Summarized: /captf/module present and lints clean for the role;
/captf/runtime present and executable; /captf/providers, if present,
follows the mirror layout and covers every required_providers entry for
the platform being checked; labels, if present, agree with
--role/--contract; the capacity labels, if present, are valid JSON of
the shapes above; no files under the reserved paths; config.User is
non-root (running as root is a warning, not an error). See
tfcapi-lint CLI: checks for
every check’s ID, severity and the roles it applies to.
See also
Runtime Environment
This page describes what a module sees once CAPTF actually runs it: the working directory, the environment, the provider mirror, the commands the runner runs and in what order, and how a failure is reported. It complements Image Contract, which is the static shape of the image itself.
For every operation, the runner replaces the image’s ENTRYPOINT/CMD
with its own binary, checks the image’s layout (below), prepares a working
directory, then runs /captf/runtime (the image’s tofu or terraform)
as a child process, one command at a time, streaming its output to the
Job’s log.
Working directory and the generated root
Before running anything, the runner checks that /captf/module holds at
least one .tf, .tf.json, .tofu or .tofu.json file at its top level
and that /captf/runtime is executable; either failure stops the Job with
error.kind: image-layout before any command runs.
It then builds a generated root at /captf/work/root from the files
mounted read-only at /captf/config (the per-run Secret): main.tf.json,
which declares a partial kubernetes backend (init’s -backend-config=
flags complete it) and calls the image’s module as module "role" { source = "../../module" } (resolving to /captf/module, a local path, so init
never fetches the module from anywhere), and terraform.tfvars.json, the
rendered contract and user variables. A restore’s root is the backend block
alone: it calls no module and declares no variable, since state push and
state list need neither. Every command in this page runs with
/captf/work/root as its working directory; the runner never uses
-chdir. What the module actually receives through these two files — the
contract inputs, user variables and where each value comes from — is on
Job Inputs.
The runner never invokes a shell: it execs /captf/runtime directly. A
module’s local-exec provisioner needs a shell in the image, or its own
interpreter, to run at all: the reference images on
Image Contract ship one, but a
distroless/static final stage does not. The image’s root filesystem is
read-only by default (readOnlyRootFilesystem: true); the only writable
paths are /captf/work (this generated root, the plan files and
TF_DATA_DIR) and /tmp. A provider that writes anywhere else needs the
module or the image to relocate it, typically through an environment
variable such as a cache directory setting.
Environment set and dropped
The runner starts from the container’s own environment (the identity’s
envFrom, the image’s ENV, and spec.jobs.env) and rewrites it before
running any command: HOME is always forced to /captf/work and
TF_DATA_DIR to /captf/work/.terraform (note that the working directory
of every command is /captf/work/root, one level below); TMPDIR is kept
if already set (the Job sets it to /tmp) and otherwise defaults to
/captf/work/tmp; and every other
TF_* and KUBE_* variable is dropped — logged by name, never by value —
except the three the Job itself sets (TF_IN_AUTOMATION, TF_INPUT,
KUBE_NAMESPACE). This stops a leaked TF_WORKSPACE, TF_CLI_ARGS_*,
TF_VAR_*, TF_LOG or KUBE_* value from moving state, rewriting a step’s
flags, overriding a rendered input, or logging provider traffic and
credentials. See Job Environment for the
full variable list, a worked example of what is kept and dropped, and the
names spec.jobs.env may not set at all.
There is no way to turn on provider debug logging in a Job. TF_LOG
falls under the dropped TF_* names above; spec.jobs.env cannot set it
either, since any name starting with TF_ or KUBE_ is rejected outright
(with an event) rather than passed through. A module author debugging a
provider needs to reproduce the run outside a Job: pull the pinned image,
and run its /captf/runtime binary by hand against a copy of the
rendered root and inputs, with TF_LOG set in that shell. The
stuck-destroy runbook
walks through getting those inputs (the durable inputs Secret’s
main.tf.json/terraform.tfvars.json, the image digest, and the
identity’s credentials) and invoking the image’s binary directly; the same
recipe works for a debug run, add -e TF_LOG=DEBUG (or TRACE) to the
docker run/podman run invocation.
Provider mirror and CLI configuration
When the image ships /captf/providers, the runner writes
/captf/work/cli.tfrc:
provider_installation {
filesystem_mirror {
path = "/captf/providers"
include = ["*/*/*"]
}
direct {
exclude = ["*/*/*"]
}
}
and sets TF_CLI_CONFIG_FILE to that path, so init installs every
provider from the image and never contacts a registry; init fails
clearly instead of silently downloading a provider missing from the
mirror. Without /captf/providers, the runner sets neither, and init
falls back to the runtime’s default direct installation, which needs
registry egress. TF_PLUGIN_CACHE_DIR is never set: it must not coincide
with a filesystem mirror.
A network mirror or a private registry cannot be configured at run
time. TF_CLI_CONFIG_FILE is runner-owned, written fresh for every Job
as shown above (or left unset), and spec.jobs.env cannot override it or
any other TF_* name (see “Environment set and dropped”); there is no
field that lets an object or its defaults point init at a network_mirror
block or a private registry host’s credentials. The supported path is a
filesystem mirror baked into the image at /captf/providers (see Image
Contract): build it from
whatever upstream, mirror or private registry the image’s build pipeline
can reach, and ship the result. There is no run-time equivalent.
Commands, and their order
version -json runs first, once per Job, to record the runtime version for
the result document; its outcome is informational and never fails the Job.
init, plan, apply and destroy (including -refresh-only) carry
-input=false -no-color -lock-timeout=<n>s (init also carries a
-backend-config= flag for each value the generated root’s partial
kubernetes backend needs); validate and show carry -json -no-color
instead; state push carries only -force -lock-timeout=<n>s; and
state list and force-unlock take neither.
Every command runs against the generated root in order, and a non-zero exit
(outside the codes a step accepts, such as plan’s 2 for “changes
present”) stops the Job at that step. When the Job carries a stale lock ID
to clear, force-unlock -force <id> runs immediately after init
(force-unlock needs an initialized backend); it accepts an already-unlocked
state as success, since a retried pod or another holder may have cleared it
already.
| Operation | Commands, in order |
|---|---|
| Apply (machine and machine-pool roles) | init → validate -json → apply -auto-approve -var-file=<tfvars> |
Apply (cluster role: always guarded, and re-planned against the approved hash under applyPolicy: Manual) | init → validate -json → plan -detailed-exitcode -var-file=<tfvars> -out=<file> → show -json <file> (only if plan exited 2) → apply <file> |
| Destroy | init → destroy -auto-approve -var-file=<tfvars> |
| Refresh | init → apply -refresh-only -auto-approve -var-file=<tfvars> |
| Drift check | init → apply -refresh-only -auto-approve -var-file=<tfvars> → plan -detailed-exitcode -refresh=false -var-file=<tfvars> -out=<file> → show -json <file> (only if plan exited 2) |
Plan preview (cluster role, applyPolicy: Manual, before approval) | init → validate -json → plan -detailed-exitcode -var-file=<tfvars> -out=<file> → show -json <file> (only if plan exited 2) |
| Restore | init → state push -force <backup> → state list |
<tfvars> is terraform.tfvars.json, in the generated root. An apply
whose plan exits 0 (no changes at all, not even to outputs) skips its
apply step entirely: the Job ends successfully without applying
anything. show -json’s output is never written to the Job’s log, since a
plan document carries every input value; every other command’s output
streams to the log as it runs. validate’s JSON diagnostics are the
exception: they are captured for the failure summary below and still
written to the log, since they carry no input values. state list’s
output (resource addresses only) is logged too. Approving a destructive
plan or a Manual-policy preview is covered in
Plan Approval; the counts, add/change/
destroy resource lists and the drift check itself are covered in
Drift and Health.
Lock timeout and stop timeout
init, plan, apply, destroy and state push all carry
-lock-timeout, defaulting to 300 seconds and overridable per object with
spec.jobs.lockTimeoutSeconds (Tuning Jobs).
On SIGTERM — a deletion, a drain, or the Job’s activeDeadlineSeconds
(default 3600 seconds) — the runner sends /captf/runtime SIGTERM, never
its provider plugins, and gives it up to 570 seconds (the pod’s 600-second
termination grace period, less a margin the runner needs to write the
result) to finish in-flight provider calls, write state and release the
backend lock before SIGKILL. A run stopped this way reports
error.kind: interrupted, not a module failure.
What a failed step reports
A run’s status.lastRun.error never carries raw process output. Its kind
is one of image-layout, step, interrupted, blocked (a guarded apply
stopped before a plan that deletes or replaces a resource) or
plan-changed (an approved apply whose new plan no longer matches the
approved hash); step names the command that failed, when there is one. A
blocked or plan-changed result changes nothing. A plan-changed
result also carries a plan summary in status.plan; a blocked result’s
summary appears in the object’s condition message instead. Both are
covered in Plan Approval.
The failure summary itself is built from the failing step’s own output,
never a raw stderr dump: for validate, from its -json diagnostics; for
every other step, from the Error: diagnostic header lines in its stderr
(ANSI escape codes stripped first, since a provider or a local-exec
child is not bound by -no-color). When neither yields a line, it falls
back to step <name> exited <code>; see the Job's logs. The summary is
capped at 512 bytes; the process exit code is the failing step’s own code,
or 1 when the runner itself could not start or was killed. The full
stderr always reaches the Job’s own pod log, uncapped, whether or not it
contributed to the summary.
Output size limits
| Limit | Value | Applies to |
|---|---|---|
| Failure summary | 512 bytes | status.lastRun.error.summary, above |
| Termination message | 4096 bytes | The whole result document; the kubelet truncates a longer one, so the runner drops fields in stages (resource-change counts, then the error tail, then plan and drift resource lists, then step history) to fit, keeping the plan’s hash and counts last |
| Plan resources listed | 50 | status.plan.resources; a plan with more sets truncated: true |
| Drift resources listed | 20 | The drift check’s resource address list |
| Stderr kept in memory per step | 64 KiB | The tail a step’s failure summary is built from; the log itself is not truncated |
See also
tfcapi-lint
tfcapi-lint checks a Terraform or OpenTofu module, and the OCI image built
from it, against the CAPTF module contract
(module contract) and the
image contract, without running init, plan or
apply. It is for module authors, and for CI pipelines that build module
images. It is released alongside the provider, with the same version.
Install
Every release attaches one binary per platform, plus a checksum file:
| Asset | Platform |
|---|---|
tfcapi-lint-linux-amd64 | Linux x86-64 |
tfcapi-lint-linux-arm64 | Linux ARM64 |
tfcapi-lint-darwin-amd64 | macOS Intel |
tfcapi-lint-darwin-arm64 | macOS Apple silicon |
tfcapi-lint-windows-amd64.exe | Windows x86-64 |
tfcapi-lint-checksums.txt | SHA-256 of every asset above |
base="https://github.com/captf-io/cluster-api-provider-terraform/releases/download/<version>"
curl -fsSLO "${base}/tfcapi-lint-<os>-<arch>"
curl -fsSLO "${base}/tfcapi-lint-checksums.txt"
sha256sum --check --ignore-missing tfcapi-lint-checksums.txt
install -m 0755 "tfcapi-lint-<os>-<arch>" /usr/local/bin/tfcapi-lint
tfcapi-lint version
<version>is the provider release you deploy, for examplev0.1.0.<os>and<arch>pick one row of the table above, for examplelinuxandamd64. On macOS, verify the checksum withshasum -a 256 -c --ignore-missing tfcapi-lint-checksums.txtinstead.
Match the tfcapi-lint release to the controller you deploy against:
tfcapi-lint version --json reports the contract versions it lints
against, ["v1alpha1"].
go install .../cmd/tfcapi-lint@<version> does not work: the repository is
a Go workspace, and cmd/tfcapi-lint depends on the api module, which
only the workspace resolves. Build from source instead (below), or use a
release binary.
Roles
Every module implements exactly one role, and --role on module and
image is required and takes one of cluster, machine or
machinepool, matching the TerraformCluster,
TerraformMachine and
TerraformMachinePool contracts. tfcapi-lint
checks the module or image against that role’s inputs, outputs and checks
only; see tfcapi-lint CLI for
which checks apply to which roles.
Lint a module
tfcapi-lint module --role machine ./machine
This reads the .tf, .tf.json, .tofu and .tofu.json files under
./machine directly; it needs neither terraform nor tofu installed,
and never contacts a registry. A clean module prints an empty finding
list and an all-zero summary.
Lint an image
tfcapi-lint image --role machine registry.example.com/acme/machine:v1.0.0
This pulls the image manifest and its layers, and checks the fixed paths
and labels the image contract requires, without
running the image. <image-ref> is a registry reference; oci:<dir>
reads a local OCI image layout instead, such as one written by
podman save --format oci-dir or skopeo copy ... oci:<dir>, with no
registry or daemon involved. By default it checks the linux/amd64
platform of a multi-platform image; --platform os/arch picks a
different one, and --all-platforms checks every platform the image
publishes.
Registry credentials
tfcapi-lint image authenticates the same way docker and podman do,
through go-containerregistry’s default keychain: it reads
~/.docker/config.json, or $DOCKER_CONFIG/config.json when that
variable is set; if neither exists, it falls back to a Podman-style
config at $REGISTRY_AUTH_FILE or
$XDG_RUNTIME_DIR/containers/auth.json. With none of those present, the
pull is anonymous. Log in with docker login or podman login against
the registry before linting a private image; --insecure allows a
plain-HTTP registry for a local or air-gapped registry that has none.
Strict mode and allowed warnings
--strict treats a warning the same as an error for the exit code, so a
module or image that is merely clean today does not silently pick up new
warnings later. --allow-warning <id> (repeatable, or a comma-separated
list) downgrades one check ID’s warnings to informational findings; it
never touches errors. Use it for a deliberate, reviewable exception, for
example a module whose provider cannot tag anything:
--allow-warning input/tags-unused. Run with --strict by default, and
add --allow-warning only for checks you have decided not to act on.
--json prints a report with a findings array (id, severity,
file, line, message) and a summary, instead of one line of text
per finding. See tfcapi-lint CLI for
every flag, and
tfcapi-lint CLI: checks for
every check ID, its severity and the roles it applies to.
Exit codes
tfcapi-lint uses its exit code to signal a CI step’s pass or fail; see
tfcapi-lint CLI: exit codes
for the full list. In short: 0 is clean, 1 is at least one error (or,
under --strict, at least one warning), 2 means the module could not
be parsed or the image could not be pulled, and 3 is a usage error.
In CI
Lint the module before building the image, then lint the built image before pushing it: the source is checked before the build, and the built layout after it.
tfcapi-lint module --role machine --strict ./module
podman build -t "$IMAGE" .
podman save --format oci-dir -o "$RUNNER_TEMP/image" "$IMAGE"
tfcapi-lint image --role machine --strict "oci:$RUNNER_TEMP/image"
podman push "$IMAGE"
To check exactly what was pushed, for example a multi-platform index built and pushed by a separate step, lint the pushed reference instead of the local layout:
tfcapi-lint image --role machine --strict --all-platforms "$IMAGE"
Building from source
make release-lint-snapshot builds all five release assets and the
checksum file into dist/ with GoReleaser, stamped with a snapshot
version; make release-lint does the same from the current git tag.
See Releasing for how these assets
reach a GitHub release.
See also
- tfcapi-lint CLI for every flag, every check ID and the exit codes.
- Module contract and image contract for what the checks enforce.
- Runtime environment for what a module sees once CAPTF actually runs it.
Control-Plane Integration
CAPTF’s cluster and machine modules are infrastructure only: they create networks, load balancers and instances, and report state back through the contract. Bringing up Kubernetes on top of that infrastructure is the job of a Cluster API control-plane provider — KubeadmControlPlane (KCP) or RKE2ControlPlane (RCP) — plus its bootstrap provider.
Those providers read specific fields from your modules and write specific fields back, on a specific schedule. Getting the details wrong produces a cluster that hangs rather than one that fails loudly.
This guide set distills the verified requirements into what you need as a module author:
kubeadm.md— everything specific to KubeadmControlPlane / the kubeadm bootstrap provider (CABPK).rke2.md— everything specific to RKE2ControlPlane / the RKE2 bootstrap provider (CAPRKE2).checklist.md— every port, health check and ordering rule from both guides, organized by topic, for a build-time or review-time reference.
Both guides cite CAPI/CAPRKE2 source paths and the contract’s own module role pages: the cluster role, the machine role and the common contract. Where this guide set and the contract state the same requirement, the contract’s wording is authoritative.
The shared creation sequence
Both control-plane providers drive the same shape of sequence against a
CAPTF cluster module and one or more CAPTF machine modules. The sequence
below is the kubeadm case; RKE2’s differences are called out inline and
detailed in rke2.md, in what it reads from your
modules and
lifecycle constraints.
- A user (or a ClusterClass topology) creates the
Cluster,TerraformCluster, control-plane object (KubeadmControlPlaneorRKE2ControlPlane) andTerraformMachineTemplate. - The CAPI Cluster controller reconciles the
infrastructureRef, setting its owner reference to theTerraformCluster. - The
TerraformClustercontroller runs the cluster module’s first apply. At this pointcontrol_plane_endpointis renderednullandcontrol_plane_initializedis renderedfalse. The module returnscontrol_plane_endpoint,failure_domains,exportsandhealth. - The controller copies
control_plane_endpointontoTerraformCluster.spec.controlPlaneEndpointandfailure_domainsontoTerraformCluster.status.failureDomains; CAPI in turn copies them ontoCluster.spec.controlPlaneEndpointandCluster.status.initialization.infrastructureProvisioned = true. - The control-plane provider gates all further work on: a valid
Cluster.spec.controlPlaneEndpointandinfrastructureProvisioned = true. Neither KCP nor RCP creates a single Machine before both are true (see what KubeadmControlPlane reads and what RKE2ControlPlane reads). - For control-plane machine
i = 1..N, in strict order — the provider creates machinei+1only after machineihas anodeRef:- The control-plane provider creates the bootstrap config (init for
i=1, join fori>1) and theMachine(labeledcluster.x-k8s.io/control-plane) plus aTerraformMachinecloned from the template. - The bootstrap provider writes the bootstrap Secret (
value,format);Machine.spec.bootstrap.dataSecretNameis set. - The
TerraformMachinecontroller reconciles machinei, gated on: the cluster’sinfrastructureProvisioned, the cluster’sexportsbeing readable, and the bootstrap Secret existing. - The machine module’s apply runs once (immutable): it receives
bootstrap_data(base64),control_plane = true,captf_cluster_outputs(the cluster’sexports),failure_domainandkubernetes_version. In this same apply the module creates the instance and registers it in the control-plane LB backend, or does neither of the latter under Pattern C below. It returnsprovider_id,addresses,failure_domain,interruptibleandhealth. - The
TerraformMachinecontroller writesspec.providerIDandstatus.addresseson theTerraformMachine, and marks it provisioned andReady. The CAPI Machine controller then copies both ontoMachine.spec.providerIDandMachine.status.addresses. - cloud-init on the instance runs
kubeadm init(i=1) orkubeadm join(i>1); post-init phases go through the endpoint (hairpin — see kubeadm.md’s networking section). Under RKE2’s default registration method the join target is the endpoint host itself, but the join cannot proceed until at least one Ready control-plane Machine already exists, sinceRCP.status.availableServerIPsstays empty without one (see what RKE2ControlPlane reads and rke2.md’s lifecycle constraints). - The Node registers with
spec.providerIDset (by a cloud-controller-manager orkubelet --provider-id); the Machine controller matchesnodeRefby an exactproviderIDmatch. - When
i = 1and the control plane finishes initializing,Cluster.status.initialization.controlPlaneInitializedflipstrue. This triggers exactly one cluster re-apply, for resources gated on a live workload API server.
- The control-plane provider creates the bootstrap config (init for
- Workers follow the same machine path (
MachineDeployment→MachineSet→Machine→TerraformMachine,control_plane = false), gated onControlPlaneInitializedbecause the bootstrap provider writes worker join data only after that. A worker fleet backed by aMachinePoolinstead goes straight fromMachinePooltoTerraformMachinePool— one infrastructure object for the whole group, with no per-replicaMachineorTerraformMachine; see Machine Pools.
LB-membership patterns
The contract requires a control-plane machine module to register (and, on destroy, deregister) its instance with the control-plane load balancer (see machine.md’s control-plane machines), but leaves how up to the module. Three patterns are sanctioned here.
Pattern A: attachment resource in the machine module’s state
The machine module creates the instance and, in the same apply, an
explicit attachment/registration resource (for example an LB target-group
attachment) pointed at the target-group or backend-pool id it received
through captf_cluster_outputs (see common.md’s
outputs). Because the attachment
lives in the machine module’s own state, destroy deregisters it as an
ordinary part of tearing that state down.
This is the pattern the contract documents directly: registration is
“done in the machine module’s own Terraform state, not the cluster
module’s” (machine.md’s control-plane
machines, citing
capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645).
The instance MUST be registered before kubeadm init/kubeadm join
finishes (KCP) or before the Machine is Ready (RKE2) — see kubeadm.md’s
module checklist and rke2.md’s module
checklist.
Typical fit: clouds whose load balancer exposes an explicit target-group/backend-pool attach API (for example AWS ALB/NLB target-group attachments, Azure Load Balancer backend-pool membership).
This is design guidance, not yet tested against a real control plane.
Pattern B: selector/tag-based backend pools
The load balancer derives its backend membership itself, from a tag or
label selector query, rather than from an explicit per-instance attach
call — “target group by tag/label”. The machine module’s only job is to
apply the right tags to the instance (via captf_tags, see common.md’s
inputs, or an additional
module-defined tag); the load balancer’s own membership scan does the
rest, and removing the instance (destroy) removes it from the pool.
The same ordering requirement applies as under Pattern A: the instance
MUST be a member of the backend before kubeadm init/join finishes on
it (KCP) or before the Machine is Ready (RKE2) — see kubeadm.md’s module
checklist and rke2.md’s module
checklist.
Typical fit: clouds or load balancers whose backend pool is defined by a selector/tag query against an instance group or autoscaling group, rather than an explicit attach call.
This is design guidance, not yet tested against a real control plane.
Pattern C: kube-vip-style VIP
No separate load-balancer resource is created by the cluster module at
all; a virtual IP is run by the control-plane nodes themselves (implicit
via kube-vip or similar), so there is “nothing LB-wise” for the machine
module to do. The cluster module still MUST emit a valid
control_plane_endpoint before any Machine is created — the same gate
applies regardless of how the endpoint is realized (see what
KubeadmControlPlane
reads and
kubeadm.md’s networking section;
what RKE2ControlPlane
reads and
rke2.md’s networking section).
Under this pattern the backend-membership and per-backend health-check
rows of checklist.md do not apply, because there is no
separate LB backend to join.
Typical fit: bare-metal, on-premises, or otherwise L2-reachable environments without a managed load balancer in front of the control plane.
This is design guidance, not yet tested against a real control plane.
See also
- The module contract — the normative cluster and machine roles this guide set builds on.
- Security model — the trust boundary a control-plane Job runs inside.
- The Kinds — how
TerraformClusterandTerraformMachinemap to Cluster API objects.
KubeadmControlPlane
What KubeadmControlPlane (KCP) and the kubeadm bootstrap provider (CABPK) require of your CAPTF cluster and machine modules, and why this is the case.
Citations below prefixed capi/ are files in the CAPI v1.14.2 tree
(https://github.com/kubernetes-sigs/cluster-api/tree/v1.14.2). See also
the shared creation sequence
and the requirements checklist for every requirement in
one place.
What KubeadmControlPlane reads from your modules
Cluster.spec.controlPlaneEndpoint. KCP has no endpoint field of its own; it only reads the Cluster’s. It gates all Machine creation on this field being valid (host != ""andport != 0) — see networking and load balancer. The endpoint can come from the user, yourTerraformCluster(copied once while the Cluster field is not yet valid), or a control-plane provider; KCP never supplies one itself (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:316,439;capi/core/reconcilers/cluster/cluster_controller_phases.go:210-238).Machine.status.addresses. Not read by KCP or CABPK at all. Only the core Machine controller copies it from the InfraMachine; apiserver certificate SANs come from kubeadm’s own node-IP detection, not from CAPI addresses (capi/core/reconcilers/machine/machine_controller_phases.go:339-346).providerID/nodeRef. The Machine controller links a Node to a Machine only on an exactproviderIDmatch (see../contract/v1alpha1/machine.md“Node providerID matching”). KCP then gates almost everything on the resultingnodeRef: etcd-member matching, static-pod health, and scale-up/scale-down preflight all require every existing control-plane Machine to have one (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/preflight.go:193-203;capi/controlplane/kubeadm/pkg/workload_cluster_conditions.go:136,390,434).- Failure domains. KCP only sees
Cluster.status.failureDomainsentries withcontrolPlane == true; a nilcontrolPlanecounts asfalse. This is why the cluster module’sfailure_domains[].control_planedefault oftruematters (capi/controlplane/kubeadm/pkg/control_plane.go:171-184). cluster_network.api_server_port. Not read by CABPK or KCP at all in CAPI v1.14.2. The real kube-apiserver bind port is the KubeadmConfiglocalAPIEndpoint.bindPort(default 6443) (capi/controlplane/kubeadm/pkg/workload_cluster.go:281-291). Treatapi_server_port ?? 6443as a documentation-only convention, not a wired-through value (see networking and load balancer).- InfraMachineTemplate rotation. KCP triggers a rollout when a
Machine’s InfraMachine carries a
cluster.x-k8s.io/cloned-from-name/-groupkindannotation different from the currentmachineTemplate.spec.infrastructureRef. Template content is never diffed — CAPTF’s template immutability already matches this expectation (capi/controlplane/kubeadm/pkg/filters.go:183-227).
What KCP and CABPK write
- Bootstrap Secret. Named after the KubeadmConfig (== the Machine
name). Keys
valueandformatare always both written;formatdefaults tocloud-config(Ignition requires the alpha feature gateKubeadmBootstrapFormatIgnition, off by default) (capi/bootstrap/kubeadm/reconcilers/kubeadmconfig/kubeadmconfig_controller.go:1405-1437;capi/api/bootstrap/kubeadm/v1beta2/kubeadmconfig_types.go:27-36,117-120;capi/feature/feature.go:108). See../contract/v1alpha1/machine.mdbootstrap_formatfor how CAPTF surfaces this. - Payload contents. Control-plane init and join payloads embed the
cluster CA, etcd CA, service-account and front-proxy key material as
files, uncompressed
(
capi/bootstrap/kubeadm/pkg/cloudinit/controlplane_init.go:63,controlplane_join.go:61). See module checklist for what this means for your module. - Labels and hooks. Every KCP Machine, InfraMachine and KubeadmConfig
gets
cluster.x-k8s.io/cluster-name,cluster.x-k8s.io/control-plane: ""and apre-terminate.delete.hook.machine.cluster.x-k8s.io/kcp-cleanup: ""annotation (capi/controlplane/kubeadm/pkg/desiredstate/desired_state.go:324-340,110-135). CAPTF’s own metadata-update path already allows these (label/annotation edits are excluded from TerraformMachine’s immutability checks). KubeadmControlPlane.status.initialization.controlPlaneInitialized. Latchestrueonce thekubeadm-configConfigMap is readable through the endpoint — the reason the endpoint’s hairpin reachability matters (see networking and load balancer) (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/status.go:176-210). This mirrors intoCluster.status.initialization.controlPlaneInitializedand drives one cluster re-apply (capi/core/reconcilers/cluster/cluster_controller_status.go:80-81,450-551).- Conditions. KCP writes
Initialized,Available,CertificatesAvailable,EtcdClusterHealthy,ControlPlaneComponentsHealthy,MachinesReady,MachinesUpToDate,RollingOut,ScalingUp/ScalingDown,Remediating,Deleting, plus per-Machine pod/etcd-member conditions (capi/api/controlplane/kubeadm/v1beta2/kubeadm_control_plane_types.go:74-421,504-517). None of these read anything from your module beyond what is described above.
Lifecycle constraints
- Init ordering. KCP creates nothing until the Cluster is both
infrastructureProvisionedand has a valid endpoint; CABPK waits on the same gate. The first control-plane Machine gets init data; every other Machine (control-plane or worker) waits onCluster.status.initialization.controlPlaneInitialized(capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:316,413,439;capi/bootstrap/kubeadm/reconcilers/kubeadmconfig/kubeadmconfig_controller.go:289-298,352-368,466-495). - Join is serialized. Scale-up preflight requires certificates
available, no Machine deleting, and every existing control-plane
Machine already healthy with a
nodeRef. A failed preflight requeues after 15s, so control-plane joins are serialized on your apply time plus kubeadm join time (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/scale.go:60-97;preflight.go:61-218). - Rollout.
maxSurgeis 0 or 1 (default 1): with 1, KCP scales up then down, one Machine at a time; with 0 it scales down first (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/update.go:103-120). Rotation is by creating a newTerraformMachineTemplatename; edits to an existing template’s content are never detected. - Deleting blocks everything. While any control-plane Machine has a
deletion timestamp, KCP pauses scale-up, scale-down, rollout and
remediation
(
capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/preflight.go:111-122). A slow or stuck machine-module destroy therefore freezes the whole control plane, not just that Machine. - Remediation gates and retries. Post-initialization remediation
needs more than one replica, no Machine still provisioning without a
Node, no Machine deleting, and preserved etcd quorum; one remediation
runs at a time
(
capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/remediation.go:138-313).remediation.maxRetrydefaults to unlimited andretryPeriodSecondsdefaults to0(retry immediately) (capi/api/controlplane/kubeadm/v1beta2/kubeadm_control_plane_types.go:631-673). See module checklist for the template defaults this motivates. - No built-in control-plane MHC. KCP does its own remediation and is
not configured through a MachineHealthCheck object. Ship a plain
MachineHealthCheck selecting
cluster.x-k8s.io/control-plane: ""(or a ClusterClasscontrolPlane.healthCheck) if you want MHC-driven remediation in addition to KCP’s own (capi/docs/book/src/tasks/automated-machine-management/healthchecking.md:76-99). - Version skew.
KubeadmControlPlane.spec.versionmust be valid semver with avprefix, and an update may move at most one minor version (admission webhook,capi/controlplane/kubeadm/webhooks/admission/kubeadmcontrolplane.go:322-328,619-662). - Cluster deletion. Workers and MachinePools are deleted first; once
only control-plane Machines remain, KCP deletes all of them in
parallel, with no drain
(
capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:720-813). Expect concurrent destroy Jobs on cluster teardown. - Failed create cleanup. KCP creates the InfraMachine before the
Machine object. If KubeadmConfig or Machine creation then fails, KCP
deletes the InfraMachine directly, with no Machine owner reference at
all
(
capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/helpers.go:160-219). CAPTF’s delete webhook already allows a delete with no Machine owner reference for exactly this reason (see../contract/v1alpha1/machine.md“Delete”) — nothing further is required of your module.
Networking and load balancer
- Endpoint before any Machine. KCP creates no Machine at all until
Cluster.spec.controlPlaneEndpoint.IsValid()(capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:438-457). - Frontend port.
control_plane_endpoint.port, any value 1-65535. - Backend port. The kube-apiserver bind port, i.e. KubeadmConfig
localAPIEndpoint.bindPort(default 6443). CABPK/KCP never readCluster.spec.clusterNetwork.apiServerPort, so treatapi_server_port ?? 6443as a convention your templates must keep in sync withbindPortyourselves (capi/controlplane/kubeadm/pkg/workload_cluster.go:281-291). - Reachability, including hairpin. The endpoint must be reachable
from the management cluster, from every control-plane and worker node
(join discovery, kubelet), and from the first control-plane Machine
itself. This last point is the hairpin requirement: kubeadm’s
getKubeConfigSpecsBasesets the server URL ofadmin.conf,super-admin.confandkubelet.confto the control-plane endpoint, not the node’s own local address, so the node that is itself an LB backend must be able to reach the frontend it was just registered behind (see../contract/v1alpha1/cluster.md“Hairpin reachability”, citing kubeadm’scmd/kubeadm/app/phases/kubeconfig/kubeconfig.go~L582-626, kubernetes/kubernetes release-1.33). - Backend membership. Your machine module MUST register (and on
destroy deregister) each control-plane instance in the LB backend from
its own state
(
capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645;../contract/v1alpha1/machine.md“Control-plane machines”). The instance MUST be in the backend beforekubeadm init/kubeadm joinfinishes on it, orcontrolPlaneInitializednever latches (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/status.go:195-207). - Health check. Either plain TCP, or HTTPS
/readyz//healthzwithout certificate verification, works; the check must be able to go green with a single backend during init. CAPD’s reference LB uses haproxyoption httpchk GET /healthzwithcheck-ssl verify noneagainst backend 6443 (capi/test/infrastructure/docker/internal/loadbalancer/config.go:81-83). - Security groups / firewall. Backend port (bindPort) from the LB,
from all nodes, and from the management cluster; TCP 2379-2380 and
10250 between control-plane nodes; 10250 from control-plane nodes to
workers. KCP itself reaches etcd only through an apiserver pod
port-forward, so the management cluster needs no direct etcd or kubelet
port
(
capi/controlplane/kubeadm/pkg/etcd_client_generator.go:53;capi/controlplane/kubeadm/pkg/proxy/dial.go:72-120).
Module checklist
Cluster module
- MUST output a valid
control_plane_endpoint(host and port) no later than the apply that makes the clusterprovisioned, when no user endpoint is given — there is no other source with KCP (capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:316,439; see what KubeadmControlPlane reads).tfcapi-lint’soutput/endpoint-never-setcheck warns on a module that never emits one (see cluster.md’sEndpointAvailablecondition and the tfcapi-lint CLI reference). - MUST keep the emitted
control_plane_endpointstable for the life of the object: CAPI never updatesCluster.spec.controlPlaneEndpointafter the first valid copy (capi/core/reconcilers/cluster/cluster_controller_phases.go:217-228), so replacing the LB behind it (new DNS name or IP) breaks every kubeconfig and the API server certificate SANs. The controller neither detects nor prevents such a replacement;lifecycle { prevent_destroy = true }on the resource behind the endpoint turns it into a failed Job instead of a lost cluster (../contract/v1alpha1/cluster.mdcontrol_plane_endpointoutput). - MUST create a stable LB or VIP reachable from the management cluster, all nodes, and the control-plane nodes themselves (hairpin, see networking and load balancer).
- MUST use a frontend on
control_plane_endpoint.portand a backend onapi_server_port ?? 6443, kept equal to the KubeadmConfigbindPortyour templates configure (capi/controlplane/kubeadm/pkg/workload_cluster.go:281-291; see networking and load balancer). - SHOULD publish the LB target/backend-pool id in
exportsfor machine modules to register against — a CAPTF convention built on the contract’sexportsfield, not a CAPI-mandated shape (../contract/v1alpha1/cluster.mdexports). - SHOULD leave
failure_domains[].control_planeat its defaulttrue(capi/controlplane/kubeadm/pkg/control_plane.go:171-184; see what KubeadmControlPlane reads); a nil value is invisible to KCP.
Machine module
- MUST register the control-plane instance in the LB target/backend-pool
from
captf_cluster_outputs, in the module’s own state, beforekubeadm init/joinfinishes on it (capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645;capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/status.go:195-207; see networking and load balancer and one of the three patterns in README.md’s LB-membership patterns). - MUST accept
bootstrap_dataopaquely as base64 and MUST NOT need to parse it, per../contract/v1alpha1/machine.mdbootstrap_data. - Control-plane payloads embed cluster CA and service-account private key
material and are not compressed by CABPK (see what KCP and CABPK
write). Payload size limits (for example a
cloud’s user-data cap) and keeping key material out of readable
instance metadata are the module’s responsibility: gzip the payload,
or stage it in a secret store and pass only a small stub through
bootstrap_data/user-data. - MUST emit
provider_idexactly matching the Node’sspec.providerID; KCP gates join preflight, etcd-member matching and remediation on the resultingnodeRef(see what KubeadmControlPlane reads). addressesare informational only for KCP — no KCP or CABPK code reads them.- MUST make
failure_domainoutput equal thefailure_domaininput when the input is non-null (../contract/v1alpha1/machine.mdfailure_domainoutput).
Templates and health checks
- Bring-up time for N control-plane replicas is approximately
N × (apply + boot + join), because every join is serialized on the
previous Machine’s
nodeRef(see lifecycle constraints). A rollout adds (apply + join + destroy) per replica. - Ship
KubeadmControlPlane.spec.remediation.maxRetry(e.g.3) and a non-zeroretryPeriodSecondsin your templates: the unbounded defaults turn a control-plane module whose apply always fails into an endless create/destroy loop of cloud resources (see lifecycle constraints). - Ship a control-plane MachineHealthCheck (selector
cluster.x-k8s.io/control-plane: "", or a ClusterClasscontrolPlane.healthCheck). ItsInfrastructureReady=Falsetimeout MUST exceed apply time plus one drift interval — the same rule as any other machine module (../contract/v1alpha1/machine.md“Health → remediation”).
Open questions
- Bootstrap payload size and secrecy. The contract leaves the gzip/stub strategy to the module (see module checklist); there is no CAPTF-side size or secrecy check today.
- Unbounded remediation retry. Combined with a deterministically
failing module, KCP’s default unlimited
maxRetrycauses cloud-resource churn. The module checklist’s template-default recommendation is the only mitigation; there is no CAPTF-side guard against it.
See also
README.md— the shared creation sequence and the LB-membership patterns.rke2.md— the same requirements under RKE2ControlPlane.checklist.md— every requirement from both guides in one place.- The machine role — the normative contract this guide builds on.
RKE2ControlPlane
What RKE2ControlPlane (RCP) and the RKE2 bootstrap provider (CAPRKE2) require of your CAPTF cluster and machine modules, and why this is the case.
Citations below prefixed rke2/ are paths in
rancher/cluster-api-provider-rke2 at tag v0.25.2; rke2docs/ is
rancher/rke2-docs (main branch); capi/ is the CAPI v1.14.2 tree
(https://github.com/kubernetes-sigs/cluster-api/tree/v1.14.2).
See also the shared creation sequence
and the requirements checklist for every requirement in
one place.
Version and compatibility
CAPRKE2 v0.25.2 is built against sigs.k8s.io/cluster-api v1.13.5
(rke2/go.mod:33) and its v1beta2 API is the contract version CAPTF also
targets. Its own e2e suite runs only against CAPI core v1.12.11 and
v1.13.5 (rke2/test/e2e/config/e2e_conf.yaml:22,33), and its
getting-started guide pins CAPI core to v1.13.5
(rke2/docs/book/src/01_user/01_getting-started.md:56).
CAPTF runs CAPI v1.14.2. The contract shape is unchanged, but CAPI v1.14 core is untested by CAPRKE2 v0.25.2 — run a compatibility smoke test before relying on it in production; see open questions.
What RKE2ControlPlane reads from your modules
-
Cluster.status.initialization.infrastructureProvisioned. RCP creates nothing — not even certificates — until this istrue; RKE2Config generates no bootstrap data until it istrueeither (rke2/controlplane/internal/controllers/rke2controlplane_controller.go:376-394;rke2/bootstrap/internal/controllers/rke2config_controller.go:171-182). -
Cluster.spec.controlPlaneEndpoint. A hard gate: RCP creates no Machine at all — including the first — while the endpoint is not valid (host and port both set) (rke2/controlplane/internal/controllers/rke2controlplane_controller.go:421-440). Only the endpoint’s host feeds the RKE2tls-sanlist and the init node’sServerURL(https://<host>:9345); the endpoint’s port is ignored for the supervisor join URL, which is always 9345 (rke2/pkg/rke2/config.go:350;rke2/bootstrap/internal/controllers/rke2config_controller.go:66,455-473). -
Cluster.spec.clusterNetwork.pods/services.cidrBlocksmap tocluster-cidr/service-cidrwhen non-empty (rke2/pkg/rke2/config.go:209-215).serviceDomainandapi_server_portare not read at all: cluster domain comes from RKE2’s ownserverConfig.clusterDomain, and the API server is always on 6443 on the node (rke2/pkg/rke2/config.go:224-225). Do not derive an LB backend port fromapi_server_portunder RKE2 beyond using it as the frontend port if your templates choose to. -
Failure domains. Same rule as KCP: only entries with
controlPlane: trueare used for scale-up/scale-down placement; a nilcontrolPlanecounts asfalse(rke2/pkg/rke2/control_plane.go:176-190). -
Machine.status.addresses. Read only for three of the fiveregistrationMethodvalues, and only from Machines whoseReadycondition is alreadyTrue(theregistrationMethodtable below;rke2/controlplane/internal/controllers/status.go:80,152,158-177). The contract’s canonical output order foraddresses(InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname — see../contract/v1alpha1/machine.mdaddresses) is what makes RCP’s first-match logic deterministic; your module’s job is only to emit the address type(s) each method requires (below), not to order them — the controller does that. -
Machine.status.nodeRef. Used for etcd leader-move/member-removal ordering and remediation;HasHealthyMachineStillProvisioning= a healthy Machine with no Node yet (rke2/pkg/rke2/workload_cluster_etcd.go:40-45;rke2/pkg/rke2/control_plane.go:527-529). -
registrationMethod(immutable once set, enforced by webhook):Method status.availableServerIPs=Address types your module must supply Join URL control-plane-endpoint(default,""≡ this)[endpoint.host]none https://<endpoint host>:9345address[spec.registrationAddress]none https://<registrationAddress>:9345internal-firstfirst InternalIPorExternalIPper Ready CP Machine, in list orderInternalIPand/orExternalIPhttps://<ip>:9345internal-only-ipsfirst InternalIPper Ready CP MachineInternalIPhttps://<ip>:9345external-only-ipsfirst ExternalIPper Ready CP MachineExternalIPhttps://<ip>:9345(
rke2/pkg/registration/registration.go:46,68-140;rke2/controlplane/api/v1beta2/rke2controlplane_types.go:97-100;rke2/controlplane/api/v1beta2/rke2controlplane_webhook.go:142-145,211-214.) A Ready Machine with no address of the required type produces the RCP status error “ready but they have no IP Address available” (rke2/controlplane/internal/controllers/status.go:175-177); see open questions. -
RKE2ControlPlane.spec.version. Must match(v\d\.\d{2}\.\d+\+rke2r\d)|^$and is copied verbatim toMachine.spec.version(rke2/controlplane/api/v1beta2/rke2controlplane_types.go:79-81;rke2/controlplane/internal/controllers/scale.go:566-567,612). See../contract/v1alpha1/machine.mdkubernetes_versionand the module checklist.
What RCP and RKE2Config write
- Bootstrap Secret. Name = RKE2Config name; keys
valueandformatare always both written (there is no case whereformatis absent);formatdefaults tocloud-config(rke2/bootstrap/internal/controllers/rke2config_controller.go:1040-1061;rke2/bootstrap/api/v1beta2/rke2config_webhook.go:79-80). See../contract/v1alpha1/machine.mdbootstrap_format. - Ignition is supported end to end (init, CP join, worker), via
Butane → Ignition 3.3
(
rke2/bootstrap/internal/controllers/rke2config_controller.go:552-556,798-802,927-931). gzipUserData: true. For cloud-config,valuebecomes raw gzip bytes (not base64, not UTF-8 text);formatstayscloud-config. For Ignition, the gzip is wrapped inside the Ignition config instead (rke2/bootstrap/internal/controllers/rke2config_controller.go:1015-1037). See machine.md’sbootstrap_datainput for how the contract resolves the binary-payload question, and the module checklist for what it means for your module.- Cluster Secrets.
<cluster>-ca,-cca(client CA),-peer-etcd,-etcd,-kubeconfig,-token. No-sa, no-proxy(contrast with KCP’s secret set, kubeadm.md’s what KCP and CABPK write) (rke2/pkg/secret/certificates.go:52-80,168-190,382-384). - Machine labels/annotations. Template labels plus forced
cluster.x-k8s.io/cluster-name,cluster.x-k8s.io/control-plane: ""; apre-terminate.delete.hook.machine.cluster.x-k8s.io/rke2-cleanupannotation on every control-plane Machine (rke2/controlplane/internal/controllers/scale.go:650-666,577-578). - RCP status (v1beta2).
initialization.controlPlaneInitialized(true once the workloadkube-system/rke2-servingSecret exists, or once every owned Machine is Ready);availableServerIPs; conditions includingEtcdClusterHealthy,ControlPlaneComponentsHealthy,Remediating(rke2/controlplane/internal/controllers/status.go:98-194;rke2/controlplane/api/v1beta2/rke2controlplane_types.go:262-316). - Node identity. CAPRKE2 never sets kubelet
provider-id,node-ipornode-nameitself; the fields exist in its config but nothing in the repo writes them (rke2/pkg/rke2/config.go:501-503). See the module checklist and machine.md’s node providerID matching — RKE2.
Lifecycle constraints
- Init. Exactly one control-plane Machine gets init data, guarded by
an init-lock ConfigMap. The init node has no
server:configured — it creates the cluster standalone and, notably, needs no LB/9345 reachability to itself during init, unlike the hairpin requirement that applies to every later join (rke2/controlplane/internal/controllers/scale.go:51-99;rke2/pkg/rke2/config.go:650-676). Don’t read this as weakening the general endpoint-reachability requirement in networking and load balancer: it applies to every Machine that joins, i.e. every control-plane Machine after the first, and to every worker. - Join (control-plane and worker), in order. Requires the Cluster’s
ControlPlaneInitializedcondition; the<cluster>-tokenSecret; andRCP.status.availableServerIPsnon-empty, which itself needs at least one Ready control-plane Machine and a reachable workload API server via the<cluster>-kubeconfigSecret (rke2/bootstrap/internal/controllers/rke2config_controller.go:230,699-703,852-856;rke2/controlplane/internal/controllers/status.go:114-146,152). Both CP joins and worker joins useavailableServerIPs[0](rke2/bootstrap/internal/controllers/rke2config_controller.go:717,867). - Rollout.
maxSurge0 or 1, default 1; scale-up then scale-down of the oldest outdated Machine in the most-populated failure domain (rke2/controlplane/api/v1beta2/rke2controlplane_types.go:548-580;rke2/controlplane/internal/controllers/scale.go:285-334). - Preflight (scale up/down, in-place). No Machine deleting; every
control-plane Machine must have
AgentHealthyandEtcdMemberHealthybothTrue(Unknownor missing blocks); a missing Node blocks by extension (rke2/controlplane/internal/controllers/scale.go:200-266). - Scale-down / deletion. Etcd leadership is forwarded to the newest
Machine, then the Machine is deleted; the
rke2-cleanuppre-terminate hook re-forwards leadership, annotates the Node for etcd removal, and waits for confirmation before releasing — InfraMachine deletion happens only after that (rke2/controlplane/internal/controllers/scale.go:141-197). - Remediation. RCP remediates control-plane Machines with
HealthCheckSucceeded=False+OwnerRemediated=False(MHC-driven). Pre-init remediation is allowed directly; post-init remediation additionally requires more than one replica, no Machine still provisioning without a Node, no Machine deleting, and preserved etcd quorum (rke2/controlplane/internal/controllers/remediation.go:97-330,366-420). Machine.md’s health → remediation section already reflects this: remediation happens for Machines whose owner acts on the MachineHealthCheck’sMachineOwnerRemediatedcondition, and RCP is such an owner for its own control-plane Machines, not only MachineSet and KubeadmControlPlane. - In-place updates. Behind the CAPI
InPlaceUpdatesfeature gate (alpha) plus exactly one registeredCanUpdateMachineextension; otherwise falls back to delete/recreate. Not something a CAPTF module needs to support, sinceTerraformMachine.spec.sourceis immutable regardless (rke2/controlplane/main.go:287). - Payload size / air-gap. Init control-plane user-data embeds four CA
key pairs plus config and optional manifest/registry/audit files.
Non-air-gapped installs
curl -sfL https://get.rke2.ioat boot, which means node internet egress;airGapped: trueexpects pre-baked artifacts in the image instead (rke2/bootstrap/internal/cloudinit/controlplane_init.go:34-36).gzipUserDataexists for size-limited clouds; see machine.md’sbootstrap_datainput for how the contract handles the resulting binary payload.
Networking and load balancer
| Path | Port | Needed when |
|---|---|---|
| LB/VIP → CP nodes | 6443/TCP (endpoint port → node 6443; RKE2 ignores apiServerPort) | always |
| LB/VIP → CP nodes | 9345/TCP supervisor, same host as the endpoint | control-plane-endpoint (default) and address registration methods |
| All nodes → CP nodes | 6443, 9345 TCP direct (after registration) | always |
| CP ↔ CP | 2379, 2380, 2381 TCP | embedded etcd (not with externalDatastoreSecret) |
| all ↔ all | 10250 TCP; NodePort range (default 30000-32767) | always |
| CNI (default canal) | 8472/UDP VXLAN, 9099/TCP | cni: canal or unset |
(rke2/controlplane/internal/controllers/rke2controlplane_controller.go:757;
rke2docs/docs/install/ha.md:42; rke2docs/docs/install/requirements.md:128,142-148,158-161.)
- No LB listener on 2379 is needed. The controller reaches etcd only
by port-forwarding through the API server; an LB entry for 2379 is not
required by anything in RCP
(
rke2/pkg/proxy/dial.go:99-106;rke2/pkg/etcd/client_generator.go:34). - Health checks. 9345: TLS
GET /v1-rke2/readyzexpecting 403 (unauthenticated) — or plain TCP. 6443: TCP, or HTTPS/healthzonly if anonymous auth is enabled (rke2/examples/templates/docker/cluster-template.yaml:60,176-194). - DNS endpoints are supported. RKE2 HA supports a DNS name /
round-robin DNS as the fixed registration address; the endpoint host is
added to
tls-sanautomatically.registrationAddressis not added totls-sanautomatically — add it viaserverConfig.tlsSanif it differs from the endpoint host (rke2docs/docs/install/ha.md:36-37;rke2/pkg/rke2/config.go:350). - Backend membership. The module owns LB membership — nothing in RCP or CAPI registers a backend for you. The instance MUST be added before its Machine is Ready (joins happen through the LB during bring-up), and the LB MUST tolerate a backend whose supervisor is not yet up (see lifecycle constraints, “Join”).
Module checklist
Cluster module
- MUST output a valid
control_plane_endpoint(host and port) no later than the apply that makes the clusterprovisioned, unless the user sets one — RCP creates no Machine at all untilCluster.spec.controlPlaneEndpoint.IsValid()(rke2/controlplane/internal/controllers/rke2controlplane_controller.go:421-440; see what RKE2ControlPlane reads). This matches the contract’s own endpoint-timing rule in../contract/v1alpha1/cluster.mdcontrol_plane_endpointoutput. - MUST keep the emitted
control_plane_endpointstable for the life of the object: CAPI never updatesCluster.spec.controlPlaneEndpointafter the first valid copy (capi/core/reconcilers/cluster/cluster_controller_phases.go:217-228), so replacing the LB behind it (new DNS name or IP) breaks every kubeconfig and the API server certificate SANs. The controller neither detects nor prevents such a replacement;lifecycle { prevent_destroy = true }on the resource behind the endpoint turns it into a failed Job instead of a lost cluster (../contract/v1alpha1/cluster.mdcontrol_plane_endpointoutput). - MUST provision two LB listeners on the same host:
endpoint.port→ CP nodes:6443, and:9345→ CP nodes:9345— not just one, as with KCP (rke2docs/docs/install/ha.md:42; see networking and load balancer). - MUST NOT rely on
cluster_network.api_server_portorservice_domainto configure RKE2; RKE2 reads neither (rke2/pkg/rke2/config.go:224-225; see what RKE2ControlPlane reads). - Backend membership is the module’s job (not CAPI’s): a target group by tag/label, or an explicit attach resource — see the three patterns in README.md’s LB-membership patterns.
- SHOULD leave
failure_domains[].control_planeat its defaulttrue(rke2/pkg/rke2/control_plane.go:176-190; see what RKE2ControlPlane reads), same as under KCP.
Machine module
- MUST accept both
bootstrap_formatvalues RCP can write,cloud-configandignition(rke2/bootstrap/internal/controllers/rke2config_controller.go:1040-1061; see what RCP and RKE2Config write). - MUST register the control-plane instance in both the 6443 and 9345 LB
target sets, and open/attach the corresponding security rules, before
the Machine becomes Ready
(
capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645; see networking and load balancer). - MUST supply
addressesof the type(s) the cluster’sregistrationMethodneeds (theregistrationMethodtable in what RKE2ControlPlane reads) when that method is anything other thancontrol-plane-endpoint/address; the controller, not the module, is responsible for output ordering (../contract/v1alpha1/machine.mdaddresses). - MUST emit
provider_idmatching exactly what the Node’s kubelet registers — either via a cloud-controller-manager, or viaagentConfig.kubelet.extraArgs: [provider-id=<value>]set from instance metadata, since CAPRKE2 sets none of this itself (see what RCP and RKE2Config write). Under RKE2 a mismatch on the first control-plane Machine blocks every subsequent join, not just that Machine’s own readiness, because every join needsavailableServerIPsnon-empty, which needs a Ready CP Machine (see what RKE2ControlPlane reads and lifecycle constraints). kubernetes_versionarrives asvX.Y.Z+rke2rN; strip the+rke2rNsuffix before using it for an image lookup or a semver comparison (see what RKE2ControlPlane reads).- Init/CP-join payloads embed four CA key pairs; on size-limited clouds
both
cloud-config+gzipUserData: trueandignition+gzipUserData: truework for reducing payload size (see what RCP and RKE2Config write). Either way the module MUST routebootstrap_datato a base64-taking argument (e.g.user_data_base64); it MUST NOTbase64decode()it, because a gzipped payload is not valid UTF-8.
Differences from KubeadmControlPlane
| Aspect | KubeadmControlPlane | RKE2ControlPlane v0.25.2 |
|---|---|---|
| LB ports | API port only | API (6443) and supervisor 9345, same host |
| Join target | CP endpoint, backend port | availableServerIPs[0]:9345 — endpoint host, registrationAddress, or a CP Machine IP |
Uses Machine.status.addresses | no | yes, for 3 of 5 registrationMethod values |
| Join gate | Cluster ControlPlaneInitialized | that, plus ≥1 Ready CP Machine (needs Node ⇒ providerID match) |
Machine.spec.version | vX.Y.Z | vX.Y.Z+rke2rN |
clusterNetwork.apiServerPort | not wired to bindPort either, but the field exists as a convention (see kubeadm.md’s what it reads) | ignored entirely; 6443 is fixed |
clusterNetwork.serviceDomain | honored by CABPK | ignored (serverConfig.clusterDomain) |
Bootstrap format | always present (cloud-config or ignition) | always present |
| Secrets | -ca, -etcd, -sa, -proxy, -kubeconfig | -ca, -cca, -etcd, -peer-etcd, -kubeconfig, -token |
| etcd removal | KCP pre-terminate hook, etcd client | rke2-cleanup pre-terminate hook, Node annotation etcd.rke2.cattle.io/remove |
| Install | kubeadm/kubelet pre-installed in image | curl get.rke2.io at boot unless air-gapped artifacts are baked in |
| Contract | v1beta2, built on CAPI v1.14 | v1beta2, built on CAPI v1.13.5 |
(Individual rows cited in the corresponding sections above and in
kubeadm.md.)
Open questions
- No
InternalIPguarantee.addressesmay legally be[](../contract/v1alpha1/machine.mdaddresses). Forinternal-only-ips/external-only-ips/internal-first, a Ready Machine without the right address type blocks every subsequent join with an RCP status error (see what RKE2ControlPlane reads). Notfcapi-lintrule warns when a module never emits anInternalIP. - CAPI v1.14 compatibility. CAPRKE2 v0.25.2 is tested only up to CAPI core v1.13.5 (see version and compatibility). CAPTF runs v1.14.2. Run a compatibility smoke test against v1.14.2 before relying on RKE2ControlPlane in production; watch for a CAPRKE2 release built on v1.14.
- Nondeterministic join target. Under IP-based
registrationMethodvalues,availableServerIPsis built by iterating a Go map, so which control-plane node a joiner targets is nondeterministic, and a stale or just-removed node’s address can be chosen until the next status refresh (see what RKE2ControlPlane reads;rke2/bootstrap/internal/controllers/rke2config_controller.go:717,867). A module cannot fix this directly; a Machine that leavesReady(for example through a MachineHealthCheck) drops out ofavailableServerIPson the next status refresh.
See also
README.md— the shared creation sequence and the LB-membership patterns.kubeadm.md— the same requirements under KubeadmControlPlane.checklist.md— every requirement from both guides in one place.- The machine role — the normative contract this guide builds on.
Requirements Checklist
Every port, health check and ordering rule a KubeadmControlPlane (KCP) or
RKE2ControlPlane (RCP) cluster needs from a CAPTF cluster or machine
module, grouped by topic. Details and citations are in
kubeadm.md and rke2.md; the shared creation
sequence and LB-membership patterns are in README.md.
n/a means the row does not apply to that provider or module role.
“Verified by” is one of:
lint:tfcapi-lintcan check it today.e2e: only an end-to-end run with a real control-plane provider catches a violation.operator: checked by whoever reviews or writes the module; no automated check exists.none: not checked anywhere today.
Some lint-worthy rules have no check today; those are called out in
kubeadm.md’s open questions and
rke2.md’s open questions. The project has no
end-to-end suite today, so no e2e row below has been verified against
a real control plane.
Endpoint and load balancer
| Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by |
|---|---|---|---|---|---|
| Endpoint valid before any Machine | KCP creates no Machine until Cluster.spec.controlPlaneEndpoint.IsValid() (host and port both set) | RCP returns before creating any Machine, including the first, while !IsValid() | MUST output a valid control_plane_endpoint no later than the apply that makes the cluster provisioned, unless the user sets one | n/a | e2e |
| LB frontend → CP backend port | endpoint.port → CP :bindPort (default 6443); one frontend | endpoint.port → CP :6443 (apiServerPort ignored); one of two frontends | MUST create a frontend on endpoint.port and a backend on api_server_port ?? 6443, kept equal to the module’s bindPort/RKE2’s fixed 6443 | n/a | operator |
| Second LB listener on 9345 | not applicable | :9345 → CP :9345, same host as the endpoint; required for control-plane-endpoint (default) and address registration methods | MUST create this second listener for RKE2 clusters | MUST register the instance in the 9345 target set alongside 6443 | operator |
| Hairpin reachability | The endpoint MUST be reachable from the first CP node itself — admin.conf/super-admin.conf/kubelet.conf all point at the endpoint, not the node’s local address | Same MUST applies to every RKE2 join. The init node itself has no server: configured and so needs no LB/9345 reachability at init time (rke2.md’s lifecycle constraints) — this does not weaken the MUST for every later join | MUST make the LB/VIP allow a backend to reach its own frontend | n/a | none |
| Endpoint stability | Once emitted, control_plane_endpoint MUST be stable for the life of the object — CAPI never updates Cluster.spec.controlPlaneEndpoint after the first valid copy | Same rule; RCP also builds tls-san and the kubeconfig Secret from the first copy | MUST NOT replace the LB behind an already-emitted endpoint (new DNS name or IP); consider lifecycle { prevent_destroy = true } on that resource | n/a | none |
Health checks and backend membership
| Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by |
|---|---|---|---|---|---|
| Health check — API port | TCP, or HTTPS /readyz//healthz without certificate verification; MUST go green with a single backend during init | TCP, or HTTPS /healthz only if anonymous auth is enabled | MUST configure a check matching one of these | n/a | operator |
| Health check — RKE2 supervisor (9345) | not applicable | TLS GET /v1-rke2/readyz expecting 403 (unauthenticated), or plain TCP | MUST configure this check for RKE2 clusters | n/a | operator |
| Backend membership timing | Instance MUST be in the LB backend before kubeadm init/kubeadm join finishes on it, or controlPlaneInitialized never latches | Instance MUST be added before its Machine becomes Ready; the LB MUST tolerate a backend whose supervisor is not yet up | n/a (registration is the machine module’s job) | MUST register/deregister the instance in its own Terraform state, in the same apply as instance creation | e2e |
| SG / firewall — control plane | Backend port (bindPort) from LB, nodes and management; TCP 2379-2380 and 10250 between CP nodes; 10250 from CP to workers | CP↔CP TCP 2379-2381; all→CP TCP 6443 and 9345; all↔all TCP 10250 and the NodePort range (default 30000-32767); CNI ports for the chosen CNI (default canal: 8472/udp, 9099/tcp); egress to get.rke2.io/GitHub releases unless air-gapped | MUST create these security groups/firewall rules | MUST attach the instance to them | operator |
Node identity and addresses
| Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by |
|---|---|---|---|---|---|
provider_id == Node providerID | MUST match exactly — the Machine controller links Node to Machine only on an exact match; KCP gates join preflight, etcd matching and remediation on the resulting nodeRef | MUST match exactly; a mismatch on the first CP Machine blocks every later join (no Ready CP Machine ⇒ availableServerIPs stays empty) | n/a | MUST emit the same value the Node’s kubelet/CCM will register (CCM sets it, or kubeletExtraArgs/agentConfig.kubelet.extraArgs: [provider-id=...]) | e2e |
addresses types and order | Not read by KCP or CABPK at all | Required address types by registrationMethod: control-plane-endpoint/address — none; internal-first — InternalIP and/or ExternalIP; internal-only-ips — InternalIP; external-only-ips — ExternalIP. Order is fixed by the controller (InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname), not the module | n/a | MUST emit the required type(s) for the cluster’s registrationMethod; MUST NOT rely on emission order | none |
kubernetes_version suffix | Arrives as plain vX.Y.Z | Arrives as vX.Y.Z+rke2rN, copied verbatim from RKE2ControlPlane.spec.version | n/a | Under RKE2, MUST strip +rke2rN before using the value for an image lookup or a semver comparison | none |
Bootstrap payload
| Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by |
|---|---|---|---|---|---|
bootstrap_format always set | Always written by CABPK (cloud-config default, ignition behind the alpha feature gate) | Always written by CAPRKE2 (cloud-config default, ignition fully supported) | n/a | MUST accept both cloud-config and ignition | e2e |
bootstrap_data is base64, including the gzip case | Payload is always UTF-8 text (cloud-config/ignition), base64-encoded by the CAPTF controller like every other bootstrap provider | With gzipUserData: true, the decoded payload is raw (non-UTF-8) gzip bytes — still base64-encoded by the CAPTF controller, same as any other payload | n/a | MUST pass bootstrap_data to a base64-taking argument, or base64decode() it only when the content is known to be UTF-8 (never for a gzipped payload) | none |
| CP bootstrap payload size and secrecy | Init/join payload embeds cluster CA, etcd CA, service-account and front-proxy key material, uncompressed | Init payload embeds four CA key pairs, uncompressed unless gzipUserData: true | n/a | MUST gzip the payload or stage it in a secret store with a small stub, and MUST keep key material out of readable instance metadata | none |
Templates, timing and compatibility
| Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by |
|---|---|---|---|---|---|
failure_domains[].control_plane default | Only entries with controlPlane == true are visible to KCP; nil counts as false | Same rule for RCP | SHOULD leave the field at its default true | n/a | none |
| Control-plane bring-up timing | N replicas ≈ N × (apply + boot + join); joins strictly serialized on the previous Machine’s nodeRef | Same shape; a join additionally requires ≥1 already-Ready CP Machine (Node with matching providerID) | n/a | n/a (informs MHC timeout sizing) | none |
| Control-plane template defaults | Templates SHOULD set remediation.maxRetry (e.g. 3) and a non-zero retryPeriodSeconds (defaults are unlimited retries, immediate retry); the control-plane MachineHealthCheck’s InfrastructureReady=False timeout MUST exceed apply time plus one drift interval (no fixed number specified here) | RCP has its own remediationStrategy{maxRetry, retryPeriod, minHealthyPeriod}; no CAPTF-recommended value is given | n/a | n/a | operator |
| Control-plane provider / CAPI compatibility | Built for and tested against CAPI v1.14.2 (this provider’s target) | CAPRKE2 v0.25.2 is tested by its own e2e suite only against CAPI core v1.12.11 and v1.13.5; CAPTF runs v1.14.2. Run a compatibility smoke test before relying on RKE2ControlPlane in production | n/a | n/a | none |
See also
README.md— the shared creation sequence and the LB-membership patterns.kubeadm.mdandrke2.md— the full guidance and citations behind each row above.- The module contract — the normative cluster and machine roles both guides build on.
Installation
This page covers installing the CAPTF provider into a management cluster
with clusterctl: what to have ready first, how to register the provider,
what the install creates, how the runner image is set, the optional
components, and how to confirm the install worked. It is for whoever
administers the management cluster, not for cluster tenants.
Before you begin
- A management cluster with Cluster API’s core, bootstrap and
control-plane providers already initialized, and a
kubectlcontext pointing at it. clusterctl, matching the version documented for the Cluster API release you run.cert-manager, with thecert-manager.io/v1API available.clusterctl initinstalls a compatiblecert-managerrelease itself when one is not already present, so a separate install step is only needed to pin a specificcert-managerversion or to install it ahead of time.- Cluster API installed at a release that implements the same contract
CAPTF does. CAPTF’s own
metadata.yamllists one release series so far, contractv1beta2; CAPTF is built and tested against Cluster API v1.14.2.
Register the provider
CAPTF is not one of clusterctl’s built-in providers, so clusterctl needs
a config file naming its release manifest. The config entry’s name is
terraform:
# clusterctl.yaml
providers:
- name: terraform
type: InfrastructureProvider
url: https://github.com/captf-io/cluster-api-provider-terraform/releases/latest/infrastructure-components.yaml
url can also name a specific tag instead of latest, or a file:// path
into a local repository built from a release’s assets; see Installing from
a local repository
for the local repository layout and for pinning a version.
CAPTF has not published a release yet, so the URL above does not resolve to anything: install from a local repository until one exists.
Install with clusterctl init
clusterctl init --config clusterctl.yaml --infrastructure terraform
clusterctl init --infrastructure terraform:vX.Y.Z pins a specific
released version instead of the newest one clusterctl can see.
If you plan to use the ClusterClass flavor, enable the ClusterTopology
feature gate before running init — it is alpha in Cluster API and off by
default:
CLUSTER_TOPOLOGY=true clusterctl init --config clusterctl.yaml --infrastructure terraform
What gets installed
Everything below lands in one namespace, captf-system; clusterctl init
also installs cert-manager itself when it is missing, in its own
namespace.
- CRDs for the seven kinds:
TerraformCluster,TerraformClusterTemplate,TerraformMachine,TerraformMachineTemplate,TerraformMachinePool,TerraformMachinePoolTemplateandTerraformClusterIdentity(the last is cluster-scoped; the rest are namespaced). See The Kinds for what each one does, and API Reference for every field. - The manager, a single-replica
Deploymentnamedcaptf-controller-manager, running as a non-root user, with aServiceAccountof the same name. It serves its webhooks on:9443, its metrics on:8443, and its health and readiness probes on:9440. - The admission webhook: a
ValidatingWebhookConfigurationnamedcaptf-validating-webhook-configuration, one rule per kind, backed by thecaptf-webhook-serviceService. Its serving certificate is acert-managerCertificate(captf-serving-cert, issued by the self-signedIssuercaptf-selfsigned-issuer), mounted into the manager pod and kept current bycert-manager’s CA injection. - RBAC for the manager itself: the
ClusterRolecaptf-manager-roleand itsClusterRoleBinding, plus a leader-electionRoleandRoleBindingscoped tocaptf-system. The manager also creates a runnerServiceAccountandRoleBindingin each tenant namespace at first use, from the staticClusterRolecaptf-runner. See RBAC for what each role grants and for the per-namespace runner setup.
None of this installs a TerraformClusterIdentity or any Terraform*
object: those come from templates you apply afterward, covered in
Identities and Credentials and Templates
and ClusterClass.
The manager image and the runner image
clusterctl init sets the manager container’s image to the release’s
image, ghcr.io/captf-io/cluster-api-provider-terraform:vX.Y.Z. The same
reference is also set as the CAPTF_MANAGER_IMAGE environment variable on
that container: it is the default of --runner-image, the image the
manager runs as the init container that injects the runner binary into
every Job it creates. Point --runner-image at a different reference
(for example a registry mirror the tenant namespaces’ Jobs can reach, or a
build carrying a patched runner binary) when the manager’s own image is
not the one you want Jobs to pull; see --runner-image
for the flag and Job Environment for how the
init container uses it. spec.source.image on a Terraform* object is
unrelated: it names the role image the Job actually runs, never the
runner.
Optional components
config/prometheus and config/network-policy are kustomize components:
neither is part of infrastructure-components.yaml, so clusterctl init
never installs them and clusterctl upgrade never touches them, and
neither is a release asset — get them from a checkout of the tag you
installed (see Register the provider), not from
the release URL. A kustomize Component builds standalone: point a bare
kustomization at one with no resources: of your own, and it emits only
that component’s own objects, carrying the captf-/captf-system names
config/default produces:
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
components:
- <path-to-checkout>/config/prometheus
- Prometheus (
config/prometheus): a metricsService, aServiceMonitor, the alerting rules, and the RBAC Prometheus needs to scrape the manager’s authenticated metrics endpoint. See Enabling the Prometheus component for what it adds and how to wire it up, and how a later upgrade affects it. - Network policy (
config/network-policy): aNetworkPolicythat restricts the manager pod to webhook, metrics and health-probe traffic inbound, and DNS, the API server and container registries outbound. It only has an effect on a CNI that enforcesNetworkPolicy, and it covers only the manager pod, not runner Jobs;config/network-policy/job-egress-sample.yamlis a separate starting point to copy into each tenant namespace for its Job pods. See Network exposure for why the manager’s and a Job’s exposure differ.
Building both at once needs both under components: in the same
kustomization. Building either on top of a from-source install of
config/default (rather than the one clusterctl init already installed)
also works, but is unnecessary for these two components: each one’s
objects stand alone and never depend on config/default’s own resources
being built alongside them.
Verifying the install
kubectl -n captf-system rollout status deployment/captf-controller-manager
kubectl get crds -l cluster.x-k8s.io/provider=infrastructure-terraform
kubectl get validatingwebhookconfigurations captf-validating-webhook-configuration
kubectl -n captf-system get certificate captf-serving-cert
The rollout command returns once the manager pod is ready. The CRD list
should show all seven kinds. The webhook configuration and the certificate
must both exist and the certificate must report Ready=True, cert-manager
was able to issue the webhook’s serving certificate: without it, webhook
calls from the API server fail closed (failurePolicy: Fail) and every
Terraform* create or update is rejected.
clusterctl init itself prints the components it installed; clusterctl describe cluster (once you have created one) reports whether CAPTF’s
objects are ready.
See also
- Upgrades for upgrading an existing install.
- Configuration for the manager flags you set on top of this default install.
- RBAC for the manager’s and runner’s permissions in full.
- Security Model for the trust boundary a
Terraform*object’s Job operates inside.
Configuration
The CAPTF manager takes every setting as a command-line flag; there is no config file. This page explains what each group of flags changes and when to change it. For the full flag list, types and defaults, see Manager Flags; this page does not repeat that table.
Before you begin
- Cluster-admin access to the management cluster, and
kubectl. - The shipped Deployment is named
captf-controller-managerin the provider’s namespace (captf-systemafterclusterctl init; see Installation).
Changing a flag
The manager’s arguments live on the manager container of the
captf-controller-manager Deployment. Because args is a plain list with no
merge key, a strategic-merge patch that only adds one flag replaces the
whole list and silently drops the others (including --leader-elect and
the diagnostics flags the shipped manifest sets). Add a flag with a JSON
patch instead, which appends to the existing list:
kubectl patch deployment captf-controller-manager -n captf-system --type=json \
-p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--<flag>=<value>"}]'
To manage the arguments as a whole (for example with a kustomize overlay
over the released manifest), patch the full args list so no existing flag
is lost:
apiVersion: apps/v1
kind: Deployment
metadata:
name: captf-controller-manager
namespace: captf-system
spec:
template:
spec:
containers:
- name: manager
args:
- --leader-elect
- --diagnostics-address=:8443
- --insecure-diagnostics=false
- --webhook-port=9443
- --<flag>=<value>
Confirm the change with:
kubectl rollout status deployment/captf-controller-manager -n captf-system
kubectl get deployment captf-controller-manager -n captf-system \
-o jsonpath='{.spec.template.spec.containers[0].args}'
The manager also logs its parsed flags once at start (FLAG: --name="value"
lines), and refuses to start on an invalid combination, so a typo or an
out-of-range value shows up in the Pod’s logs and status rather than
running with a wrong value. clusterctl upgrade apply reinstalls the
provider’s manifests, so a hand-patched flag does not survive an upgrade
unless the patch is reapplied; see Upgrades.
Namespace scoping and --watch-filter
--namespace restricts the manager to one namespace: it only watches and
reconciles namespaced objects (TerraformCluster, TerraformMachine,
TerraformMachinePool, their *Template kinds, and the Secrets and
ConfigMaps CAPTF reads) in that namespace. TerraformClusterIdentity is
cluster-scoped and is always watched everywhere, regardless of
--namespace. Unset (the default, and what clusterctl init installs),
the manager watches every namespace. A single manager instance therefore
watches one namespace or all of them, never a chosen set: running a
second, differently-namespaced manager instance does not split load or
ownership the way it might look like it should. The leader-election Lease
name (controller-leader-election-captf) is a fixed constant, not
namespaced or parameterized by --namespace, and the shipped install is a
singleton — one ValidatingWebhookConfiguration, one webhook Service,
one Certificate. Two manager Deployments with --leader-elect and
different --namespace values both contend for the same Lease in the same
manager namespace, so only one of them ever holds leadership and
reconciles anything; the other sits idle regardless of which namespace it
was pointed at.
To share one management cluster across several provider instances or
tenants, use --watch-filter instead: it limits reconciling to objects
labeled cluster.x-k8s.io/watch-filter: <value>; unset, the manager
reconciles every object it can see. It is the convention CAPI and its
other providers use for exactly this, and each instance still watches (and
its webhook still admits) every namespace. The filter applies to the
Terraform* objects themselves; a Job or Secret one of them owns is still
acted on regardless of its own labels.
The ValidatingWebhookConfiguration has no namespaceSelector: it admits
a Terraform* create or update in any namespace, whether or not any
manager instance watches that namespace or the object’s
watch-filter label matches one. A Terraform* object created in a
namespace no running manager watches (outside --namespace’s one
namespace, or not labeled for any instance’s --watch-filter) is admitted
normally and then never reconciled: no Job ever starts for it, and its
conditions never move past whatever they were set to (or left unset) at
admission.
Both flags apply to the orphan sweep too: --namespace limits
which namespaces it sweeps, and it always ignores --watch-filter when it
decides whether a namespace still holds a Terraform* object, so another
instance’s objects (including one this manager does not watch) keep a
namespace from being swept.
Per-kind concurrency
--terraformcluster-concurrency, --terraformmachine-concurrency,
--terraformmachinepool-concurrency and
--terraformmachinetemplate-concurrency each cap how many objects of that
kind the manager reconciles at once (default 10). Raise one when that
kind’s objects queue behind each other under load — reconciles are
lightweight (a handful of API reads and a status patch, unless a Job needs
starting) and mostly wait on Job completion, so a higher number rarely
costs much CPU. Lower one to reduce the manager’s burst of API calls
against a small or rate-limited API server.
Leader election
--leader-elect (default false) turns on leader election so that only
one of several manager replicas reconciles at a time; enable it whenever
you run more than one replica, so a rolling update or a crash never leaves
two managers reconciling the same objects together. The shipped Deployment
runs one replica but sets --leader-elect anyway, ready for a scale-up.
--leader-elect-lease-duration, --leader-elect-renew-deadline and
--leader-elect-retry-period tune how fast a crashed leader is detected
and replaced; the defaults (15s/10s/2s) match kube-controller-manager and
rarely need changing.
Sync period and the orphan sweep
--sync-period (default 10 minutes) is the minimum interval at which the
manager’s informers re-enqueue every cached object for reconciliation, on
top of the normal event-driven reconciles; it reads only the local cache,
never the API server, so it does not set the cadence of drift or health
checks, which run on their own schedule
(see Drift and Health), and it does not
recover a watch event that never reached the cache. Lowering it corrects a
missed requeue sooner at the cost of more reconciles; raising it does the
opposite.
The same interval drives the orphan sweep, so lowering --sync-period also
cleans up an orphaned namespace’s runner objects sooner. See
The orphan sweep for what it removes and how it
decides.
--cluster-operation-gate
Keeps a TerraformCluster‘s apply, destroy or restore from running at the
same time as its machines’ and machine pools’ applies, destroys or
restores, through a per-Cluster write Lease; the per-object run Lease that
keeps two Jobs from starting for the same object is always on and cannot be
turned off. Default true. Turn it off only if you accept a cluster’s own
operation and its machines’ operations running concurrently — see
Run leases and the cluster operation gate
for what the gate does and how the two sides wait for each other.
--runner-image
The image of the init container that copies the runner binary into every
Job; it must be an image that contains a /runner binary, which in practice
means a CAPTF manager or runner image, not a module image. Unset, it
defaults to $CAPTF_MANAGER_IMAGE, the manager’s own image, which the
shipped Deployment sets to whatever image it runs — so most installs never
need to set the flag or the environment variable by hand. Set it to pin
the init container to a specific published image independently of the
manager’s own, for example while testing a new manager build against the
current runner. The manager refuses to start when neither the flag nor the
environment variable resolves to a valid image reference. See
Job Environment for where the copied binary
ends up and Security Model for why images
are pinned by digest once a Job has run.
--runner-events
Default true. Have each Job’s runner post its own progress (RunStarted,
Step*, PlanSummary, ResourcesChanged, RunFinished) as Events on the
Terraform* object the Job is for, related to the Job itself; emission is
best effort and never fails or slows the run. Turn it off to reduce Event
volume on a cluster with many objects, or if the runner ClusterRole in your
installation does not grant events create (see RBAC). See
Events for the full list.
--state-backups
How many state backups to keep per object (default 5); every new state
serial the manager observes is copied into a captf-state-backup-* Secret,
and older copies beyond the count are pruned in the same pass.
--state-backups=0 takes no new backups but leaves existing ones in place
and restorable. Lower it to reduce the Secret count and storage in a large
installation; raise it for a longer recovery window. See
Terraform State for the backup
naming and what is skipped, and
State Restore for the restore procedure.
--drift-default-interval
The drift check interval (default 30 minutes) an object falls back to when
neither it nor, for a machine or pool, its TerraformCluster’s
spec.defaults.drift sets one. For a TerraformCluster or
TerraformMachine, it has no effect once spec.drift.intervalSeconds (own
or inherited) is set, including to 0, which disables drift. A
TerraformMachinePool’s drift can never be disabled: a 0, its own or
inherited, falls back to this default instead. Change the default to shift
the fleet-wide drift cadence without touching every object; see
Drift for setting an interval per object and
Drift and Health for how the schedule
and its jitter work.
Diagnostics address, authentication and TLS
--diagnostics-address (default :8443) is where the manager serves
Prometheus metrics, authenticated and authorized against the API server by
default. --insecure-diagnostics (default false) turns that off and
serves plain HTTP with no authentication instead; use it for local
development only, never for a manager reachable from anything but your own
workstation. --tls-min-version, --tls-cipher-suites and
--tls-curve-preferences constrain the TLS the metrics and webhook servers
negotiate, the same flags and defaults as the CAPI core providers. See
Observability for what the endpoint serves and how a
scraper authenticates to it.
--webhook-port (default 9443) is where the manager serves admission
webhooks; --webhook-cert-dir, --webhook-cert-name and
--webhook-key-name say where it finds the serving certificate that
cert-manager issues. The shipped manifests wire all of this together; see
Installation.
Logging
--logging-format (default text) also accepts json. -v sets the log
verbosity (default 2); raise it while diagnosing a problem and lower it
back afterward, since higher verbosities log more of each reconcile.
--vmodule overrides the verbosity for individual source files, and only
works with the text format. --feature-gates takes a comma-separated
key=value list of the logging feature gates (ContextualLogging,
LoggingBetaOptions, LoggingAlphaOptions); CAPTF itself registers no
feature of its own. See Manager Flags
for the full flag and feature gate list.
See also
- Manager Flags — every flag, its type and its default.
- Installation — installing the provider and its webhook certificate.
- Upgrades — what a provider upgrade changes.
- RBAC — the orphan sweep and the runner’s permissions.
- Observability — metrics, alerts and events.
RBAC
CAPTF ships a fixed set of ClusterRoles, one Role and the objects that bind them, and creates one more RoleBinding per tenant namespace as it goes. This page covers every one of them: what the manager itself can do, what the runner can do inside a tenant namespace, how a namespace gets its runner identity, and how CAPTF cleans that identity up again. For what a Terraform or OpenTofu module can do with the runner’s access once it has it, see Security model.
Before you begin
- Cluster-admin access to the management cluster, and
kubectl. - The manager’s own ClusterRole and ClusterRoleBinding are named
captf-manager-roleandcaptf-manager-rolebindingafterclusterctl init(see Installation); the runner ClusterRole iscaptf-runnerregardless of which cluster or namespace it runs in.
The manager’s ClusterRole
The manager’s ServiceAccount, captf-controller-manager in the provider’s
namespace, is bound to one ClusterRole scoped to exactly what reconciling
Terraform* objects needs:
| Resource | Verbs | Why |
|---|---|---|
terraformclusters, terraformmachines, terraformmachinepools, terraformclusteridentities, their *Template kinds | get, list, watch, create, update, patch, delete | Reconciles and owns every kind it serves |
The /status subresources of every kind above except terraformmachinepooltemplates (which has none) | get, update, patch | Writes status separately from the spec |
The /finalizers subresources of terraformclusters, terraformmachines, terraformmachinepools and terraformclusteridentities (never a *Template kind) | update | Adds and removes its own finalizers |
clusters, clusters/status, machines/status, machinepools/status | get, list, watch | Reads the Cluster API objects a Terraform* object belongs to |
machines | get, list, watch, patch | Reads Machines and patches the cluster.x-k8s.io/remediate-machine annotation it sets to request remediation |
machinepools | get, list, watch, patch | Reads MachinePools and writes back an autoscaled pool’s observed replicas |
secrets | get, list, watch, create, update, patch, delete | Reads and writes every Secret in Secrets |
namespaces | get, list, watch | Evaluates a TerraformClusterIdentity’s allowedNamespaces selector |
configmaps | get, list, watch | Reads spec.variablesFrom ConfigMap sources labeled captf.io/variables=true |
serviceaccounts | get, list, watch, create, delete | Creates the default runner ServiceAccount, reads any override, and the orphan sweep deletes unused ones |
rolebindings | get, list, watch, create, update, delete | Manages the per-namespace runner RoleBinding below |
clusterroles, resource name captf-runner | bind | Lets the manager bind that ClusterRole without holding its permissions itself |
leases | get, list, watch, create, update, delete | Its own run and cluster-operation-gate leases, and cleaning up the backend’s state lock lease on delete |
jobs | get, list, watch, create, patch, delete | Creates, watches and prunes the runner Jobs |
jobs/finalizers | update | Granted alongside jobs; the controller itself never sets a finalizer on a Job |
pods | get, list, watch | Reads a Job’s pod status for its outcome and the image digest it ran |
pods/log | get, list, watch | Granted alongside pods; the controller itself never reads a pod’s logs |
events, events.k8s.io/events | create, patch | Records reconcile events on Terraform* objects |
authentication.k8s.io/tokenreviews | create | Backs the authenticated diagnostics endpoint |
authorization.k8s.io/subjectaccessreviews | create | Backs the authenticated diagnostics endpoint and the identity webhook’s Secret-read check |
The manager is never labeled to aggregate into a broader ClusterRole, and it holds no permission on any resource outside this table.
The leader-election Role
captf-leader-election-role is a namespaced Role, scoped to the provider’s
namespace, bound to the manager’s ServiceAccount by
captf-leader-election-rolebinding. It grants leases (get, list, watch,
create, update, patch, delete) and events (create, patch):
controller-runtime’s leader election uses a Lease, one per manager, and only
matters when --leader-elect is on (see
Leader election).
The runner ClusterRole
captf-runner is the one ClusterRole every tenant namespace’s runner
ServiceAccount is bound to. It is cluster-scoped and static: the manager
never edits its rules, only binds it per namespace
(below).
| Resource | Verbs | Why |
|---|---|---|
secrets | get, list, create, update, delete | The Terraform/OpenTofu Kubernetes state backend’s own requirements |
leases | get, create, update | Takes, releases and force-unlocks the state lock |
events.k8s.io/events | create | Reports run progress on the object that owns the Job |
The Secret verbs are exactly what the state backend uses, no more: get,
create and update read and write the state; list is needed because the
backend lists its state chunks by label on every read and write, and lists
workspaces during init; delete trims chunks when the state shrinks. None
of those verbs can be scoped by name or label, so they apply to every Secret
in the namespace — this is the basis of the trust boundary described in
Security model. leases omits delete:
the backend deletes its lock Lease only when a workspace is deleted, which
the runner never does, and the manager deletes it itself once an object’s
state is gone. events.k8s.io/events grants create only, never patch or
get: the runner emits a series of one-shot events
(Reading events) and never reads or
updates one; with --runner-events=false the Jobs pass no event target and
nothing calls this permission.
No Role is ever created for the runner. The manager binds this ClusterRole
through the bind verb, restricted to its name, so the manager does not
need to hold the ClusterRole’s own permissions itself.
The per-namespace runner ServiceAccount and RoleBinding
The first time a namespace needs a Job, the manager makes sure it has both:
- a ServiceAccount,
captf-runnerby default, labeledcaptf.io/managed=true. An operator who pre-creates it themselves keeps whatever else they put on it. - a RoleBinding, always named
captf-runner, whoseroleRefis fixed to thecaptf-runnerClusterRole and whose subjects are recomputed on every reconcile: the union of every ServiceAccount someTerraform*object in the namespace runs its Job as.
A RoleBinding named captf-runner that already exists without
captf.io/managed=true is left alone and never modified: the manager
reports an error rather than take over a hand-written binding of the same
name. One it does manage is re-created, not patched, if something changes
its roleRef — roleRef is immutable in Kubernetes — and otherwise only
has its subject list updated.
Custom ServiceAccounts and the runner opt-in
spec.jobs.serviceAccountName (or the same field under a
TerraformCluster’s spec.defaults.jobs) can name a ServiceAccount other
than captf-runner for a Job to run as. Naming captf-runner itself is not
an override and needs nothing further. Naming anything else needs that
ServiceAccount to already exist and carry the label
captf.io/runner=true; without both, the condition
RunnerRBACReady=False/ServiceAccountNotOptedIn is set and no Job runs.
Once it opts in, the manager adds it to the captf-runner RoleBinding’s
subjects alongside every other ServiceAccount the namespace’s objects use;
removing the label removes it from the binding again on the next reconcile.
The label is consent, not authorization: whoever can label a ServiceAccount
opts it in, whether or not they administer the namespace. It grants little
by itself, because anyone who can already create Pods running as that
ServiceAccount can read the namespace’s Secrets through a volume mount
anyway; what it prevents is a ServiceAccount gaining the runner’s Secret
write access merely by being named in spec.jobs.serviceAccountName.
The orphan sweep
At manager start, and then every --sync-period while this manager leads
(see Sync period and the orphan
sweep), the manager
deletes the captf.io/managed=true ServiceAccounts, RoleBindings and Leases
of every namespace holding no TerraformCluster, TerraformMachine and
TerraformMachinePool — what a clusterctl move leaves
behind in the source namespace, since clusterctl strips finalizers before
deleting the source objects and the controller’s own delete-time cleanup
never runs there. The decision is made through the manager’s uncached reader
with no --watch-filter applied, so another manager instance’s objects,
even ones this manager does not watch, still keep a namespace from being
swept. Nothing is deleted for a single object here, so the sweep logs what
it removed rather than recording an event.
A namespace that still holds objects keeps its runner ServiceAccount and
RoleBinding, but the RoleBinding’s subjects are pruned to the
ServiceAccounts still in use: a Terraform* object’s own
spec.jobs.serviceAccountName, else, for a machine or pool, its cluster’s
spec.defaults.jobs.serviceAccountName, else captf-runner. This is how
switching spec.jobs.serviceAccountName from one opted-in ServiceAccount to
another eventually drops the old one’s Secret access, rather than leaving it
bound until the whole namespace empties out. When a machine or pool’s
cluster cannot be resolved yet, every cluster default in the namespace is
kept rather than pruned, so a resolution delay never flaps the binding.
The metrics-reader ClusterRole
captf-metrics-reader lets a Prometheus instance read the manager’s
authenticated /metrics endpoint; it ships only with the opt-in Prometheus
component, not with the base install. See Enabling the Prometheus
component for what it
binds to and how to point it at your own Prometheus.
See also
Secrets
This page inventories every Secret CAPTF reads or writes: what it is named, which namespace it lives in, who creates and deletes it, what it contains, and whether it carries sensitive data. Read it before deciding who may read Secrets in a tenant namespace, or before writing a NetworkPolicy or admission policy that assumes only some Secrets matter. For what the runner’s own access to these Secrets means once a Job is running, see Security model.
Before you begin
getaccess to Secrets in the namespaces you administer, to inspect any of these.- Knowing which kind (
TerraformCluster,TerraformMachineorTerraformMachinePool), object name or identity a Secret belongs to; several of the name patterns below embed one of these.
Every Secret CAPTF reads or writes
| Name pattern | Namespace | Owner | Created / deleted | Contents | Sensitive |
|---|---|---|---|---|---|
Identity’s source Secret (any name, TerraformClusterIdentity.spec.secretRef) | secretRef.namespace, any namespace | The operator | By the operator; CAPTF never creates, updates or deletes it, and removes any owner reference an older CAPTF version left on it | Cloud credentials (arbitrary keys and values, provider SDK conventions or file contents) | Yes |
captf-creds-<identity> (mirror) | Each namespace allowedNamespaces permits and that has resolved the identity | The Terraform* object(s) using the identity there, as non-blocking owner references | Created the first time an object in the namespace resolves the identity; rewritten whenever the source Secret’s data changes; deleted when the namespace stops being allowed, or when its last owning object is removed | A byte-for-byte copy of the source Secret’s data | Yes |
captf-inputs-c-<name>, captf-inputs-m-<name>, captf-inputs-mp-<name> (durable inputs) | The object’s namespace | The TerraformCluster, TerraformMachine or TerraformMachinePool itself | Created before the object’s first Job; rewritten on every reconcile that re-renders inputs; deleted once the object’s state and infrastructure are gone | The rendered root module and tfvars, plus the pinned image, image digest and identity | Yes |
captf-run-<job> (per-run) | The object’s namespace | The Job | Created just before the Job’s pod starts, from that reconcile’s rendered inputs; deleted once the controller has read the finished Job’s result | The same rendered root module and tfvars the Job runs with | Yes |
tfstate-default-<suffix> and its -part-N chunks (state) | The object’s namespace | Unowned until the first successful apply, then the object, as a non-blocking owner reference | Created by the Terraform/OpenTofu Kubernetes state backend itself, not by CAPTF; deleted by CAPTF after a successful destroy, or immediately on deleting an object that never applied | The compressed Terraform state, plus the backend’s own workspace labels | Yes |
captf-state-backup-<suffix>-<serial> and its -part-N chunks (state backups) | The object’s namespace | The Terraform* object, as a non-blocking owner reference (never the state) | Created after a reconcile observes a state serial not backed up yet; pruned to a configured retention on the same pass; never deleted by the destroy cleanup above | A verbatim copy of the state Secrets they were taken from, at that serial | Yes |
Bootstrap data Secret (any name, Machine.spec.bootstrap.dataSecretName or MachinePool.spec.template.spec.bootstrap.dataSecretName) | The Machine’s or MachinePool’s namespace | The bootstrap provider (for example a KubeadmConfig), not CAPTF | By the bootstrap provider; CAPTF only reads it, uncached, and never labels, updates or deletes it | The bootstrap data (value) and its format (format) | Yes |
spec.variablesFrom source (any name, labeled captf.io/variables=true) | The object’s namespace | The operator | By the operator; CAPTF only reads it, and only while the label is present — an unlabeled Secret counts as missing | Arbitrary keys treated as module variable values | Yes |
Image pull secret (any name, named in spec.jobs.imagePullSecrets) | The object’s namespace | The operator | By the operator; for a TerraformCluster, TerraformMachine or TerraformMachinePool, CAPTF only references its name on the Job’s pod, never reading or writing its contents. For a TerraformMachineTemplate, CAPTF also reads its own pull Secrets’ contents to authenticate the image inspection that resolves capacity, but never writes them | A kubernetes.io/dockerconfigjson registry credential | Yes |
captf-webhook-service-cert (webhook serving certificate) | The provider’s namespace | cert-manager, through its Certificate object | By cert-manager, on issuance and renewal; CAPTF never creates, reads or deletes it | A TLS key pair and CA bundle for the admission webhook | Yes |
A name pattern that would exceed the 253-character Secret name limit — the mirror and the durable inputs Secret, both built from a user-chosen name — is shortened to its prefix plus a hash, still deterministic and unique.
What the manager caches
The manager’s main cache holds only Secrets labeled captf.io/managed=true:
credential mirrors, durable and per-run inputs, state and state backups —
the five Secrets above that carry that label, whether CAPTF or the state
backend created them. A
second, separate cache backs spec.variablesFrom watches; it holds Secrets
(and ConfigMaps) labeled captf.io/variables=true, with their data stripped
out before they are stored, so no variable value ever sits in memory there.
Everything else — the identity’s source Secret, bootstrap data, image pull
secrets, and a spec.variablesFrom Secret’s actual content when it is
resolved — is read directly from the API server on demand and never
watched.
What never enters logs or status
The durable and per-run inputs Secrets carry bootstrap data and module
variable values in clear by design, and CAPTF never logs their data. State
and its backups are read the same way: never logged. A spec.variablesFrom
value marked sensitive is redacted from the Job’s own plan and apply output
and from the controller’s trace-level logs, but that redaction does not
reach the Secrets in this table: like every other input, the value is still
written in clear into the durable and per-run inputs Secrets and into the
state.
This is a separate guarantee from what CAPTF keeps out of status and
events, which never carry Secret contents at all; see what CAPTF keeps out
of status, events and
logs.
See also
Observability
This page is for operators running the manager: what its metrics endpoint serves and who may read it, how to wire up Prometheus, what each alert means and where to look first, and how to read conditions, events and logs once you are looking at a specific object or Job.
The diagnostics endpoint
The manager serves Prometheus metrics over HTTPS, authenticated and authorized against the API server by default:
- The bearer token in the request is checked with a
TokenReview. - The token’s access is checked with a
SubjectAccessReviewforgeton the non-resource URL/metrics.
--insecure-diagnostics turns both checks off and serves plain HTTP
instead; use it only for local development.
Configuration
covers that flag and the address, TLS version and cipher flags, and
Manager Flags
lists them with their defaults.
| Request | Result |
|---|---|
GET /metrics without a token | 401 Unauthorized |
GET /metrics with a token authorized for get on /metrics | 200, serving captf_build_info and the rest of the series |
There is no flag for the metrics server’s own certificate: mount a Secret
with tls.crt and tls.key at /tmp/k8s-metrics-server/serving-certs/ to
serve your own, or leave it unmounted and the manager generates a
self-signed certificate at startup.
The same authenticated endpoint also serves a profiler
(GET /debug/pprof/*) and an endpoint to change the log level at runtime
(PUT /debug/flags/v; see Logs and verbosity) once
--insecure-diagnostics is off. Each request is authorized the same way,
against its own non-resource URL and the HTTP method lowercased as the
verb (get for the profiler, put for the log level). Nothing in the
shipped manifests grants either, only /metrics, so bind your own
ClusterRole to reach them.
The series themselves are listed in Metrics.
Enabling the Prometheus component
config/prometheus is an opt-in kustomize component: a metrics Service,
a ServiceMonitor and a PrometheusRule with the eleven alerts below,
plus a ClusterRole that lets a Prometheus instance read the diagnostics
endpoint. It is not part of infrastructure-components.yaml and is not a
release asset, so clusterctl init never installs it and clusterctl upgrade never touches it: get config/prometheus from a checkout of the
tag you installed, build it, and apply it yourself, and repeat that after
every upgrade to pick up any change to the alert rules. See
RBAC for what the ClusterRole
grants.
Before you begin:
- the Prometheus Operator CRDs (
ServiceMonitor,PrometheusRule) are installed in the cluster; - you know the namespace and name of the ServiceAccount your Prometheus
scrapes with, if it is not kube-prometheus’s default
(
monitoring/prometheus-k8s).
-
Build the component on its own — it needs no other resources, and its objects already carry the
captf-names and thecaptf-systemnamespace thatconfig/defaultproduces:apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization components: - <path-to-checkout>/config/prometheuskustomize build <path-to-this-kustomization> >captf-prometheus.yaml -
If your Prometheus does not run as
monitoring/prometheus-k8s, patch thecaptf-metrics-readerClusterRoleBinding’s subject to your ServiceAccount instead, before applying. -
Apply the built manifest with
kubectl apply -f captf-prometheus.yaml. If your provider install itself uses renamed objects (a non-defaultnamePrefixor namespace), patch theServiceMonitor’s andPrometheusRule’s selectors and theServicereference to match, since a component does not inherit an overlay’s name or namespace transformers.
The ServiceMonitor scrapes over HTTPS with insecureSkipVerify set,
since the metrics server’s certificate is self-signed by default; once you
mount a CA-signed certificate, set tlsConfig.ca instead and drop
insecureSkipVerify.
To confirm it works without waiting on Prometheus, read the endpoint yourself with the same kind of token Prometheus uses:
kubectl port-forward -n captf-system svc/captf-controller-manager-metrics 8443:8443 &
curl -sk -H "Authorization: Bearer $(kubectl create token <serviceaccount> -n <namespace> --duration=5m)" \
https://localhost:8443/metrics | grep captf_build_info
<serviceaccount> and <namespace> name a ServiceAccount already bound
to captf-metrics-reader (or your own equivalent grant). A captf_build_info
line back confirms both the endpoint and the token’s authorization; 401
or 403 means the token or the binding, not the manager.
make promtool-check and make promtool-test validate the alert rules
before you ship a change to them; see
Make Targets.
Useful queries
# p95 Job duration by op, over the last day, for Jobs that ran their course
histogram_quantile(0.95, sum by (le, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1d])))
# Slowest runner steps (p90) by kind, op and step
histogram_quantile(0.9, sum by (le, kind, op, step) (rate(captf_job_step_duration_seconds_bucket[6h])))
# Failures by error kind and step
sum by (op, error_kind, step) (increase(captf_job_errors_total[1d]))
# p90 queue time (creation to the source container's start) by op
histogram_quantile(0.9, sum by (le, op) (rate(captf_job_queue_seconds_bucket[1h])))
# The ten largest states, and how close each is to one Secret's 1 MiB
topk(10, captf_state_bytes)
captf_state_bytes / (1024 * 1024)
# Objects whose drift check or health refresh is overdue
time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 3600
No Grafana dashboard ships with the component; the queries above are the panels one would build. The full series list, with every label, is in Metrics.
Alerts
config/prometheus ships eleven alerts on the series above. Two things
help in reading them:
- The failure counters count transitions, not reconciles: they go up once
when an object or Job newly enters the bad state, so
increase(...) > 0means something newly broke, not that it is still broken. CAPTFClusterDrift,CAPTFStateNearSecretLimit,CAPTFInputsNearLimitandCAPTFNoRecentSuccessname the object directly, withnamespaceandnamelabels. The rest aggregate bykind(withoporreason), exceptCAPTFReconcileErrors, which aggregates bycontrolleralone; find the affected object through its conditions, as each section below says.
Every rule’s expression, for and severity are in
Alerts; each heading below links there.
CAPTFJobFailing
Jobs of a kind and op failed or hit their deadline more than twice in 30
minutes. Find the objects with ApplyJobSucceeded=False or
DriftJobSucceeded=False and read status.lastRun and the Job’s logs:
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | .type == ("ApplyJobSucceeded", "DriftJobSucceeded") and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"'
See Failing Jobs and the rule.
CAPTFDestroyStuck
An object’s destroy keeps failing while it deletes; its finalizer and state stay in place, so nothing is orphaned. See Stuck Destroy and the rule.
CAPTFClusterDrift
A TerraformCluster’s last drift check found changes and it has stayed
that way for an hour. With drift.action: Report an operator decides
next; with Remediate the remediation apply is failing, check
ApplyJobSucceeded. See Drift and
the rule.
CAPTFStateUnreadable
A state Secret turned unreadable. Find the object with
StateReadable=False and read that condition’s reason and message:
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | .type == "StateReadable" and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"'
See Unreadable State and the rule.
CAPTFForceUnlocks
A Job force-unlocked a state lock whose holder pod no longer existed. Find out why the previous Job died before it happens again. See Stale State Lock and the rule.
CAPTFReconcileErrors
A controller keeps returning errors from its reconcile loop. Read the manager’s logs for the failing kind. See Reconcile Errors and the rule.
CAPTFJobSlow
Jobs of a kind and op are taking longer than expected at the 90th
percentile. Compare it with the Jobs’ activeDeadlineSeconds and find the
slow step with captf_job_step_duration_seconds. See
Slow Jobs and
the rule.
CAPTFJobQueueSlow
Jobs are waiting too long between creation and the source container starting: scheduling, image pulls or the runner’s own init copy. Look for Pending runner pods and their events. See Slow Jobs and the rule.
CAPTFStateNearSecretLimit
An object’s compressed state is approaching the 1 MiB a Secret can hold.
captf_state_resources shows how many resources it manages. See
Size Limits and
the rule.
CAPTFInputsNearLimit
An object’s rendered inputs are approaching the size no Job will start past. See Size Limits and the rule.
CAPTFNoRecentSuccess
An object’s scheduled drift check or health refresh has not succeeded in
six hours. Read DriftJobSucceeded and status.lastRun, the same as for
a failing Job. See Failing Jobs and
the rule.
Reading conditions
kubectl describe on any TerraformCluster, TerraformMachine,
TerraformMachinePool or TerraformClusterIdentity shows its conditions:
a type, a status of True, False or Unknown, a reason and a message.
Most conditions are normal polarity (True is healthy); a few are
inverted, such as Deleting. Unknown most often means the object is
waiting on something else to finish, not that anything failed: it does
not fail Ready. See retry backoff
for how a wait like this is treated.
Every condition type CAPTF sets, its polarity, and every reason and message it can carry are in Conditions.
Reading events
Every stage of an object’s life emits a Kubernetes Event on it: the manager once per transition, and a Job’s runner in real time while it runs. Read them in order with:
kubectl events --for terraformcluster/<name> -n <namespace>
kubectl events --for terraformmachine/<name> -n <namespace>
kubectl events --for terraformmachinepool/<name> -n <namespace>
kubectl events --for terraformclusteridentity/<name> -n default
TerraformClusterIdentity is cluster-scoped, but its events still land
in the default namespace.
Add -o wide for a SOURCE column that distinguishes the manager’s
events from a runner’s. Notes never carry credentials, tfvars, output,
plan values or raw stderr: a step failure’s note is the runner’s curated
summary, not its log output. Every reason, its type and what it means are
in Events.
Runner events
A Job’s runner posts its own progress (RunStarted, StepStarted,
StepSucceeded, StepFailed, PlanSummary, ResourcesChanged,
RunFinished) as Events on the object the Job is for, related to the Job
itself. Emission is best effort: each request has its own short timeout,
and after a few consecutive failures the runner stops emitting for the
rest of that run without failing or slowing it. Turning --runner-events
off (see Configuration) skips them
entirely, which is worth doing on a large fleet since a single scheduled
drift check alone produces several of them. The runner’s own events
create grant is in RBAC.
Logs and verbosity
The manager logs at a default verbosity where the usual reconcile flow is visible; raising it shows more detail down to per-request tracing, and lowering it keeps only errors and irreversible actions such as force unlocks. Credentials, bootstrap data, tfvars content and output values are never logged, whatever the level.
Change the level without restarting the manager through the same authenticated diagnostics endpoint:
curl -sk -X PUT -H "Authorization: Bearer <token>" --data '<level>' \
https://<address>/debug/flags/v
The token needs its own authorization for put on /debug/flags/v; the
shipped manifests do not grant it. --v, --vmodule and
--logging-format set the level, per-file overrides and the output
format at startup instead; see
Configuration and
Manager Flags.
Read the manager’s own logs, and a Job’s, with:
kubectl logs -n captf-system deploy/captf-controller-manager -c manager
kubectl logs -n <namespace> job/<name> -c source
See also
- Metrics — every series, its type, labels and meaning.
- Alerts — every rule’s expression,
forand severity. - Conditions — every condition, reason and message.
- Events — every event reason, type and meaning.
- Configuration — the flags behind the diagnostics endpoint, runner events and logging.
- RBAC — the manager’s and Prometheus’s RBAC.
- Runbooks — the recovery procedures the alerts above link to.
Upgrades
This page covers upgrading an installed CAPTF provider with clusterctl,
how CAPTF versions its contract, what an upgrade can trigger on existing
objects, where to read what changed, and how to roll back. See
Installation for installing CAPTF the first time.
Before you begin
- CAPTF already installed with
clusterctl init(see Installation), and aclusterctl.yamlnaming the provider (name: terraform). clusterctl, at a version that supports upgrading the Cluster API release you run.
Check what changed first
Before upgrading, read what changed between your installed version and the target one:
- The provider’s own release notes on its GitHub release, generated from the commits since the previous tag.
- The module contract changelog, when the entries mention a contract or inputs-hash change (below); it records every contract-visible change, newest first, whether or not the contract’s own version number moved.
Upgrade with clusterctl
clusterctl upgrade plan
lists the provider versions clusterctl can upgrade each installed component to, grouped by the Cluster API contract they implement. Apply a plan, or name a version for CAPTF directly:
clusterctl upgrade apply --contract v1beta2
clusterctl upgrade apply --infrastructure terraform:vX.Y.Z
CAPTF has not published a release yet, so there is nothing for
clusterctl upgrade to find until one exists; until then, changing
versions means reinstalling from a local repository, the same way as a
first install (see Installing from a local
repository).
Contract versioning
CAPTF’s metadata.yaml lists its release series for clusterctl: each
entry is a (major, minor) pair and the Cluster API contract it
implements. The list is append-only — an entry, once published, is never
removed or changed — so clusterctl upgrade plan can always resolve an
older installed version’s contract. CAPTF has one series so far, 0.1,
implementing contract v1beta2; a later series only appears once a
release under it exists, and only ever adds to the list.
The contract version in metadata.yaml is separate from the module
contract version (v1alpha1, Module Contract):
the first is what Cluster API’s own core and other providers require of
CAPTF, the second is what CAPTF requires of a module image. A provider
upgrade can change either, both, or neither.
What an upgrade can trigger
An upgrade replaces the installed manifests — CRDs, RBAC, the manager and
its webhooks — but does not touch a Terraform* object’s spec. Two things
it can still trigger on existing objects, both driven by the manager’s own
reconcile after it restarts on the new version:
- A one-time re-apply of every provisioned mutable object. The inputs
hash (
captf.io/inputs-hash) covers a hash scheme identifier along with the contract version, role, image reference and rendered inputs (see What is (and isn’t) in the inputs hash). A release that changes what the hash covers bumps the scheme, so every existingTerraformCluster’s andTerraformMachinePool’s inputs hash changes even though nothing in its spec did; the object’s next reconcile sees a hash that no longer matches its state and re-applies once (aTerraformClusterwithspec.applyPolicy: Manualplans once instead and waits for approval; see Plan Approval). ATerraformMachineis unaffected: its inputs hash is never recomputed once it is provisioned. - A one-time cleanup of a
TerraformClusterIdentity’s credentials Secret. Earlier versions put an ownerRef on the Secret named inspec.secretRef, so deleting the identity garbage-collected it; current versions do not own that Secret. The identity’s reconcile removes a leftover ownerRef the first time it runs after the upgrade, independent of whether anything uses the identity. Let the upgraded manager reconcile every identity once before deleting one, so this cleanup runs first; see Identities and Credentials.
Neither is specific to the two examples above: any change the contract changelog records as covering the inputs hash or an identity’s Secret ownership behaves the same way on the next release that ships it. Read the changelog entries for the version you are moving to, since only they say whether either applies.
Rolling back
clusterctl has no dedicated rollback command: rolling back means
installing the older version’s manifests the same way you install any
version, clusterctl upgrade apply --infrastructure terraform:vX.Y.Z
naming the earlier tag, or reinstalling from that tag’s local repository.
Rolling back reverses an inputs-hash scheme change the same way upgrading
applies one: the older manager computes the older scheme, so every
provisioned mutable object re-applies once again on its first reconcile
after the rollback. v1alpha1 (the module contract, distinct from the
metadata.yaml contract above) carries no compatibility guarantee between
its own changes (Versioning),
so a rollback that crosses a contract-changing release can also mean the
older manager and the module images or generated roots from the newer one
disagree about a field name or an input; check the contract changelog for
the versions in between before rolling back across one.
Runbooks
Each runbook below covers one symptom: what it looks like, why it happens, how to check, how to fix it, and how to confirm the fix worked. Alerts and condition reasons are cross-references, not a substitute for reading the page: start from whichever alert fired or condition you see, but read the whole runbook before you act.
- Failing Jobs — a
TerraformCluster,TerraformMachineorTerraformMachinePoolwhose Jobs keep failing.CAPTFJobFailing,CAPTFNoRecentSuccess;ApplyJobSucceeded=FalseandDriftJobSucceeded=False. - Stuck Destroy — a
destroyJob that cannot succeed, and how to remove the object’s finalizer safely.CAPTFDestroyStuck;ApplyJobSucceeded=False/DestroyFailed. - Unreadable State — a state Secret CAPTF cannot
parse.
CAPTFStateUnreadable;StateReadable=False/StateCorrupt,StateInconsistentorStateEncrypted. - State Restore — restoring a Terraform or OpenTofu
state from a CAPTF-managed backup.
StateReadable=False/StateLost;RestoreJobSucceeded. - Stale State Lock — clearing a state lock left behind by
a killed or evicted runner.
CAPTFForceUnlocks;StateReadable=False/StateLocked. - Size Limits — a state or rendered inputs approaching
the Secret size limit.
CAPTFStateNearSecretLimit,CAPTFInputsNearLimit;ApplyJobSucceeded=False/InputsTooLarge. - Slow Jobs — Jobs that take a long time to run, or a long
time to start.
CAPTFJobSlow,CAPTFJobQueueSlow. - Reconcile Errors — the controller itself failing
to reconcile.
CAPTFReconcileErrors. - Identities and Credentials — an object
that cannot resolve or mirror its
TerraformClusterIdentity.IdentityAllowed=False,CredentialsMirrored=False. - Webhook Unavailable — writes to a
Terraform*object failing because the admission webhook cannot be reached. - clusterctl move — moving a Cluster’s
Terraform*objects withclusterctl move: the procedure, what does and does not come along, and cleaning up what a move leaves behind.
Every runbook above applies to TerraformCluster, TerraformMachine and
TerraformMachinePool alike unless it says otherwise.
See also
- Observability — the metrics and alerts these runbooks are reached from.
- Conditions reference — every condition type and reason named above.
Failing Jobs
This page helps you diagnose a Job that failed, or an object whose apply,
destroy, drift check or health refresh keeps failing. It covers the
CAPTFJobFailing and CAPTFNoRecentSuccess alerts and reading a failure
directly from an object’s status, for a TerraformCluster,
TerraformMachine or TerraformMachinePool.
Before you begin
kubectlaccess to read the failing object and its Job’s pod logs, in its namespace, on the management cluster.- The object’s kind, namespace and name.
CAPTFJobFailingcarries onlykindandop, so find the object through its conditions (step 1);CAPTFNoRecentSuccessalready carriesnamespaceandname.
1. Find the failing objects
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(any(.status.conditions[]?; (.type=="ApplyJobSucceeded" or .type=="DriftJobSucceeded") and .status=="False")) | "\(.kind) \(.metadata.namespace)/\(.metadata.name)"'
ApplyJobSucceeded covers apply and destroy; DriftJobSucceeded covers
drift checks and health refreshes. CAPTFNoRecentSuccess fires when a
scheduled drift check or refresh has not succeeded for six hours, which
usually means one of these Jobs is failing repeatedly rather than missing
once.
2. Read the condition’s reason
The reason on ApplyJobSucceeded or DriftJobSucceeded says where to
look next; see Conditions for the full
list of reasons each condition can carry. The ones that matter here:
ApplyFailed,DestroyFailedorDriftJobFailed: a Job ran and its runtime failed partway through.status.lastRuncarries the detail — go to step 3.ImagePullFailed: a container stayed inErrImagePullorImagePullBackOffuntil the Job’s deadline. No step ever ran, sostatus.lastRuncarries nothing useful — go to Image pull failures.ImageInvalid: the image does not follow the module contract. Go to Image layout errors.JobDeadlineExceeded: the Job’s pod ran pastactiveDeadlineSecondswithout finishing. Go to Deadline exceeded.DestructivePlanBlockedorPlanChanged: not a failure — an apply stopped on purpose, waiting for an approval, and does not count towardCAPTFJobFailing. See Plan Approval.
3. Read status.lastRun for a completed run
When the runner actually started and produced a result — every reason
above except ImagePullFailed and a JobDeadlineExceeded that hit before
any step ran — status.lastRun.error.kind classifies what happened. See
the full field list in API Reference.
step: a runtime command failed.error.stepnames it — usuallyinit,validate,plan,show-json,apply,apply-refresh-onlyordestroy; occasionallyforce-unlock(retrying past a stale lock left by a previous run; see the stale state lock runbook) orprepare(the runner’s own environment setup, before any runtime command ran) — anderror.summarygives the runner’s curated one-line reason, at most 512 bytes — the module’s or the cloud provider’s own error, never raw output. Read the rest from the Job’s pod logs (step 4).image-layout: the image does not follow the module contract; the same cause asImageInvalidabove. See Image layout errors.interrupted: the runner was sentSIGTERMbefore it could finish — a node drain, an eviction, the Job being deleted, or the deadline. It is retried without counting toward backoff; see Retry backoff.blockedorplan-changed: see Plan Approval.
status.lastRun.steps lists every step the runner completed, with its
exit code and duration — useful for finding a slow step even when the run
as a whole succeeded.
4. Read the Job’s pod logs
kubectl get job -n <ns> <job-name>
kubectl logs -n <ns> job/<job-name> -c source
<job-name> is status.lastRun.job for a finished run, or
status.activeJob.name for one still running. The source container
runs the pinned image and the module; a separate init container only
copies the runner binary into it and rarely fails on its own (that
failure surfaces as an image pull failure on the manager’s --runner-image
instead, a cluster-wide problem rather than one specific to this object).
status.lastRun.error.summary is a curated, size-limited copy of the
failing step’s output, not the raw log, because the object’s status is
readable by anyone who can get it — the full log is only in
kubectl logs.
By default the newest three failed Jobs of each operation are kept
(spec.jobs.failedJobsHistoryLimit; see
Tuning Jobs), so the pod and its logs
are usually still there for a fresh alert.
kubectl events --for <kind>/<name> -n <ns> --types=Warning also carries a
curated summary of each failure as a JobFailed or StepFailed event,
even after the Job itself is pruned.
Image pull failures
ImagePullFailed means a container was stuck pulling its image until the
Job’s deadline — never that the module ran and failed. Two images can be
at fault:
- the
sourcecontainer’s image,spec.source.image— check for a typo, a missing pull secret, or a tag that was deleted from the registry; - the runner init container’s image, the manager’s
--runner-image(defaults to the manager’s own image) — a cluster-wide problem, not specific to this object.
kubectl get pods -n <ns> -l batch.kubernetes.io/job-name=<job-name>
kubectl describe pod -n <ns> <pod-name>
The Events section names the image and the pull error
(ErrImagePull/ImagePullBackOff). Fix the image reference or the pull
credentials (spec.jobs.imagePullSecrets; see
Tuning Jobs), then let the next
reconcile retry — no manual retrigger is needed.
Image layout errors
ImageInvalid (status.lastRun.error.kind: image-layout) means the
runner started but the image itself does not follow the module contract:
it is missing the module directory the contract requires, or its entry
point is not executable. This is a problem with the image, not with the
object’s spec or a transient failure — no retry fixes it. See
Image Contract for the layout an
image must follow. Republish the image and point spec.source.image at
the new tag or digest to try again.
Deadline exceeded
JobDeadlineExceeded means the Job’s pod ran past
spec.jobs.activeDeadlineSeconds (one hour unless set) without
finishing. Kubernetes sends the runner SIGTERM; it gets most of the
Job’s termination grace period to finish an in-flight provider call and
write state before being killed, so infrastructure created before the
deadline is not lost even though the Job is marked failed. Common causes:
- a module that is simply slow relative to its deadline — raise
spec.jobs.activeDeadlineSeconds; see Tuning Jobs; - a state lock wait: the first step that locks the state (
plan,apply,apply-refresh-onlyordestroy—initnever takes the lock) waits at mostspec.jobs.lockTimeoutSeconds(five minutes unless set) for it before failing on its own — well inside the default deadline, so a lock wait alone should rarely be the cause; see Slow Jobs and the stale state lock runbook if it is; - a provider call that never returns — the module’s or the cloud API’s problem; the log up to where it stopped is the only record when the runner is killed before it can write a result.
Retry backoff
A Job that failed for step, image-layout or a pull failure counts
toward that operation’s retry backoff; a blocked or plan-changed Job
never counts, since it waits for an approval instead of a timer. See
The Reconcile Lifecycle for
the delay itself. A deadline counts too, unless the runner caught its
SIGTERM in time to report itself interrupted
(status.lastRun.error.kind: interrupted) — the common case, which
retries on the next reconcile with no wait, the same as any other
interruption. There is no way to skip a real backoff by hand; retrying
immediately against an unresolved cause (a bad image, a still-held lock)
only fails again the same way.
Confirm it worked
kubectl get <kind> -n <ns> <name> -o json \
| jq '.status.conditions[] | select(.type=="ApplyJobSucceeded" or .type=="DriftJobSucceeded")'
The condition returns to status: "True", and the next
status.lastRun.error is empty.
See also
- Slow Jobs — a Job that runs but takes far longer than expected.
- Stale State Lock — a lock held long enough to fail a Job on its own.
- Plan Approval — approving a blocked or changed plan.
- Tuning Jobs — deadlines, lock timeouts, history limits and pull secrets.
Stuck Destroy
CAPTF has no skip-destroy annotation: if a destroy Job keeps failing, the
object stays with ApplyJobSucceeded=False/DestroyFailed (so Ready=False)
and the controller retries with backoff forever. This applies the same way
to a TerraformCluster, a TerraformMachine and a TerraformMachinePool.
This page walks through recovering: back up what the destroy would remove,
clean up the cloud resources another way, then remove the object’s
finalizer by hand.
Removing the finalizer garbage-collects the state Secrets, the state backups and the durable inputs Secret through their owner references — the only record of the live cloud resources — so back them up, or un-own them, before you do that.
A stuck destroy on a control-plane machine also blocks KubeadmControlPlane/RKE2ControlPlane remediation, scale and upgrade until it is resolved.
Before you begin
get,patchanddeleteaccess to Secrets and to the stuck object, in its namespace, on the management cluster.- Replace
<ns>,<name>and<kind>below with the object’s namespace, name and Kind (terraformcluster,terraformmachineorterraformmachinepool; the lowercase singular works withkubectl). Run every command against the management cluster.
1. Find the state suffix
State Secret and backup Secret names are keyed by a suffix the controller derives from the object’s namespace, kind and name. Read it back from the object’s own status rather than recomputing it:
kubectl get <kind> -n <ns> <name> -o jsonpath='{.status.stateSecretSuffix}'
Save the output as <suffix> for the commands below.
2. Back up the state and inputs Secrets
kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix> -o yaml > backup-state.yaml
kubectl get secret -n <ns> -l captf.io/state-backup=true,captf.io/state-backup-suffix=<suffix> -o yaml > backup-state-backups.yaml
kubectl get secret -n <ns> captf-inputs-<kindshort>-<name> -o yaml > backup-inputs.yaml
<kindshort> is c for a TerraformCluster, m for a TerraformMachine, mp
for a TerraformMachinePool. The first selector matches the base state
Secret and every -part-N chunk of a large Terraform state in one call;
the second matches every backup Secret (and its chunks) taken of this
object — see Terraform State and
the state restore runbook. Both kinds of Secret are
owned by the object and are garbage-collected along with it once its
finalizer is removed.
backup-inputs.yaml’s captf.io/image-digest annotation (see
Annotations, labels and finalizers)
names the exact image that ran the last successful apply — you will need it
in step 4.
3. Un-own the Secrets, if you want them to survive finalizer removal
kubectl patch takes names, not a label selector, so patch each Secret by
name:
for s in $(kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix> -o name); do
kubectl patch -n <ns> "$s" --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]'
done
for s in $(kubectl get secret -n <ns> -l captf.io/state-backup=true,captf.io/state-backup-suffix=<suffix> -o name); do
kubectl patch -n <ns> "$s" --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]'
done
kubectl patch secret -n <ns> captf-inputs-<kindshort>-<name> \
--type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]'
Skip this if you would rather rely on step 2’s backup-*.yaml files: with
those in hand you don’t need the live Secrets to survive finalizer removal.
4. Clean up the cloud resources
Either clean up out of band (the cloud console, or the module’s own
tooling), or run the pinned image’s terraform/tofu binary by hand
against the backed-up state and inputs. The image’s own binary lives at
/captf/runtime
(Image Contract); the
Job’s own use of it, including its exact -backend-config flags, is in
Job Environment.
docker run --rm --entrypoint /captf/runtime \
-v "$PWD/root:/captf/work/root" -w /captf/work/root \
<image>@<digest> init -input=false -no-color \
-backend-config=secret_suffix=<suffix> -backend-config=namespace=<ns> \
-backend-config=in_cluster_config=true -backend-config=labels=<labels>
docker run --rm --entrypoint /captf/runtime \
-v "$PWD/root:/captf/work/root" -w /captf/work/root \
<image>@<digest> destroy -auto-approve -input=false -no-color \
-var-file=terraform.tfvars.json
<image>@<digest>is the repository frombackup-inputs.yaml’scaptf.io/imageannotation (drop any:tag) plus@and the digest from itscaptf.io/image-digestannotation.<labels>must be the same HCL object the Job itself would pass, orinitreads an empty state: the backend selects state by this whole map. Copy it frombackup-state.yaml’s base Secret.metadata.labels, droppingtfstate,tfstateSecretSuffixandtfstateWorkspace(the backend sets those itself), as{"key"="value",...}; see Annotations, labels and finalizers for what each remaining key means../rootneedsmain.tf.jsonandterraform.tfvars.jsonfrombackup-inputs.yaml’s data (the durable inputs Secret’s rendered root module and tfvars), and the identity’s credentials as environment variables (-e <KEY>=<value>for each key of the mirrored credentials Secret, or a file for a file-based one — see Identities and credentials).-backend-config=in_cluster_config=truereads the pod’s own ServiceAccount token, so this only works run from inside the cluster (for example, a debug pod using thecaptf-runnerServiceAccount). Running it from a workstation needs a kubeconfig-based backend configuration instead.- Restoring
backup-state.yamlfirst (recreate the Secrets it holds) and letting thekubernetesbackend read that live state also works, and skips reconstructing the backend configuration by hand.
5. Remove the finalizer
kubectl patch <kind> -n <ns> <name> --type=json \
-p '[{"op":"remove","path":"/metadata/finalizers"}]'
Use the lowercase, plural CRD resource name (for example,
terraformmachines) if your kubectl version needs it instead of the
Kind. This removes the whole metadata.finalizers array: safe for a
Terraform* object, which carries only CAPTF’s own finalizer, but check
.metadata.finalizers first if something else may have added one, and
remove that entry by index instead.
Confirm it worked
kubectl get <kind> -n <ns> <name>
kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix>
The first command reports NotFound once the object is gone. The second
shows nothing unless you un-owned the Secrets in step 3, in which case they
are exactly what you chose to keep.
The durable inputs Secret is missing
If the object reports ApplyJobSucceeded=False/DestroyFailed with the
message “The durable inputs Secret is missing, so destroy cannot be
rendered; see
https://captf.io/docs/operator-guide/runbooks/stuck-destroy.html”, the
durable inputs Secret (captf-inputs-<kindshort>-<name>: the rendered root
module, tfvars and pinned image from the last successful apply — see Job
Inputs) was deleted or never written.
TerraformMachine is the persistent case: it is immutable and never falls
back to re-rendering current inputs for a destroy, since the Machine and
its bootstrap Secret a rebuild would need are usually already gone by the
time destroy runs, so once its durable Secret is gone the condition never
clears on its own. TerraformCluster and TerraformMachinePool (both
mutable) fall back to building current inputs instead; they show the same
message only while that build is gated (for example, waiting on a
dependency that is itself being deleted), and it usually clears once the
gate does. The controller retries forever either way; it never invents
inputs to destroy with.
To recover:
- If you have a backup (
backup-inputs.yamlfrom a previous run of step 2 above, or any earlier copy of the durable inputs Secret), recreate it withkubectl apply -f backup-inputs.yamlafter removingmetadata.uidandmetadata.resourceVersionfrom the YAML: reapplying the exact object restores its owner reference, labels and annotations, including the pinnedcaptf.io/imageandcaptf.io/image-digest. The next reconcile reads it and starts the destroy Job. - If you have no backup, the controller cannot destroy the object’s
resources. Back up the state (step 2), then clean up out of band (step
4): without the rendered
main.tf.jsonandterraform.tfvars.json, running the module by hand needs reconstructing them, but the state lists every resource the module created.status.source.imageandstatus.source.imageDigeststill name the image the last Job ran. Then remove the finalizer (step 5). Removing the finalizer without cleaning up the cloud resources first abandons them: they stay in the cloud, unmanaged, with nothing in Kubernetes recording that they ever existed.
See also
- Terraform State — how state, its backups and their owner references work.
- State restore runbook — restoring a state instead of destroying it.
- Job Inputs — the durable inputs Secret.
Runbook: state that will not read
StateReadable reports whether the controller could read a
TerraformCluster, TerraformMachine or TerraformMachinePool’s state
this reconcile. While it is False or Unknown, no plan, apply, drift
check or refresh Job runs for the object, and Ready follows it down.
This runbook covers every reason StateReadable (and its True
counterpart) can carry, what each means, and what to do. For the
CAPTFStateUnreadable alert itself, see
Observability.
Replace <ns>, <kind> and <name> below with the object’s namespace,
kind (terraformcluster, terraformmachine or terraformmachinepool) and
name.
Find the reason
kubectl get <kind> -n <ns> <name> \
-o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}: {.message}{"\n"}{end}'
Also check for a Warning event: a reason that turns False emits one,
named either StateLost, StateLocked or, for the three read errors
below, StateUnreadable.
kubectl events --for <kind>/<name> -n <ns>
StateRead
The healthy, True reason: the state was read this reconcile. Nothing to
do.
StateNotFound
Unknown. No state Secret exists yet, because the object has never
completed an apply, or (for a mutable kind) it was deleted along with the
state on a previous destroy. This is not an error: a new object reports it
until its first successful apply, and it never fires
CAPTFStateUnreadable. If an object that used to be provisioned reports
this instead of StateRead, its state Secret is gone; see StateLost
below, which is the reason a provisioned object with a missing state
Secret carries instead.
StateLost
False. Either the state Secret of a provisioned object is missing, or
its state has no recorded inputs hash although the object is provisioned
(only possible for TerraformMachine, which is immutable and so never
rebuilds inputs to re-apply from). Either way, applying again is not safe:
for the missing-Secret case, a second apply would create a second set of
resources next to the live ones instead of managing the ones the lost
state recorded; for the missing-inputs-hash case, there is nothing to
re-apply. No Job runs, and the condition does not clear on its own.
Fix: restore the object’s newest usable state backup with the
captf.io/restore-state annotation. See the
state restore runbook, including how to list the
backups in status.stateBackups. If no backup exists, StateLost never
clears; the object’s resources still exist in the cloud but nothing in
Kubernetes can manage them until a state is restored, or reconstructed out
of band with manual recovery.
StateLocked
False. The state’s lock Lease is held by something other than this
object’s own runner — for example a workstation running terraform state rm, or a Job that died without releasing it. The condition’s message
names the holder, its operation and when it took the lock; every Job for
the object waits lockTimeoutSeconds (spec.jobs.lockTimeoutSeconds,
defaulting to 300 seconds — see Job tuning)
for it and then fails.
The controller already force-unlocks a lock whose holder pod is gone, the
next time it starts a Job for the object; most StateLocked conditions
clear on their own. See the stale-lock runbook for how
that detection works, how to tell a live holder from a stale one by hand,
and manual force-unlock.
StateEncrypted
False. The state Secret carries OpenTofu’s client-side state encryption
envelope. CAPTF v1 cannot read an encrypted state at all — not the
outputs, not the resource count, nothing — so this condition never clears
on its own and the object gets no drift checks, health checks or further
applies. Because an encrypted state fails the reader before it is ever
parsed, the manager never took a backup of it either: there is no
CAPTF-side backup to restore.
Fix: disable client-side state encryption for this object’s module (drop
the encryption configuration so future writes are plain state), and use a
workstation that holds the decryption key to decrypt the current state
and push the plaintext back with terraform state push (or tofu state push) against the same kubernetes backend configuration — the backend
config, not the state format, is what CAPTF’s reader needs to match; see
manual recovery for the
exact secret_suffix, namespace and labels the object’s backend uses.
Once the pushed state is unencrypted, the next reconcile reads it normally.
StateCorrupt
False. The reader could not treat the Secret’s payload as
gzip-compressed Terraform state JSON: the gzip or JSON decoding failed, the
decompressed size exceeded the reader’s cap, or (for a multi-Secret
chunked state) there were more chunks than the reader accepts. An
unsupported state file version is reported the same way. The condition’s
message names which of these it was. See
Chunking and size caps
for the reader’s exact limits.
A corrupt state is never backed up by the manager (backups copy only a state the reader could parse), so a backup taken before the corruption is your most recent recoverable copy. Fix: restore the newest backup from before the corruption with the state restore runbook. If nothing wrote a backup before the state became corrupt, there is no CAPTF-side recovery: rebuild the state out of band with manual recovery, using the resource list from the last known-good backup or the cloud provider’s own inventory as your guide.
StateInconsistent
False. The set of state Secrets for the object’s suffix does not add up
to one complete, contiguous state: a chunk is missing or duplicated, a
Secret in the set has an unexpected name, or the base Secret has no
tfstate data key. The condition’s message names which of these it was.
This can be transient: a runner Job writes a multi-chunk state one Secret
at a time, so a reconcile that reads mid-write sees an incomplete set and
the next reconcile, after the Job finishes, usually reads a complete one.
If it persists past the Job finishing, something outside CAPTF edited or
deleted one of the chunk Secrets by hand. As with StateCorrupt, an
inconsistent state is never backed up, so restore the newest backup from
before the inconsistency appeared with the
state restore runbook; the reader’s own limits are at
Chunking and size caps.
Confirm it worked
kubectl get <kind> -n <ns> <name> \
-o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}{"\n"}{end}'
Expect True/StateRead. status.observedStateSerial moves to the
restored or rebuilt state’s serial, and the object’s next drift check or
apply proceeds from it.
See also
State Restore
The kubernetes backend keeps only the latest state for each object; there
is no history to roll back to inside the backend itself. CAPTF keeps
versioned backups of every object’s state instead, and can push one
back into the backend when you ask it to. This applies the same way to a
TerraformCluster, a TerraformMachine and a TerraformMachinePool: all
three keep backups and restore the same way.
Before you begin
getaccess to the object and its Secrets, andpatchaccess to the object (kubectl annotatepatches it), in its namespace, on the management cluster.- Replace
<ns>,<name>and<kind>below with the object’s namespace, name and Kind (terraformcluster,terraformmachineorterraformmachinepool; the lowercase singular works withkubectl).<suffix>isstatus.stateSecretSuffix. Run every command against the management cluster.
When to use it
StateReadable=False/StateLost: the state Secret of a provisioned object is gone.StateReadable=False/StateCorruptorStateInconsistentthat does not clear on its own (a chunk deleted or overwritten by something other than the runner).- A state overwritten with the wrong content — someone ran
terraform state pushorstate rmfrom a workstation against the wrong object.
Restore is for disasters. It rewrites the object’s state, so every resource created after the backup’s serial is no longer in the state: Terraform forgets it, and it keeps running unmanaged in the cloud. The next plan shows the difference (for a machine, the next drift check). Do not use a restore to undo an ordinary change: change the inputs back instead.
For how backups are taken, how many are kept, and what is never backed up, see Terraform State.
List the backups
status.stateBackups lists up to 16 complete backups, newest first
(serial, time taken, compressed bytes — see
reference/api.md for its fields). It is
refreshed when a backup is taken or pruned and when a restore is
requested.
kubectl get <kind> <name> -n <ns> -o jsonpath='{.status.stateBackups}'
kubectl get secrets -n <ns> -l captf.io/state-backup=true,captf.io/state-backup-suffix=<suffix> \
-o custom-columns=NAME:.metadata.name,SERIAL:.metadata.annotations.captf\.io/state-backup-serial,TAKEN:.metadata.annotations.captf\.io/state-backup-taken-at
<kind> is terraformcluster, terraformmachine or
terraformmachinepool; <suffix> is status.stateSecretSuffix.
Restore
Pick the serial and annotate the object:
kubectl annotate <kind> <name> -n <ns> captf.io/restore-state=<serial> --overwrite
What happens:
- The controller finds the newest complete backup of that serial. None,
or a value that is not a serial:
RestoreJobSucceeded=False/RestoreBackupNotFound, no Job. - A restore takes precedence over apply, drift and refresh, not over
deletion: a deleting object never restores. It waits for a running Job
to finish, then takes the object’s run lease (and, for a
TerraformClusterunder the cluster operation gate, the cluster write lease; a machine’s or pool’s restore waits for itsTerraformCluster’s apply or destroy) like an apply — see run leases and the cluster operation gate. A wait shows asRestoreJobSucceeded=Unknown/WaitingForRunLease(or the cluster-gate reasons). - A restore Job starts. Its root module declares only the
kubernetesbackend, so it needs neither the durable inputs nor any provider. It runsinit(backend config as usual),state push -force(holding the state lock;-forcebecause the backup’s serial is older, or the lineage differs, or there is no state at all), andstate list. The Job fails if the listed state has no managed resource although the backup had some. - On success:
RestoreJobSucceeded=True/StateRestored, oneStateRestoredevent, the annotation is removed,status.lastRunnames the restore Job, and the next reconcile reads the restored state as usual, adopting the inputs hash the backup carried (StateAdopted— see Events). The push writes back exactly the backup’s serial rather than advancing it, and that serial is already backed up with this same content, so nothing new is written tostatus.stateBackups. - On failure:
RestoreJobSucceeded=False/RestoreFailed, oneStateRestoreFailedwarning event. The restore is not retried for the same serial: set another serial, or delete the failed restore Job to try the same one again.
Confirm
kubectl get <kind> <name> -n <ns> \
-o jsonpath='{range .status.conditions[*]}{.type}={.status}/{.reason}{"\n"}{end}'
kubectl logs job/<restore-job-name> -n <ns> -c source # state list output
kubectl events --for <kind>/<name> -n <ns>
Expect StateReadable=True/StateRead, status.observedStateSerial at
exactly the restored serial (the push writes it back unchanged, it does
not advance it), and RestoreJobSucceeded=True/StateRestored.
Then look at the difference between the restored state and reality before anything applies it:
- A
TerraformClusterorTerraformMachinePool(both mutable) whose current inputs differ from the restored state’s inputs hash applies them right after the restore (InputsChanged). Only theTerraformCluster’s apply is guarded: a plan that deletes or replaces anything stops for approval (see The destructive-plan guard), so read the blocked Job’s plan before approving it; aTerraformMachinePoolapplies unguarded. Do not pause the object to prevent this: a paused object starts no Job, the restore included. - Run a drift check to list what the restored state no longer matches;
with drift action
Reportnothing is changed. See Drift. - Resources created after the backup are not in the state. Import them
(
terraform importfrom a workstation against the object’s backend) or delete them in the cloud.
Manual recovery with no backup
If StateLost, StateCorrupt or StateInconsistent never clears because
no backup was taken before the loss (see the
state-unreadable runbook), there is no CAPTF-side
restore: reconstruct the state from a workstation against the object’s own
backend, then let the next reconcile read it back.
The object’s backend is the kubernetes backend, configured exactly:
terraform {
backend "kubernetes" {
secret_suffix = "<suffix>" # status.stateSecretSuffix
namespace = "<ns>" # the object's own namespace
in_cluster_config = false # true only from inside the cluster
config_path = "~/.kube/config" # a kubeconfig for the management cluster
labels = {
"captf.infrastructure.cluster.x-k8s.io/owner-kind" = "<kind>"
"captf.infrastructure.cluster.x-k8s.io/owner-name" = "<owner-name>"
"cluster.x-k8s.io/cluster-name" = "<cluster-name>"
"captf.io/managed" = "true"
"clusterctl.cluster.x-k8s.io/move" = ""
}
}
}
<kind>is exactlyTerraformCluster,TerraformMachineorTerraformMachinePool.<owner-name>and<cluster-name>are the object’s own name and its Cluster’s name, each verbatim if 63 characters or fewer, or the first 16 hex characters of its sha256 otherwise (the Kubernetes label value limit) — the same rulestatus.stateSecretSuffix’s own derivation follows; see Terraform State.- Getting the
labelsmap wrong does not fail loudly: the backend uses it, together with the suffix and workspace, to find the object’s existing state Secrets, so a mismatched value makes it list none and behave as if the object had no state at all, rather than erroring. Match it exactly, including the empty string on the last key.
Run terraform init with this backend block, then either terraform import each resource the cloud provider’s own inventory shows is missing,
or terraform state push <file> a state file reconstructed another way.
Once it is pushed, the next reconcile reads it back the same as a restored
backup (StateReadable turns True/StateRead): run a drift check before
anything applies again, since the object’s own record of its last-applied
inputs was lost along with the state, and the first reconcile after the
push may see the current inputs as changed.
Caveats
- Restoring an older state makes Terraform forget every resource created after that serial: those resources become unmanaged and are neither updated nor destroyed with the object.
- A backup is taken of what the backend stored. A state that was already wrong when it was written (a bad apply) is backed up just as faithfully.
- A backup older than the object’s current inputs makes a
TerraformClusterorTerraformMachinePoolre-apply its current inputs after the restore. - A paused object (or Cluster) starts no restore until it is resumed.
- Backups share the namespace’s Secret quota with everything else.
See also
- Terraform State — the backend, Secret naming and how backups are taken and pruned.
- Stuck destroy runbook — recovering when the object cannot be destroyed at all.
- Conditions reference
— every
RestoreJobSucceededreason.
Stale State Lock
Both Terraform and OpenTofu lock the kubernetes backend for the duration
of a run and release the lock on a clean exit
(Terraform State). Only a SIGKILL or a
node loss leaves the lock held with nothing left to release it — a stale
lock. This page covers finding a held lock, telling a stale one from a live
one, and clearing it. It applies the same way to a TerraformCluster, a
TerraformMachine and a TerraformMachinePool: all three use the same
backend and the same lock mechanics.
Before you begin
getaccess to Leases and Pods, andpatchaccess to the object, in its namespace, on the management cluster.- Docker or Podman, and pull access to the object’s pinned runtime image, if the automatic force-unlock in Diagnosis does not apply and you need Fix’s manual run.
- Replace
<ns>,<name>and<kind>below with the object’s namespace, name and Kind (terraformcluster,terraformmachineorterraformmachinepool; the lowercase singular works withkubectl). Run every command against the management cluster.
Symptoms
- The object reports
StateReadable=False/StateLocked, naming the holder, its operation and when it took the lock. - A Job runs longer than usual, then fails: it waited out
lockTimeoutSeconds(spec.jobs.lockTimeoutSeconds, default 300 s; see Tuning Jobs) for a lock that never cleared. - The
CAPTFForceUnlocksalert fired: the controller already force-unlocked a stale lock and moved on: not itself a problem, but worth checking why the previous Job died.
Cause
The lock is a coordination.k8s.io/v1 Lease named
lock-tfstate-default-<suffix> in the object’s namespace
(<suffix> is status.stateSecretSuffix). spec.holderIdentity carries
the lock ID; the app.terraform.io/lock-info annotation carries the JSON
lock info (ID, Operation, Who, Version, Created). Who is
<user>@<hostname>; in a Job pod the hostname is the pod name.
Job pods get a 600-second termination grace period, and on SIGTERM (a
drain, an eviction, or the Job’s activeDeadlineSeconds) the runner gives
the runtime up to 570 s to finish in-flight provider calls, write state and
release the lock before it is killed. So an ordinary deletion or drain does
not leave a stale lock; only a hard kill or a lost node does.
Diagnosis
Before creating a Job, the controller already checks the lock and acts on what it finds:
- No holder: proceeds; nothing to do.
- Holder present, and its pod is one of the object’s own runner Job pods
and still exists (not finished): the lock is live; the next Job waits out
lockTimeoutSecondsfor it. - Holder present, and its pod is one of the object’s own runner Job pods
but no longer exists: stale. The controller passes
--force-unlockto the next Job, which force-unlocks it afterinit, emits an event, and continues. - Holder present, but it is not recognizable as one of the object’s own
runner pods — the
Whofield has no@, or its hostname is not a pod of this object’s Jobs, for example a workstation runningterraform state rm: never force-unlocked automatically, whatever else is true. The object reportsStateReadable=False/StateLockednaming the holder, its operation and when it took the lock.
In most cases waiting for the object’s next reconcile (it retries with backoff) resolves a stale lock without any manual step. Manual inspection and force-unlock below are for the last case, where the controller will never act on its own.
To inspect the lock by hand:
kubectl get lease -n <ns> lock-tfstate-default-<suffix> -o yaml
Read spec.holderIdentity (the lock ID) and the app.terraform.io/lock-info
annotation. The part of Who after the last @ is the pod name; confirm
whether it still exists and has finished:
kubectl get pod -n <ns> <pod-name-from-who>
- Pod exists and has not finished: the lock is live. Do not force unlock; a concurrent run against the same state would be unsafe. Let the holder finish, or find out why it is still running.
- Pod is gone, or finished (
Succeeded/Failed): if it is one of this object’s own runner pods, the controller force-unlocks it automatically the next time it creates a Job (trigger a reconcile, for example by waiting for the next requeue, or by editing an annotation). If it is not one of this object’s own runner pods, force-unlock by hand below. Whohas no@(holder unknown): the controller never force-unlocks this automatically. Confirm independently that no process still holds the lock, then force-unlock by hand.
Fix
Run the pinned image’s own terraform/tofu binary against the backend,
the same pattern as the stuck-destroy runbook’s manual
recovery:
docker run --rm --entrypoint /captf/runtime \
-v "$PWD/root:/captf/work/root" -w /captf/work/root \
<image>@<digest> init -input=false -no-color \
-backend-config=secret_suffix=<suffix> -backend-config=namespace=<ns> \
-backend-config=in_cluster_config=true -backend-config=labels=<labels>
docker run --rm --entrypoint /captf/runtime \
-v "$PWD/root:/captf/work/root" -w /captf/work/root \
<image>@<digest> force-unlock -force <lock-id>
<image>@<digest> comes from the object’s durable inputs Secret
(captf.io/image/captf.io/image-digest; drop any :tag from image,
append @ and the digest). <labels> must be the same HCL object the Job
itself would pass, or init reads an empty state: read it off the state
Secret’s own .metadata.labels the same way as the stuck-destroy
runbook. force-unlock
touches no resources, so ./root/main.tf.json needs only a backend
declaration, not the full rendered module:
{
"terraform": {
"backend": {
"kubernetes": {}
}
}
}
(an empty ./root/terraform.tfvars.json, {}, avoids a missing-file
warning, though force-unlock never reads it). force-unlock needs an
initialized backend first, unlike apply or destroy, so init runs
first with the same -backend-config flags the Job itself would pass —
see Job Environment
for the exact set, and note that in_cluster_config=true needs this run
from inside the cluster. This clears the holder and the lock-info
annotation; it does not delete the Lease object itself — only a workspace
delete removes it, and the default workspace (the only one CAPTF ever
uses) cannot be deleted.
Confirm
kubectl get lease -n <ns> lock-tfstate-default-<suffix> \
-o jsonpath='{.spec.holderIdentity}'
Empty output means the lock is clear. The object’s next reconcile starts a
Job normally; StateReadable clears to True/StateRead once the
reconcile after that Job finishes reads the state.
See also
- Terraform State — the backend, its Secret naming and how the lock fits in.
- Stuck destroy runbook — the same manual-run pattern, for a destroy that cannot succeed.
- Conditions reference —
every
StateReadablereason.
Runbook: state or inputs near a size limit
CAPTF stores an object’s state and its rendered module inputs in
Kubernetes Secrets, which cap out well before an unbounded Terraform or
OpenTofu project would. Two alerts warn before that cap:
CAPTFStateNearSecretLimit and CAPTFInputsNearLimit; a third condition
reason, InputsTooLarge, marks the point where the cap is already hit.
This runbook covers all three: what each limit is, how to find the object
tripping it, and how to shrink what it is trying to store.
Replace <ns>, <kind> and <name> below with the object’s namespace,
kind (terraformcluster, terraformmachine or terraformmachinepool)
and name.
CAPTFStateNearSecretLimit: the state is approaching a Secret’s size cap
The limit. A Kubernetes Secret holds at most 1 MiB. OpenTofu’s
kubernetes backend writes the whole state into a single Secret and fails
to save a state larger than that; Terraform’s splits an oversized state
across additional -part-N Secrets instead, up to the reader’s own cap of
32 chunks. Either way, captf_state_bytes (the compressed size) crossing
900 KiB fires the alert — for OpenTofu that is a warning before a hard
failure, for Terraform a warning about growth. See
Chunking and size caps
for the exact numbers and how CAPTF reads a chunked state.
Find it.
kubectl get <kind> -n <ns> <name> -o jsonpath='{.status.stateSecretSuffix}{"\n"}'
kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix> -o name
One result means an unchunked (or not yet chunked) state; more than one
means Terraform has already split it into -part-N Secrets. For the exact
compressed size, read captf_state_bytes{namespace="<ns>",name="<name>"}
from Prometheus, or sum the chunks by hand:
kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix> \
-o jsonpath='{range .items[*]}{.data.tfstate}{"\n"}{end}' | while read -r c; do
echo "$c" | base64 -d | wc -c
done
Compare the total against captf_state_resources (also per object) to see
how many managed resources it is spread across.
Shrink it. The state holds every managed resource’s full attribute set, not just what you set in configuration:
- Split a large module into more than one
TerraformClusterorTerraformMachine(or use aTerraformMachinePoolfor repeated instances instead of one resource per machine): each gets its own state. - Drop attributes you do not need from the state by not managing them: large inline blocks such as embedded certificates, rendered cloud-init or user-data templates, or full API responses stored in a resource’s computed attributes are common culprits. Move that content to a source the module reads at apply time (a Secret, a bucket object) instead of a resource argument Terraform tracks verbatim.
- For a resource with many similar instances (
countorfor_each), each instance’s full attribute set is stored once; reducing the count reduces the state proportionally.
CAPTFInputsNearLimit and InputsTooLarge: the rendered inputs are approaching or over the cap
The limit. The rendered main.tf.json plus terraform.tfvars.json
(what the module actually runs against) is capped at 1,000,000 bytes,
reported per object on captf_inputs_bytes. CAPTFInputsNearLimit warns
above 900,000 bytes; at or above the cap, no Job starts and
ApplyJobSucceeded turns False with reason InputsTooLarge. See
Limits for how the size is computed
and what counts toward it.
Find it.
kubectl get <kind> -n <ns> <name> \
-o jsonpath='{range .status.conditions[?(@.type=="ApplyJobSucceeded")]}{.status}/{.reason}: {.message}{"\n"}{end}'
InputsTooLarge’s message names the operation that could not render. The
captf_inputs_bytes gauge is set to the rendered size even when it was too
large to run, so it is set for this object right up to the cap.
Shrink it. The three contributors are bootstrap data, cluster exports and user variables:
bootstrap_data(TerraformMachineandTerraformMachinePoolonly): the bootstrap provider’s cloud-init or ignition content, carried base64-encoded (roughly 4/3 its raw size). Trim what the bootstrap provider generates — fewer files, less embedded content — through its own configuration (KubeadmConfig, RKE2Config, or your bootstrap provider’s equivalent), not through CAPTF.captf_cluster_outputs(machines and pools): whatever the cluster module’sexportsoutput publishes, handed to every machine and pool of the cluster. See What a cluster hands to its machines and pools. Publish only what the machine or pool module actually needs fromexports, not the cluster module’s full internal state.- User variables (
spec.variables,spec.variablesFrom): trim large inline values, especially ones duplicated across many objects that could instead reference aConfigMaporSecretyour module reads directly at run time rather than passing through tfvars. See Module Variables for how variables are set and merged.
Confirm it worked
kubectl get <kind> -n <ns> <name> \
-o jsonpath='{.status.conditions[?(@.type=="ApplyJobSucceeded")].reason}{"\n"}'
For InputsTooLarge, the next reconcile after the inputs shrink re-renders
and, once it fits, starts the Job; the alerts clear once
captf_state_bytes or captf_inputs_bytes drops back under 900 KiB / 900,000
bytes for their for window.
See also
Slow Jobs
This page helps you find out why a TerraformCluster, TerraformMachine
or TerraformMachinePool’s Jobs are taking a long time — either to run,
or to start. It covers the CAPTFJobSlow and CAPTFJobQueueSlow alerts.
The two measure different things and point at different causes, so start
by working out which one fired.
Before you begin
kubectlaccess to the object’s namespace and to Leases and pods there, on the management cluster.- Both alerts carry
kindandop, not an object name; find the specific object throughstatus.activeJobor the Jobs of that kind and op.
1. Tell the two alerts apart
CAPTFJobSlowmeasures the wall time of a completed Job, start to finish. Go to A Job that runs long.CAPTFJobQueueSlowmeasures the time from a Job’s creation to its module container actually starting. Go to A Job slow to start.
Neither counts time spent waiting before a Job is even created — that
wait shows up as a condition stuck in an Unknown state instead. If
status.activeJob names no Job yet the object is not idle either, go to
Waiting for a lease first.
Waiting for a lease
Before creating a Job, the controller takes the object’s own run lease
(captf-run-<suffix>), and, for an apply, destroy or restore that names a
cluster, the Cluster’s write lease (captf-cluster-<hash>) as well. See
Run leases and the cluster operation
gate
for how the two work together, and Configuration
for the --cluster-operation-gate flag. While either is held by someone
else, ApplyJobSucceeded (or DriftJobSucceeded for a refresh or drift,
RestoreJobSucceeded for a restore — see
Restoring State) stays Unknown, with one of these
reasons, and no Job exists yet:
WaitingForRunLease: another live Job — usually another manager instance’s — already holds this object’s run lease.WaitingForClusterOperation: aTerraformMachine’s orTerraformMachinePool’s operation is waiting for itsTerraformCluster’s to finish first.WaitingForMachineOperations: aTerraformCluster‘s operation is waiting for its machines’ and machine pools’ applies and destroys, already in flight, to finish first; any new one of theirs waits for the cluster in turn.
The condition’s message names the Lease and the Job (or manager) holding it, and the controller checks again every 30 seconds — this is normal and usually resolves on its own once the holder finishes.
kubectl get lease -n <ns> -l captf.io/lease=run,captf.infrastructure.cluster.x-k8s.io/owner-name=<name>
kubectl get lease -n <ns> -l captf.io/lease=cluster
spec.holderIdentity is the holder’s Job name; the
captf.io/lease-op and captf.io/lease-acquired-at annotations say what
it is running and since when. A lease is only ever taken over by the
controller itself, once its holder Job has finished, does not exist after
a short grace period, or is older than its own deadline plus a backstop
— never delete or edit a Lease by hand: if its holder Job is genuinely
stuck, that is a failing Job or a
slow one to chase down instead, not a lease
problem.
A wait that outlasts the holder’s activeDeadlineSeconds means the
holder cannot finish; look at that Job directly.
A Job that runs long
CAPTFJobSlow fires on the Job’s own duration, so the module (or the
cloud it calls) is doing the work, not the controller withholding
anything. Narrow it down:
# p90 duration by step, for this kind and op
histogram_quantile(0.9, sum by (le, step) (rate(captf_job_step_duration_seconds_bucket{kind="<kind>",op="<op>"}[6h])))
status.lastRun.steps on the object itself shows the same breakdown for
its most recent run, without a Prometheus query. Once you know which step
is slow:
plan,apply,apply-refresh-onlyordestroyis slow: these are the steps that lock the state (initnever does), so a slow but successful one may have been contending for the lock — every locking step waits at mostspec.jobs.lockTimeoutSeconds(five minutes unless set) for it before failing on its own. CheckStateReadablefor reasonStateLockedand see the stale state lock runbook. Otherwise it is a large or slow module, or a cloud provider that is itself slow — nothing in the controller to tune. Compare againstspec.jobs.activeDeadlineSeconds(Tuning Jobs) before it starts hitting the deadline and turning into a failing Job instead of a slow one.- every step is proportionally slow: check the Job pod’s own resource
use against
spec.jobs.resources(Tuning Jobs) — a Terraform or OpenTofu process with several providers is memory-hungry, and throttling or swapping shows up as slowness everywhere, not one step.
A Job slow to start
CAPTFJobQueueSlow covers everything between the Job existing and its
module container running: pod scheduling, both image pulls (the runner
init container’s and the module’s own), and the runner binary copy.
kubectl get pods -n <ns> -l captf.infrastructure.cluster.x-k8s.io/op=<op> --field-selector=status.phase=Pending
kubectl describe pod -n <ns> <pod-name>
Read the pod’s Events:
FailedScheduling: exhausted node capacity, aResourceQuota, or a scheduling constraint (affinity, taints) the runner’s pod cannot satisfy. Checkspec.jobs.resources,spec.jobs.podSecurityContextoverrides and the namespace’s quota; see Tuning Jobs.Pulling/ErrImagePull/ImagePullBackOff: a slow, rate-limited or unreachable registry. A pull that never succeeds eventually hits the Job’s deadline and becomes a failing Job instead; a pull that succeeds but is merely slow is this alert’s whole story.
A controller’s own concurrency limit
(--terraformcluster-concurrency, --terraformmachine-concurrency,
--terraformmachinepool-concurrency, ten each unless set; see
Configuration) does not appear here: it bounds how
many objects of a kind reconcile at once, not how fast a Job someone
already created gets scheduled and pulled. A concurrency limit too low
for the fleet shows up as objects going longer between reconciles overall
— check the object’s own status.observedGeneration and event timestamps
for that, not this alert.
Confirm it worked
kubectl get <kind> -n <ns> <name> -o json | jq '.status.lastRun.steps'
The next run’s step durations (or captf_job_duration_seconds and
captf_job_queue_seconds for the kind and op) drop back under the
alerts’ thresholds — 30 minutes p90 duration, 5 minutes p90 queue time.
See also
- Failing Jobs — a Job that stops instead of merely running long.
- Stale State Lock — clearing a lock a Job is waiting on.
- The Reconcile Lifecycle — run leases and the cluster operation gate in full.
- Tuning Jobs — deadlines, lock timeouts and resources.
Reconcile Errors
This page helps you diagnose the CAPTFReconcileErrors alert: one of the
manager’s own controllers is repeatedly failing to reconcile, as opposed
to an object’s Job failing (see Failing Jobs). A
reconcile error means the controller could not even finish deciding what
to do — it is not a statement about any one object’s infrastructure.
Before you begin
kubectlaccess to the manager’s Deployment and its logs, in thecaptf-systemnamespace (or wherever it is installed), on the management cluster.
1. Read the manager’s logs
kubectl logs -n captf-system deploy/captf-controller-manager --tail=200
Each reconcile that returns an error logs a line carrying controller
(one of terraformcluster, terraformmachine, terraformmachinepool,
terraformmachinetemplate or terraformclusteridentity — the alert’s
controller=~"terraform.*" matches all five), the object’s namespace
and name, and the error itself. Reconciles for a TerraformMachine,
TerraformCluster or TerraformMachinePool also carry that object’s own
key (TerraformMachine, TerraformCluster, TerraformMachinePool) plus
its owning Cluster, Machine or MachinePool where one exists, so you
can filter by any of those instead of grepping free text.
The default log level (-v=2) already includes the manager’s own flow
logging in addition to errors, so raising -v is rarely needed just to
see that a reconcile failed; it helps to see why a decision was made
before the error, not to see the error itself. --diagnostics-address’s
endpoint can change the running level without a restart when
--insecure-diagnostics is not set; see
Manager Flags.
2. Narrow down which object, and how often
sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m]))
matches the alert’s own query. If it is one object erroring repeatedly, its logs name it on every line; if it is spread across many objects of one kind, look for a cause common to the whole namespace or cluster (quota, RBAC, an API server problem) rather than the object’s own spec.
Common causes
- The Kubernetes API server is unavailable or throttling requests: a
Get,Create,UpdateorPatchfailed with a server-side error (500,503, or a429from client-side rate limiting). Transient — the error clears once the API server does; if it does not, check the API server’s own health. - The manager’s own patch was rejected by the validating webhook: the
manager itself writes a
Terraform*object’s finalizer and a handful of spec fields (for example aTerraformMachine’sproviderID, aTerraformCluster’scontrolPlaneEndpoint) through the same admission path a user’skubectl applygoes through. If the webhook cannot be reached, that patch fails and shows up here. See Webhook Unavailable. - The manager’s own RBAC is missing a permission: the
ClusterRoleit runs as (captf-manager-role) was edited by hand, or an upgrade changed what it needs and the installedClusterRolewas not updated to match. The error names the verb, resource and group a403Forbiddenresponse refused. See RBAC for what the manager needs and creates. - A namespace
ResourceQuotablocks a create: the manager creates Jobs, Secrets, Leases, and a per-namespace runnerServiceAccountandRoleBindingon an object’s behalf. A quota on any of those object counts in the tenant namespace surfaces as aForbiddencreate error here rather than anywhere on theTerraform*object’s own status. - A conflicting concurrent write: two updates to the same object
raced (for example, the manager and an operator editing it at the same
moment). The condition-patching helper already retries a conflicting
status write itself; a
Conflictthat still reaches the log usually clears on the next reconcile, which controller-runtime’s own per-item backoff already schedules — distinct from a Job’s retry backoff (see Failing Jobs), and not something to act on unless it repeats for the same object.
None of these are the same as an object’s Job failing: a Job failure is
recorded on the object’s own status and conditions and never increments
controller_runtime_reconcile_errors_total — only an error the
reconciler itself returns does. See Failing Jobs if
what you are chasing is instead a failing apply, destroy, drift check or
refresh.
Confirm it worked
kubectl logs -n captf-system deploy/captf-controller-manager --tail=50 --follow
No further Reconciler error lines appear for the affected controller,
and sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m]))
returns to 0.
See also
- Failing Jobs — a Job or an object’s status, as opposed to the controller itself.
- Webhook Unavailable — a specific, common cause of reconcile errors.
- RBAC — what the manager’s own
ClusterRolegrants.
Runbook: identity, credentials or RBAC not ready
Before a TerraformCluster, TerraformMachine or TerraformMachinePool
can run a Job, the controller must resolve its
identity, mirror that identity’s
credentials into the object’s namespace, and make sure a runner
ServiceAccount exists and is bound to the runner ClusterRole. Three
conditions cover those steps: IdentityAllowed, CredentialsMirrored and
RunnerRBACReady. While any of them is not True, no Job starts. This
runbook covers every False and Unknown reason they carry, and how to
fix each.
Replace <ns>, <kind> and <name> below with the object’s namespace,
kind (terraformcluster, terraformmachine or terraformmachinepool) and
name.
Find the reason
kubectl get <kind> -n <ns> <name> -o jsonpath='{range .status.conditions[?(@.type=="IdentityAllowed" || @.type=="CredentialsMirrored" || @.type=="RunnerRBACReady")]}{.type}={.status}/{.reason}: {.message}{"\n"}{end}'
IdentityAllowed
IdentityNotFound
False. Either the object (and, for a machine or pool, its
TerraformCluster’s spec.defaults.identityRef) sets no identityRef at
all, or identityRef.name names a TerraformClusterIdentity that does
not exist. Fix: set identityRef.name (on the object, or on the
cluster’s spec.defaults for a machine or pool that should inherit it) to
an existing TerraformClusterIdentity, or create the one it already
names. See Reference it.
NamespaceNotAllowed
False. The named identity exists, but its spec.allowedNamespaces
does not include this object’s namespace. Fix: add the namespace to the
identity’s allowedNamespaces (its list, or a selector matching one
of the namespace’s labels). See
Create the identity
for the field’s exact semantics, including the empty-list-versus-empty-
selector distinction.
SecretNotFound
False. The identity exists and allows this namespace, but its
spec.secretRef Secret does not exist (or was deleted after the identity
was created and admission’s SubjectAccessReview check passed). Fix:
create the credentials Secret at the namespace and name spec.secretRef
names, or point spec.secretRef at one that exists. See
Create the credentials Secret.
IdentityCheckFailed
Unknown. A transient error while checking the identity or its Secret —
reading the TerraformClusterIdentity, evaluating a selector against
the namespace’s labels, or reading the credentials Secret all failed for
a reason other than not-found (an API server error, for example). Unlike
the three False reasons above, this is not a configuration problem: the
reconcile itself returns an error and retries with backoff, so it usually
clears on its own. If it persists, read the manager’s logs for the
underlying error.
CredentialsMirrored
MirrorPending
Unknown. Set whenever IdentityAllowed is False — no identity
resolved yet, the namespace is not (or no longer) allowed, or the
credentials Secret is missing — and before the first mirror is ever
written. Not an error condition in itself: fix the IdentityAllowed
reason above, and CredentialsMirrored follows it.
MirrorFailed
False. Either the mirror Secret captf-creds-<identity> could not be
created or updated (an API error, named in the message), or a Secret with
that exact name already exists in the namespace but is not a mirror of
this identity: it lacks the captf.io/mirrored label, or its
captf.io/identity annotation names a different identity. The controller
never overwrites a Secret it does not recognize as its own mirror, so this
does not clear on its own.
Fix the conflict case by renaming or removing whatever created the
conflicting Secret — most often a Secret created by hand or by another
tool using the same name CAPTF would mirror to. See
How credentials reach a Job
for the exact mirror name (captf-creds-<identity>, or a truncated hash
form for a very long identity name) and what the mirror carries. For any
other MirrorFailed message, it names the underlying API error; retry
after fixing that (for example, a namespace quota or a webhook rejecting
the write).
RunnerRBACReady
RBACFailed
False. Creating or updating the runner ServiceAccount or the
captf-runner RoleBinding failed. The message names the error. One
specific cause: a RoleBinding named captf-runner already exists in the
namespace without captf.io/managed=true — the controller never modifies
a RoleBinding it does not own, since a binding it did not create could
carry subjects or a RoleRef from something else. Fix that case by
renaming or removing the conflicting RoleBinding; CAPTF then creates its
own. For any other message, it is an API error (permissions, quota, a
webhook); the manager’s own RBAC to manage these objects is set up as part
of installation. See
RBAC for what the controller creates and why.
ServiceAccountNotOptedIn
False. The object’s effective spec.jobs.serviceAccountName names a
ServiceAccount other than the default captf-runner, and that
ServiceAccount either does not exist or does not carry
captf.io/runner=true. CAPTF treats that label as the namespace’s consent
to bind the ServiceAccount to the runner ClusterRole; it is never
inferred. Fix: label the ServiceAccount (kubectl label serviceaccount -n <ns> <name> captf.io/runner=true), create it if it does not exist, or
remove the override from spec.jobs.serviceAccountName (and the
cluster’s spec.defaults.jobs.serviceAccountName, if that is where it
came from) to use the default captf-runner instead. See
Job tuning for
spec.jobs.serviceAccountName and its default inheritance.
Confirm it worked
kubectl get <kind> -n <ns> <name> -o jsonpath='{range .status.conditions[?(@.type=="IdentityAllowed" || @.type=="CredentialsMirrored" || @.type=="RunnerRBACReady")]}{.type}={.status}/{.reason}{"\n"}{end}'
Expect IdentityAllowed=True/IdentityAllowed,
CredentialsMirrored=True/Mirrored and
RunnerRBACReady=True/RBACReady. The next reconcile after all three are
True starts a Job if one is otherwise due.
See also
Webhook Unavailable
CAPTF validates every TerraformCluster, TerraformClusterIdentity,
TerraformClusterTemplate, TerraformMachine, TerraformMachineTemplate,
TerraformMachinePool and TerraformMachinePoolTemplate write with an
admission webhook, and every one of those webhook rules has
failurePolicy: Fail: if the webhook cannot be reached, the write is
refused rather than let through unchecked. This page covers recognizing
that, and getting the webhook serving again.
There is no dedicated alert for this: it shows up as errors on
kubectl apply (or on anything else writing a Terraform* object) and,
because the manager’s own writes go through the same path, often as
CAPTFReconcileErrors
as well.
Before you begin
kubectlaccess to thecaptf-systemnamespace (or wherever CAPTF is installed): its Deployment, Service, Endpoints and thecert-managerCertificateandSecretit depends on.
What is blocked while the webhook is down
Every CREATE and UPDATE of the seven kinds above is blocked: nothing
new can be created, and no existing one can be changed — including by
Cluster API’s own controllers. In practice that reaches further than a
person running kubectl apply:
- a
MachineDeploymentorKubeadmControlPlane/RKE2ControlPlanescaling up cannot create theTerraformMachines for the new replicas; clusterctl movefails, since it creates and updatesTerraform*objects on the target cluster;- the manager’s own reconciles that write to a
Terraform*object’s metadata or spec — adding the finalizer to a new object, removing it once deletion finishes, or writing back a resolvedproviderIDorcontrolPlaneEndpoint— fail the same way, and surface as CAPTFReconcileErrors. - deleting a
TerraformClusterIdentityor mostTerraformMachines is also blocked: those two kinds validateDELETEas well asCREATE/UPDATE. The identity’s check normally refuses a delete while it is still in use or its credentials are still mirrored somewhere; the machine’s normally redirects a direct delete through its ownerMachineinstead, to go through drain. Either check can also be the thing allowing a delete that would otherwise be refused (an identity no longer in use, a machine whose ownerMachineis already gone), so while the webhook is down those deletes are blocked outright rather than falling back to permissive.
What keeps working: everything that only patches an object’s status
subresource — a Job finishing, status.lastRun, conditions, drift and
health sampling — since none of the webhook rules cover the status
subresource. Existing, already-running Jobs finish normally; only
changes to an object’s metadata or spec are affected.
1. Confirm the webhook, not something else, is the cause
kubectl apply against a Terraform* object returns an error naming the
webhook by name (captf-validating-webhook-configuration). Map the error
text to a cause:
| Error text | Likely cause |
|---|---|
... failed calling webhook ...: no endpoints available for service "captf-webhook-service" | No manager pod is Ready; go to step 2. |
... failed calling webhook ...: context deadline exceeded or connection refused | The Service or pod is reachable but not serving on the expected port, or a NetworkPolicy blocks it; go to step 3. |
... x509: certificate signed by unknown authority | The webhook’s caBundle is empty or stale; go to step 4. |
2. Check the manager pod is Ready
kubectl get pods -n captf-system -l control-plane=controller-manager
kubectl get endpoints -n captf-system captf-webhook-service
The manager’s readiness probe includes its webhook server: a pod that
has not finished starting the webhook server, or whose serving
certificate failed to load, reports NotReady and drops out of the
captf-webhook-service Endpoints — which is exactly why
kubectl get endpoints shows nothing while this is the cause. Read the
pod’s own logs and events for why it is not ready (a crash, an image
pull failure, or the certificate problem in step 4).
3. Check the Service and NetworkPolicy
kubectl get svc -n captf-system captf-webhook-service -o yaml
kubectl get networkpolicy -n captf-system
The Service forwards port 443 to the manager container’s :9443; a
Ready pod with a populated Endpoints list but a webhook that still
cannot be reached from the API server usually means a NetworkPolicy
(none is installed by default; see Installation)
blocking traffic to that port, or the Service’s selector no longer
matching the pod’s labels after a manual edit.
4. Check the certificate
kubectl get certificate -n captf-system captf-serving-cert
kubectl describe certificate -n captf-system captf-serving-cert
kubectl get secret -n captf-system captf-webhook-service-cert
cert-manager issues the webhook’s serving certificate as the Secret
captf-webhook-service-cert, mounted into the manager pod, from the
Certificate captf-serving-cert. The
ValidatingWebhookConfiguration’s caBundle is kept current by
cert-manager’s CA injector, driven by the
cert-manager.io/inject-ca-from: captf-system/captf-serving-cert
annotation on captf-validating-webhook-configuration itself — check
that annotation is still present (a kubectl apply of a stripped-down
copy of the manifest can remove it) and that the Certificate reports
Ready. cert-manager not running at all, or its CRDs missing, leaves
the Certificate object present but never issued.
See Installation for how these names and the
cert-manager dependency fit together, and confirm cert-manager itself
is healthy in its own namespace if the Certificate never becomes
Ready.
Confirm it worked
kubectl apply --dry-run=server -f - <<'EOF'
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformClusterIdentity
metadata:
name: webhook-check
spec:
secretRef:
name: webhook-check
namespace: default
EOF
A server-side dry run still goes through admission. It succeeding (or failing with a validation error about the object’s own content, not a webhook connectivity error) confirms the webhook is reachable again; a blocked create or scale-up elsewhere in the cluster starts progressing on its own once it does.
See also
- Reconcile Errors — the manager’s own writes failing for the same reason.
- Installation — what
clusterctl initinstalls, and the exact resource names used above.
clusterctl move
clusterctl move refuses to run against a Cluster that is already paused
— it stops with an error naming the paused Clusters — and it pauses the
Cluster itself for the duration of the move and unpauses it afterwards. So
do not leave the source Cluster paused when you run it. This page walks
through moving a Cluster’s Terraform* objects safely, what does and does
not come along, and cleaning up what a move leaves behind.
Before you begin
clusterctlconfigured with both the source and target management cluster’s kubeconfigs.- The target cluster running the same or a newer CAPTF version.
- Replace
<ns>and<name>with the Cluster’s namespace and name below.
Procedure
The documented sequence pauses first only to drain in-flight Jobs, then
unpauses before invoking clusterctl move (which pauses and unpauses it
again on its own):
-
Pause the Cluster so the controller starts no new Job on any
Terraform*object of it:kubectl patch cluster -n <ns> <name> --type=merge -p '{"spec":{"paused":true}}'A Job already running is left to finish, unless it never started at all — its per-run inputs Secret is missing and every pod it created is still
Pendinga minute after the Job was created — in which case the paused reconcile deletes it so the next reconcile starts it again. The paused reconcile still does Job bookkeeping and clearsclusterctl.cluster.x-k8s.io/block-moveonce no Job is active, even while paused (the clusterctl move block). -
Wait until no
Terraform*object of the Cluster carriesblock-move:kubectl get terraformclusters,terraformmachines,terraformmachinepools -n <ns> \ -l cluster.x-k8s.io/cluster-name=<name> \ -o custom-columns='KIND:.kind,NAME:.metadata.name,BLOCK-MOVE:.metadata.annotations.clusterctl\.cluster\.x-k8s\.io/block-move'Repeat until every row’s
BLOCK-MOVEcolumn is<none>. The annotation is set before a Job is created and removed once no Job is active, so a cleared annotation means the object has no Job in flight. -
Unpause the Cluster — required, or step 4 fails immediately:
kubectl patch cluster -n <ns> <name> --type=merge -p '{"spec":{"paused":null}}' -
Run
clusterctl moveto the target management cluster:clusterctl move --namespace <ns> --to-kubeconfig <target-kubeconfig>
clusterctl move’s own backoff does not wait for long Jobs (about two
minutes, while jobs.activeDeadlineSeconds defaults to one hour), which is
why step 2 waits explicitly before the move itself starts.
Moving every namespace
clusterctl move always moves one namespace: an unset --namespace falls
back to your kubeconfig context’s current namespace, not to every
namespace on the cluster, and there is no flag that moves all of them in
one invocation. Moving every namespace means repeating the procedure above
once per namespace.
A TerraformClusterIdentity is cluster-scoped and several namespaces can
reference the same one. This is safe to move namespace by namespace:
clusterctl move never deletes a cluster-scoped object from the source,
only namespaced ones, so the first namespace’s move copies the identity to
the target and leaves it in place on the source; each later namespace’s
move finds it already on the target and skips re-creating it rather than
erroring. You do not need to move the namespaces that share an identity in
any particular order, and you still copy its credentials Secret to the
target only once (see below),
however many namespaces that use it you move.
What moves and what doesn’t
clusterctl move discovers CRD-kind objects, ConfigMaps and Secrets in a
Cluster’s owner reference chain, copies them to the target, then strips
finalizers and deletes them from the source. The controller’s own delete
path therefore never runs on the source.
| Moves | Doesn’t move |
|---|---|
State Secrets (tfstate-default-<suffix> and its chunks) — owned by the object | The state lock Lease (lock-tfstate-default-<suffix>) — not a discovered kind. Recreated by the backend at the target’s next init. |
State backups (captf-state-backup-<suffix>-<serial> and their chunks) — owned by the object | The runner ServiceAccount (captf-runner) and its RoleBinding — not discovered kinds. Recreated by the controller before the target’s first Job. |
The durable inputs Secret (captf-inputs-<kindshort>-<name>) — owned by the object | The run lease and, for a TerraformCluster, the cluster write lease — not discovered kinds. A lease whose Job was active during the move is left on the source; see run leases. |
The mirrored credentials Secret (captf-creds-<identity>) | The identity’s own credentials Secret — deliberately not owned and not labeled for move. Copy it to the target yourself; see below. |
Every variablesFrom ConfigMap or Secret — deliberately not owned. Recreate it on the target, or label it for move yourself (see Module Variables). |
The Secrets in the “Moves” column carry the clusterctl.cluster.x-k8s.io/move
label themselves, which is what makes clusterctl move discover and carry
them alongside the Terraform* object that owns them, on top of the
owner-reference chain it always follows; see
Annotations and labels for that
label and every other one this page’s objects carry.
status is never restored by move: it is rebuilt from state on the
target’s first reconcile.
A TerraformCluster or TerraformMachinePool (both mutable) re-reads its
variablesFrom sources on every reconcile, so it waits at
DependenciesReady=False/VariablesSourceNotFound on the target until the
source exists there, whether or not it was already provisioned. A
TerraformMachine (immutable) stops needing its sources once provisioned,
so only one moved before its first apply waits.
The identity’s credentials Secret does not move
The TerraformClusterIdentity object itself is cluster-scoped and carries
a move-hierarchy label, so clusterctl moves it as part of the global
hierarchy every namespace’s move includes. Its credentials Secret is
different: it is namespace-scoped and used by every namespace the identity
allows, so CAPTF deliberately removes any owner reference it might have
added and never labels it for clusterctl move. Labeling it for move
yourself would make clusterctl delete it from the source once the move
of the namespace it lives in finished — deleting credentials that other
allowed namespaces on the source still use, and clusterctl move never
moves more than one namespace per invocation regardless (see Moving every
namespace above), so there is no invocation that
“finishes last” to safely hang the deletion off. Copy it yourself every
time, on every namespace’s move; it is a duplicate, not a move.
Because the identity object moves but its Secret does not, the identity
reports Ready=False/SecretNotFound on the target until you copy the
Secret there yourself, keeping its name and the keys the identity names;
every TerraformCluster, TerraformMachine and TerraformMachinePool that
uses it reports IdentityAllowed=False/SecretNotFound for the same
reason. kubectl get -o yaml | kubectl apply -f - carries the source
Secret’s resourceVersion and uid along, which apply rejects against a
target that has no existing object with that resourceVersion; strip the
identifying metadata first:
kubectl --kubeconfig <source-kubeconfig> get secret -n <ns> <secret-name> -o json \
| jq 'del(.metadata.resourceVersion, .metadata.uid, .metadata.creationTimestamp, .metadata.managedFields, .metadata.ownerReferences)' \
| kubectl --kubeconfig <target-kubeconfig> apply -f -
See Identities and Credentials for creating, rotating and revoking that Secret.
TerraformMachine deletion during a move
clusterctl annotates each object with
clusterctl.cluster.x-k8s.io/delete-for-move before deleting it on the
source. The TerraformMachine delete webhook honors that annotation only
while the machine’s Cluster (its cluster.x-k8s.io/cluster-name label) has
spec.paused: true, which clusterctl move sets before it deletes
anything. On an unpaused Cluster the annotation changes nothing: deleting a
TerraformMachine whose Machine is live is refused, because it would skip
drain and the lifecycle hooks. This check exists only on TerraformMachine:
a TerraformMachinePool is never delete-guarded, moved or not.
Confirm it worked
kubectl --kubeconfig <target-kubeconfig> get terraformclusters,terraformmachines,terraformmachinepools -n <ns>
kubectl --kubeconfig <target-kubeconfig> get cluster -n <ns> <name> -o jsonpath='{.spec.paused}'
The first command lists the Terraform* objects on the target with the
same names they had on the source; the second prints nothing (or false),
confirming clusterctl move unpaused the Cluster once it finished. Expect
Ready=False until you complete the manual steps
above and any
variablesFrom source the objects need, then a Job starts on the target
and the object reaches the same Ready state it had on the source. See the
other runbooks for any condition that does not clear on its
own.
Orphaned objects left on the source
After a successful move, the source namespace still holds the runner
ServiceAccount and RoleBinding (captf-runner), and any run or cluster
write lease whose Job was active during the move or had not been bookkept
yet (captf-run-<suffix>, captf-cluster-<hash>) — all labeled
captf.io/managed=true. Jobs and their per-run Secrets are not left
behind: they have owner references to the deleted source objects, so the
garbage collector removes them.
The manager sweeps these automatically: on start, and then every
--sync-period (default 10 minutes), it removes the captf.io/managed=true
ServiceAccounts, RoleBindings and Leases of every namespace that holds no
TerraformCluster, TerraformMachine or TerraformMachinePool — exactly
what a completed move leaves behind. See RBAC for the full
mechanism, including how it treats a namespace an administrator still
manages by hand.
To clean up without waiting for the next sync, once the namespace holds no
Terraform* object:
kubectl delete lease,serviceaccount,rolebinding -l captf.io/managed=true -n <ns>
Run it only once the namespace holds no TerraformCluster,
TerraformMachine or TerraformMachinePool: while any remain, their
runner and lock are still in use.
See also
- The Reconcile Lifecycle — the
block-moveannotation and run leases in full. - Identities and Credentials — creating and rotating the credentials Secret this page tells you to copy.
- RBAC — the runner ServiceAccount and RoleBinding the sweep manages.
- Module Variables —
variablesFromsources and why they don’t move either.
API Reference
CAPTF’s CRDs are one group/version, infrastructure.cluster.x-k8s.io/v1alpha1,
with seven kinds:
- TerraformCluster — the InfraCluster of Cluster API, provisioned by a Terraform/OpenTofu module.
- TerraformClusterTemplate — a template for TerraformClusters, used by ClusterClass.
- TerraformMachine — the InfraMachine of Cluster API, provisioned by a Terraform/OpenTofu module.
- TerraformMachineTemplate — a template for TerraformMachines, used by MachineDeployments, MachineSets, KubeadmControlPlane and ClusterClass.
- TerraformMachinePool — the InfraMachinePool of Cluster API, provisioned by a Terraform/OpenTofu module.
- TerraformMachinePoolTemplate — a template for TerraformMachinePools, used by MachinePools.
- TerraformClusterIdentity — cluster-scoped cloud credentials for TerraformClusters, TerraformMachines and TerraformMachinePools in allowed namespaces.
Every kind’s spec.variables and spec.variablesFrom, and the module
contract they feed, are documented in Job Inputs.
infrastructure.cluster.x-k8s.io/v1alpha1
Package v1alpha1 contains the API Schema definitions for CAPTF (the cluster-api-provider-terraform), the infrastructure.cluster.x-k8s.io v1alpha1 group version. CAPTF turns a Cluster API cluster, machine or machine pool into one Terraform or OpenTofu module run: TerraformCluster, TerraformMachine and TerraformMachinePool carry the desired state that CAPI’s core controllers create and drive, TerraformClusterTemplate, TerraformMachineTemplate and TerraformMachinePoolTemplate each hold a template (their spec.template) from which a TerraformCluster, TerraformMachine or TerraformMachinePool is created — the first typically referenced by a ClusterClass, the second by a MachineDeployment, MachineSet or a control-plane provider, the third by a MachinePool — and TerraformClusterIdentity holds the cloud credentials a cluster’s module run is allowed to use, mirrored into the namespaces that reference it. Seven kinds in all, each registered, with its List type, in groupversion_info.go.
Every object’s spec.source names the one OCI image that bundles the module’s Terraform/OpenTofu code and the runtime binary that runs it (the image-contract types, Source and JobPolicy, live in common_types.go), and spec.identityRef, spec.variables and spec.variablesFrom feed that module’s inputs. The controllers in internal/controllers render those inputs, run the image as a Kubernetes Job and translate its outputs and health into the object’s status: status.conditions (see conditions_consts.go for every condition type and reason CAPTF sets, and their Ready-summarization rules) and the run-tracking status types common_types.go shares across the TerraformCluster, TerraformMachine and TerraformMachinePool kinds — ActiveJob, LastRun (with its RunStep and RunError), SourceStatus and StateBackup — alongside the drift and remediation policy types DriftPolicy (cluster), MachineDriftPolicy (also reused for a cluster’s per-machine and per-pool defaults), MachinePoolDriftPolicy (pool, never fully disabled) and MachineRemediation (machine).
zz_generated.deepcopy.go is controller-gen output; regenerate it with make generate, never hand-edit it. groupversion_info.go registers the group
version and every kind with the runtime scheme.
Resource Types
- TerraformCluster
- TerraformClusterIdentity
- TerraformClusterTemplate
- TerraformMachine
- TerraformMachinePool
- TerraformMachinePoolTemplate
- TerraformMachineTemplate
ActiveJob
ActiveJob identifies the Job currently running for an object.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
name string | name of the Job. | Yes | MaxLength: 63 MinLength: 1 | |
operation Operation | operation the Job runs. | Yes | Enum: [apply destroy drift refresh restore plan] | |
attempt integer | attempt is the operation’s Job sequence number, the a<N> in the Job name, starting at 1. It counts every Job of the operation still retained, not retries: the 40th refresh is attempt 40. | Yes | Minimum: 1 | |
startTime Time | startTime of the Job. | No |
AllowedNamespaces
AllowedNamespaces selects namespaces allowed to use an identity. At least one of list and selector must be set.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
list string array | list of namespace names. | No | MaxItems: 100 MinItems: 1 items:MaxLength: 63 items:MinLength: 1 items:Pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ | |
selector LabelSelector | selector matches namespace labels. An empty selector ({}) matches every namespace. | No |
ApplyPolicy
Underlying type: string
ApplyPolicy decides whether a TerraformCluster applies a change on its own or waits until its plan is approved.
Validation:
- Enum: [Automatic Manual]
Appears in:
| Field | Description |
|---|---|
Automatic | ApplyPolicyAutomatic applies every change as soon as it is seen, only guarded against destructive plans. |
Manual | ApplyPolicyManual plans every change first (a plan Job), shows the plan in status.plan and applies it only once ApprovePlanAnnotation names its hash. The first apply of a new cluster is not gated. |
Architecture
Underlying type: string
Architecture is a node CPU architecture, as reported for scale from zero.
Validation:
- Enum: [amd64 arm64 s390x ppc64le]
Appears in:
| Field | Description |
|---|---|
amd64 | ArchitectureAmd64 is amd64. |
arm64 | ArchitectureArm64 is arm64. |
s390x | ArchitectureS390x is s390x. |
ppc64le | ArchitecturePpc64le is ppc64le. |
CapacitySource
CapacitySource records which image the capacity was resolved from.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
image string | image is the spec image reference last resolved. | Yes | MaxLength: 512 MinLength: 1 |
DriftAction
Underlying type: string
DriftAction is what the controller does when a drift check finds changes.
Validation:
- Enum: [Report Remediate]
Appears in:
| Field | Description |
|---|---|
Report | DriftActionReport records drift in the DriftDetected condition only. |
Remediate | DriftActionRemediate applies the current inputs to remove the drift. |
DriftPolicy
DriftPolicy configures periodic drift detection of a TerraformCluster.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
intervalSeconds integer | intervalSeconds between drift checks, in seconds. Defaults to the manager’s –drift-default-interval (30m), applied at reconcile. 0 disables drift checks, and with them every health sample after provisioning: health is re-read only by a refresh or drift run. | No | Minimum: 0 | |
action DriftAction | action taken when drift is found: Report or Remediate. Defaults to Report, applied at reconcile. Remediate applies the current inputs automatically, reverting every out-of-band change. | No | Enum: [Report Remediate] |
DriftSummary
DriftSummary summarizes the plan of a drift run that found changes.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
add integer | add is the number of resources the plan would create. | No | Minimum: 0 | |
change integer | change is the number of resources the plan would update in place. | No | Minimum: 0 | |
destroy integer | destroy is the number of resources the plan would destroy, counting replacements. | No | Minimum: 0 | |
resources string array | resources are the addresses of the drifted resources, at most 20. | No | MaxItems: 20 MinItems: 1 items:MaxLength: 512 items:MinLength: 1 |
HealthState
Underlying type: string
HealthState mirrors the health.state enum of the module contract (internal/contract.HealthState, https://captf.io/docs/module-author/contract/v1alpha1/common.html), for MachinePoolInstance.State.
Validation:
- Enum: [pending running degraded stopped terminated unknown]
Appears in:
| Field | Description |
|---|---|
pending | HealthStatePending is the contract’s “pending” health state. |
running | HealthStateRunning is the contract’s “running” health state. |
degraded | HealthStateDegraded is the contract’s “degraded” health state. |
stopped | HealthStateStopped is the contract’s “stopped” health state. |
terminated | HealthStateTerminated is the contract’s “terminated” health state. |
unknown | HealthStateUnknown is the contract’s “unknown” health state. |
IdentityReference
IdentityReference names a cluster-scoped TerraformClusterIdentity.
Appears in:
- TerraformClusterDefaults
- TerraformClusterSpec
- TerraformMachinePoolSpec
- TerraformMachineSpec
- WorkspaceSpec
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
name string | name of the TerraformClusterIdentity. | Yes | MaxLength: 253 MinLength: 1 |
Initialization
Initialization holds the v1beta2 contract’s initialization status.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
provisioned boolean | provisioned is true once the infrastructure is provisioned: derived from state until it first holds, then latched for the object’s life. | No |
JobPolicy
JobPolicy tunes the Kubernetes Jobs that run the module. Every field is optional. On a TerraformMachine or TerraformMachinePool the policy is merged field by field with TerraformCluster.spec.defaults.jobs: a field the machine or pool sets wins, an unset one comes from the defaults, and a field neither sets gets the controller’s built-in default. env is merged by name (the machine’s or pool’s wins on the same name) and imagePullSecrets is the union (the machine’s or pool’s first); resources, securityContext and podSecurityContext are replaced as a whole. Defaults are resolved at reconcile time and never persisted, so a provider upgrade reaches existing objects. Jobs never retry pods (backoffLimit 0) and never get a TTL: the controller owns retries and prunes finished Jobs itself.
Appears in:
- TerraformClusterDefaults
- TerraformClusterSpec
- TerraformMachinePoolSpec
- TerraformMachineSpec
- WorkspaceSpec
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
successfulJobsHistoryLimit integer | successfulJobsHistoryLimit is how many succeeded Jobs to keep per object and operation. The newest succeeded Job of each operation is kept even at 0. Defaults to 3, applied at reconcile. | No | Maximum: 100 Minimum: 0 | |
failedJobsHistoryLimit integer | failedJobsHistoryLimit is how many failed Jobs to keep per object and operation. The newest failed Job of an operation is kept even at 0 while no newer Job of that operation succeeded: retry backoff counts it. Defaults to 3, applied at reconcile. | No | Maximum: 100 Minimum: 0 | |
activeDeadlineSeconds integer | activeDeadlineSeconds bounds a Job’s run time, in seconds, at most one day. Defaults to 3600, applied at reconcile when unset (0). When a policy sets both, lockTimeoutSeconds must be less than activeDeadlineSeconds. | No | Maximum: 86400 Minimum: 1 | |
serviceAccountName string | serviceAccountName overrides the runner ServiceAccount. When unset the controller creates captf-runner, bound to the static captf-runner ClusterRole. An override ServiceAccount must exist and carry the label captf.io/runner=true, or no Job is created. | No | MaxLength: 253 MinLength: 1 | |
lockTimeoutSeconds integer | lockTimeoutSeconds is passed to the runtime as -lock-timeout, in seconds. Defaults to 300, applied at reconcile. | No | Maximum: 3600 Minimum: 0 | |
imagePullSecrets LocalObjectReference array | imagePullSecrets for the Job pod: they cover the source image and the runner init image. | No | MaxItems: 10 MinItems: 1 | |
resources ResourceRequirements | resources of the main container. | No | ||
env EnvVar array | env adds environment variables to the main container. It cannot override the TF_* and KUBE_* variables the runner sets. | No | MaxItems: 64 MinItems: 1 | |
securityContext SecurityContext | securityContext of the main container. Defaults, applied when the Job is built: seccompProfile RuntimeDefault, capabilities drop ALL, allowPrivilegeEscalation false, readOnlyRootFilesystem true. runAsNonRoot is not defaulted. The webhook rejects privileged: true, allowPrivilegeEscalation: true and any capabilities.add: the container holds cloud credentials. | No | ||
podSecurityContext PodSecurityContext | podSecurityContext of the Job pod. Defaults, applied when the Job is built: seccompProfile RuntimeDefault. | No |
LastRun
LastRun is the result of the most recent completed Job, copied from the runner’s termination message.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
job string | job is the name of the Job. | Yes | MaxLength: 63 MinLength: 1 | |
operation Operation | operation the Job ran. | Yes | Enum: [apply destroy drift refresh restore plan] | |
steps RunStep array | steps the runner executed, in order. | No | MaxItems: 16 MinItems: 1 | |
error RunError | error is set when the run failed. | No | ||
drift DriftSummary | drift is set when a drift run found changes. | No | MinProperties: 1 |
MachineDriftPolicy
MachineDriftPolicy configures periodic drift detection of a TerraformMachine. Drift on a machine is always reported, never remediated: the machine is immutable infrastructure, replaced by a rollout.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
intervalSeconds integer | intervalSeconds between drift checks, in seconds. Defaults to the manager’s –drift-default-interval (30m), applied at reconcile. 0 disables drift checks. Unless remediation.annotateMachine is true (which refreshes at remediation.healthCheckIntervalSeconds), that also stops every health sample after provisioning. | No | Minimum: 0 |
MachinePoolDriftPolicy
MachinePoolDriftPolicy configures periodic drift detection of a TerraformMachinePool. Unlike MachineDriftPolicy, 0 is rejected by the CRD schema: for a pool it is membership refresh (TerraformMachinePoolSpec.MembershipRefreshIntervalSeconds), not drift, that keeps status fresh (https://captf.io/docs/module-author/contract/v1alpha1/machinepool.html “Membership refresh”), and the drift Job itself feeds the refreshed replicas into its plan (machinepool.md “Drift order”), so disabling drift would also stop that refresh from ever reaching a plan.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
intervalSeconds integer | intervalSeconds between drift checks, in seconds. Defaults to the manager’s –drift-default-interval (30m), applied at reconcile when unset (0). Unlike a machine’s or the cluster’s, a pool’s drift cannot be disabled, so 0 always means “use the default”, never “disabled”. | No | Minimum: 1 | |
action DriftAction | action taken when drift is found: Report or Remediate. Defaults to Report, applied at reconcile. Unlike a machine’s, a pool’s drift may be remediated: the group’s instances are not immutable infrastructure. | No | Enum: [Report Remediate] |
MachinePoolInstance
MachinePoolInstance is one entry of a TerraformMachinePool’s status.instances, mapped from the module’s instances output (https://captf.io/docs/module-author/contract/v1alpha1/machinepool.html “instances”).
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
providerID string | providerID of the instance. | Yes | MaxLength: 512 MinLength: 1 | |
instanceID string | instanceID is a provider-defined identifier, distinct from providerID when the module has one to give. | No | MaxLength: 256 MinLength: 1 | |
addresses MachineAddress array | addresses of the instance. | No | MaxItems: 256 MinItems: 1 | |
failureDomain string | failureDomain the instance actually runs in. | No | MaxLength: 256 MinLength: 1 | |
state HealthState | state of the instance, the health.state enum. | No | Enum: [pending running degraded stopped terminated unknown] |
MachineRemediation
MachineRemediation configures how a TerraformMachine signals an unhealthy instance to Cluster API beyond its Ready condition.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
annotateMachine boolean | annotateMachine sets cluster.x-k8s.io/remediate-machine on the owner Machine once the instance has been unhealthy for unhealthyThreshold consecutive samples, or at once when it is terminated. CAPTF removes the annotation it set once the instance reads Healthy again and the Machine is not being deleted; an annotation set by anyone else is left alone. It has an effect only when a MachineHealthCheck selects the Machine; a single-replica control plane refuses the remediation. Defaults to false. | No | ||
unhealthyThreshold integer | unhealthyThreshold is the number of consecutive unhealthy health samples before the Machine is annotated. A sample is one completed refresh or drift Job. Defaults to 3, applied at reconcile when unset (0). A terminated instance counts on the first sample. | No | Maximum: 100 Minimum: 1 | |
healthCheckIntervalSeconds integer | healthCheckIntervalSeconds is how often a provisioned machine is refreshed to sample its health while annotateMachine is true, independent of drift.intervalSeconds. Defaults to 300, applied at reconcile when unset (0). With annotateMachine false it is ignored and health is re-read only at the drift (or refresh) cadence, so drift.intervalSeconds 0 then stops health sampling after provisioning. | No | Maximum: 86400 Minimum: 60 |
NodeInfo
NodeInfo describes the nodes the template creates, for Cluster Autoscaler scale from zero. It comes from the image label io.captf.node-info.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
architecture Architecture | architecture of the node’s CPU. | No | Enum: [amd64 arm64 s390x ppc64le] | |
operatingSystem string | operatingSystem of the node, e.g. linux. | No | MaxLength: 64 MinLength: 1 |
Operation
Underlying type: string
Operation is one of the operations a Job runs.
Validation:
- Enum: [apply destroy drift refresh restore plan]
Appears in:
| Field | Description |
|---|---|
apply | OperationApply creates or updates the infrastructure. |
destroy | OperationDestroy destroys the infrastructure. |
drift | OperationDrift refreshes state and plans to detect drift. |
refresh | OperationRefresh refreshes state and outputs only. |
restore | OperationRestore pushes a state backup back into the backend (RestoreStateAnnotation). |
plan | OperationPlan plans a TerraformCluster’s change for review and applies nothing (applyPolicy Manual). |
PlanPreview
PlanPreview summarizes a plan for review: counts and the address and action of each changed resource, never a value.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
inputsHash string | inputsHash is the hash of the inputs the plan was made for. | Yes | MaxLength: 128 MinLength: 1 | |
job string | job is the Job that made the plan: a plan Job, or an approved apply that found the plan changed. | Yes | MaxLength: 63 MinLength: 1 | |
planHash string | planHash fingerprints the plan’s changes: the value of the captf.io/approve-plan annotation that approves it. | Yes | MaxLength: 128 MinLength: 1 | |
add integer | add is the number of resources the plan creates. | No | Minimum: 0 | |
change integer | change is the number of resources the plan updates in place. | No | Minimum: 0 | |
destroy integer | destroy is the number of resources the plan destroys, counting replacements. | No | Minimum: 0 | |
resources string array | resources are “<address> (<action>)” of the changed resources, sorted by address, at most 50; action is create, update, delete, replace, read or forget. | No | MaxItems: 50 MinItems: 1 items:MaxLength: 600 items:MinLength: 1 | |
truncated boolean | truncated is true when resources lists fewer resources than the plan changes. | No | ||
createdAt Time | createdAt is when the plan was made. | No |
RunError
RunError describes why a run failed.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
kind RunErrorKind | kind of failure. | Yes | Enum: [image-layout step interrupted blocked plan-changed] | |
step string | step that failed, for kind step. | No | MaxLength: 64 MinLength: 1 | |
summary string | summary is the runner’s short description of the failure, at most 512 bytes. It is not raw stderr: status is readable by everyone who can get the object, so the full output stays in the Job’s logs. | No | MaxLength: 512 MinLength: 1 |
RunErrorKind
Underlying type: string
RunErrorKind classifies a failed run.
Validation:
- Enum: [image-layout step interrupted blocked plan-changed]
Appears in:
| Field | Description |
|---|---|
image-layout | RunErrorKindImageLayout means the image does not follow the image contract. |
step | RunErrorKindStep means a runtime step failed. |
interrupted | RunErrorKindInterrupted means the step was stopped from outside (the pod got SIGTERM: a drain, an eviction, a Job deletion or its deadline), not that the module failed. |
blocked | RunErrorKindBlocked means a TerraformCluster apply stopped before a plan that deletes or replaces resources, because the captf.io/approve-destructive-plan annotation does not name the inputs hash it renders. Nothing was changed. |
plan-changed | RunErrorKindPlanChanged means a TerraformCluster apply approved for one plan (applyPolicy Manual, captf.io/approve-plan) planned other changes and stopped before applying them. Nothing was changed; the new plan waits for its own approval in status.plan. |
RunStep
RunStep is one runtime command the runner executed.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
name string | name of the step, e.g. init, validate, plan, apply, apply-refresh-only. | Yes | MaxLength: 64 MinLength: 1 | |
exitCode integer | exitCode of the step. | Yes | ||
durationMilliseconds integer | durationMilliseconds is the step’s wall time, in milliseconds. | No | Minimum: 0 |
SecretReference
SecretReference names a Secret in a given namespace.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
name string | name of the Secret. | Yes | MaxLength: 253 MinLength: 1 Pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*$ | |
namespace string | namespace of the Secret. | Yes | MaxLength: 63 MinLength: 1 Pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ |
Source
Source is the deliverable: one OCI image that bundles the role module’s Terraform/OpenTofu code and the runtime binary. There is no separate module source and no separate runtime image. The image layout is a fixed-path contract: /captf/module, /captf/runtime and an optional /captf/providers mirror. The runner always execs /captf/runtime; pull secrets for the image are jobs.imagePullSecrets.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
image string | image is the OCI image reference, registry/repo:tag or registry/repo@sha256:digest. The tag or digest is the module version. Referencing an image grants its publisher Secret-read and cloud-credential access in this namespace. | Yes | MaxLength: 512 MinLength: 1 | |
imagePullPolicy PullPolicy | imagePullPolicy for the image. Defaults to IfNotPresent, applied when the Job is built; Always is recommended for mutable tags. | No | Enum: [IfNotPresent Always Never] |
SourceStatus
SourceStatus records what the last Job actually ran.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
image string | image is the reference that was run last, as given in the spec. | No | MaxLength: 512 MinLength: 1 | |
imageDigest string | imageDigest is the digest the container runtime resolved the image to (pod status imageID). Informational: the pinned copy lives on the durable inputs Secret as captf.io/image-digest. | No | MaxLength: 512 MinLength: 1 | |
runtimeVersion string | runtimeVersion reported by <command> version -json. | No | MaxLength: 64 MinLength: 1 |
StateBackup
StateBackup is one versioned copy of the object’s Terraform state that the controller keeps in a captf-state-backup-* Secret.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
serial integer | serial is the state serial the backup holds; set it as the captf.io/restore-state annotation to restore it. | Yes | Minimum: 1 | |
takenAt Time | takenAt is when the controller copied the state. | Yes | ||
bytes integer | bytes is the compressed size of the backup summed over its Secrets. | Yes | Minimum: 1 |
TemplateMeta
TemplateMeta holds the metadata field every *TemplateResource copies onto the object it creates. TerraformClusterTemplateResource and TerraformMachineTemplateResource embed it.
Appears in:
- TerraformClusterTemplateResource
- TerraformMachinePoolTemplateResource
- TerraformMachineTemplateResource
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No |
TerraformCluster
TerraformCluster is the Schema for the terraformclusters API: the InfraCluster of Cluster API, provisioned by a Terraform/OpenTofu module.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformCluster | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformClusterSpec | spec is the desired state of the TerraformCluster. | Yes | MinProperties: 1 | |
status TerraformClusterStatus | status is the observed state of the TerraformCluster. | No | MinProperties: 1 |
TerraformClusterDefaults
TerraformClusterDefaults are values the TerraformMachines and TerraformMachinePools of a cluster inherit when they do not set them. There is no source: every role names its own image.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
identityRef IdentityReference | identityRef is used by machines and pools without their own identityRef. When unset, such machines and pools use spec.identityRef. | No | ||
jobs JobPolicy | jobs is merged field by field under each machine’s or pool’s jobs policy (see JobPolicy). | No | ||
drift MachineDriftPolicy | drift is merged field by field under each machine’s or pool’s drift policy. A pool’s drift is never fully disabled: an inherited intervalSeconds of 0 disables a machine’s drift checks but not a pool’s, which then uses the controller’s default interval. | No |
TerraformClusterIdentity
TerraformClusterIdentity is the Schema for the terraformclusteridentities API: cluster-scoped cloud credentials for TerraformClusters, TerraformMachines and TerraformMachinePools in allowed namespaces.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformClusterIdentity | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformClusterIdentitySpec | spec is the desired state of the TerraformClusterIdentity. | Yes | MinProperties: 1 | |
status TerraformClusterIdentityStatus | status is the observed state of the TerraformClusterIdentity. | No | MinProperties: 1 |
TerraformClusterIdentitySpec
TerraformClusterIdentitySpec is the desired state of a TerraformClusterIdentity: cloud credentials and who may use them.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
secretRef SecretReference | secretRef names the Secret holding the credentials. It is mirrored into each allowed namespace that uses this identity and delivered to Jobs as environment variables and files. | Yes | ||
allowedNamespaces AllowedNamespaces | allowedNamespaces restricts which namespaces may reference this identity. Unset allows no namespace; selector: \{\} allows everynamespace; list and selector are ORed. An empty object is rejected. | No |
TerraformClusterIdentityStatus
TerraformClusterIdentityStatus is the observed state of a TerraformClusterIdentity. The manager fills it: whether the credentials Secret exists, and where it is mirrored.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
conditions Condition array | conditions of the TerraformClusterIdentity. Ready is True when the credentials Secret exists (SecretFound), False when it does not (SecretNotFound). | No | MaxItems: 32 | |
namespaces string array | namespaces where a mirror of the credentials Secret currently exists. | No | MaxItems: 1000 items:MaxLength: 63 items:MinLength: 1 |
TerraformClusterSpec
TerraformClusterSpec is the desired state of a TerraformCluster: the cluster-role module image and how to run it.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
controlPlaneEndpoint APIEndpoint | controlPlaneEndpoint is the endpoint of the cluster’s API server. A value set by the user is passed to the module as its control_plane_endpoint input; otherwise the controller writes the module’s output here once. Once it has a host it is immutable (the webhook enforces this). | No | ||
source Source | source is the role image: module code and runtime. | Yes | ||
identityRef IdentityReference | identityRef names the TerraformClusterIdentity whose credentials this object’s Jobs use. Whether it is required, and where it falls back to when unset, depends on the kind. | No | ||
jobs JobPolicy | jobs tunes the Jobs that run this object’s module. On a kind that inherits defaults it is merged field by field over them (see JobPolicy). | No | ||
variables RawExtension | variables are module variables, a JSON object: each key becomes a named argument of the role module, converted by the module’s declared type. Keys are Terraform identifiers; captf_ names and the role’s contract inputs are reserved. Inline variables win over variablesFrom. Whether a change re-applies or is rejected as immutable depends on the kind. A key the module does not declare fails the apply (“Unsupported argument”). | No | MaxProperties: 256 MinProperties: 1 Type: object | |
variablesFrom VariablesSource array | variablesFrom reads module variables from ConfigMaps and Secrets in this namespace labeled captf.io/variables=true, in list order: a later source wins on the same key, and inline variables win over all of them. Whether and when a change to a referenced source takes effect depends on the kind. | No | ExactlyOneOf: [configMapRef secretRef] MaxItems: 16 MinItems: 1 | |
drift DriftPolicy | drift configures drift detection for this cluster. | No | ||
applyPolicy ApplyPolicy | applyPolicy decides when a change is applied. Automatic (the default, applied at reconcile) applies every change of the inputs, and a drift remediation, as soon as it is seen; only a plan that deletes or replaces resources waits for captf.io/approve-destructive-plan. Manual runs a plan Job first, reports the plan in status.plan and waits until the captf.io/approve-plan annotation names its hash; the apply then runs only if it plans the same changes again. The first apply of a new cluster (no state yet) is never gated. Mutable. | No | Enum: [Automatic Manual] | |
defaults TerraformClusterDefaults | defaults are inherited by the TerraformMachines and TerraformMachinePools of this cluster, field by field: a field a machine or pool sets wins, an unset one comes from here. They do not apply to the TerraformCluster itself. | No |
TerraformClusterStatus
TerraformClusterStatus is the observed state of a TerraformCluster. Nothing here is load-bearing: every value is rebuilt from spec, the state Secret, the durable inputs Secret or the Job list, because clusterctl move does not restore status.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
conditions Condition array | conditions of the TerraformCluster. Ready is mirrored by Cluster API into the Cluster’s InfrastructureReady condition. | No | MaxItems: 32 | |
initialization Initialization | initialization is the v1beta2 contract’s initialization status. | No | MinProperties: 1 | |
observedGeneration integer | observedGeneration is the generation this status was computed for. | No | Minimum: 1 | |
activeJob ActiveJob | activeJob is the Job currently running for this object, if any. | No | ||
lastRun LastRun | lastRun is the result of the most recent completed Job. | No | ||
lastDriftCheck Time | lastDriftCheck is when the last drift check completed. | No | ||
lastRefresh Time | lastRefresh is when the last refresh or drift check completed, or, for a kind whose apply itself can give a definite health reading, when that apply finished instead (that reading stands in for the refresh after the apply). | No | ||
pendingRefreshes integer | pendingRefreshes counts the consecutive health samples (completed refresh or drift Jobs) that read pending since the last other reading or the last apply; unset otherwise. It spaces the refreshes while health is pending: 30s, then 1m, 2m, 4m and at most 5m. It lives in status only, so it restarts at 0 (30s) after clusterctl move. | No | Minimum: 1 | |
observedStateSerial integer | observedStateSerial is the Terraform state serial the outputs were read from. | No | Minimum: 1 | |
stateSecretSuffix string | stateSecretSuffix is the kubernetes backend secret_suffix of this object’s state. Informational: the controller derives it deterministically. | No | MaxLength: 63 MinLength: 1 | |
source SourceStatus | source records what the last Job actually ran. | No | MinProperties: 1 | |
stateBackups StateBackup array | stateBackups are the state backups the controller keeps (newest first), as of the last backup, prune or restore request. | No | MaxItems: 16 MinItems: 1 | |
failureDomains FailureDomain array | failureDomains reported by the module’s failure_domains output. | No | MaxItems: 100 MinItems: 1 | |
plan PlanPreview | plan is the plan of the change waiting for approval under applyPolicy Manual; empty when none waits. Approve it by setting the captf.io/approve-plan annotation to plan.planHash. | No |
TerraformClusterTemplate
TerraformClusterTemplate is the Schema for the terraformclustertemplates API: a template for TerraformClusters, used by ClusterClass.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformClusterTemplate | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformClusterTemplateSpec | spec is the desired state of the TerraformClusterTemplate. | Yes |
TerraformClusterTemplateResource
TerraformClusterTemplateResource describes the TerraformCluster created from a template.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformClusterSpec | spec of the TerraformCluster created from this template. | Yes | MinProperties: 1 |
TerraformClusterTemplateSpec
TerraformClusterTemplateSpec is the desired state of a TerraformClusterTemplate.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
template TerraformClusterTemplateResource | template is the TerraformCluster created from this template. | Yes | MinProperties: 1 |
TerraformMachine
TerraformMachine is the Schema for the terraformmachines API: the InfraMachine of Cluster API, provisioned by a Terraform/OpenTofu module.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformMachine | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformMachineSpec | spec is the desired state of the TerraformMachine. | Yes | MinProperties: 1 | |
status TerraformMachineStatus | status is the observed state of the TerraformMachine. | No | MinProperties: 1 |
TerraformMachinePool
TerraformMachinePool is the Schema for the terraformmachinepools API: the InfraMachinePool of Cluster API, provisioned by a Terraform/OpenTofu module.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformMachinePool | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformMachinePoolSpec | spec is the desired state of the TerraformMachinePool. | Yes | MinProperties: 1 | |
status TerraformMachinePoolStatus | status is the observed state of the TerraformMachinePool. | No | MinProperties: 1 |
TerraformMachinePoolSpec
TerraformMachinePoolSpec is the desired state of a TerraformMachinePool: the machinepool-role module image and how to run it. Unlike a TerraformMachine, every field here is mutable: the pool is re-applied on a spec change, a replica change or the bootstrap Secret’s rotation (https://captf.io/docs/module-author/contract/v1alpha1/machinepool.html “Lifecycle”).
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
providerID string | providerID is the scaling group’s provider ID, set by the controller from the module’s provider_id output. Optional in the InfraMachinePool contract; may stay unset for group-less implementations. | No | MaxLength: 512 MinLength: 1 | |
providerIDList string array | providerIDList are the provider IDs of every non-terminated member of the group, set by the controller from the module’s provider_id_list output. Each entry must equal the corresponding Node’s spec.providerID. | No | MaxItems: 10000 items:MaxLength: 512 items:MinLength: 1 | |
source Source | source is the role image: module code and runtime. | Yes | ||
identityRef IdentityReference | identityRef names the TerraformClusterIdentity whose credentials this object’s Jobs use. Whether it is required, and where it falls back to when unset, depends on the kind. | No | ||
jobs JobPolicy | jobs tunes the Jobs that run this object’s module. On a kind that inherits defaults it is merged field by field over them (see JobPolicy). | No | ||
variables RawExtension | variables are module variables, a JSON object: each key becomes a named argument of the role module, converted by the module’s declared type. Keys are Terraform identifiers; captf_ names and the role’s contract inputs are reserved. Inline variables win over variablesFrom. Whether a change re-applies or is rejected as immutable depends on the kind. A key the module does not declare fails the apply (“Unsupported argument”). | No | MaxProperties: 256 MinProperties: 1 Type: object | |
variablesFrom VariablesSource array | variablesFrom reads module variables from ConfigMaps and Secrets in this namespace labeled captf.io/variables=true, in list order: a later source wins on the same key, and inline variables win over all of them. Whether and when a change to a referenced source takes effect depends on the kind. | No | ExactlyOneOf: [configMapRef secretRef] MaxItems: 16 MinItems: 1 | |
drift MachinePoolDriftPolicy | drift is merged field by field over the cluster’s defaults.drift. Unlike a machine’s, a pool’s drift may be remediated. | No | ||
membershipRefreshIntervalSeconds integer | membershipRefreshIntervalSeconds is how often the controller runsapply -refresh-only to pick up group membership changes (new ordeparted instances) between applies, in seconds (https://captf.io/docs/module-author/contract/v1alpha1/machinepool.html “Membership refresh”). 0 (unset) means 60, applied at reconcile; the CRD schema’s minimum of 15 makes 0 itself an invalid setting, so it unambiguously means unset, the same convention as activeDeadlineSeconds and unhealthyThreshold (kube-api-linter optionalfields: WhenRequired). | No | Maximum: 86400 Minimum: 15 |
TerraformMachinePoolStatus
TerraformMachinePoolStatus is the observed state of a TerraformMachinePool. Nothing here is load-bearing: every value is rebuilt after clusterctl move.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
conditions Condition array | conditions of the TerraformMachinePool. Ready is mirrored by Cluster API into the MachinePool’s InfrastructureReady condition. | No | MaxItems: 32 | |
initialization Initialization | initialization is the v1beta2 contract’s initialization status. | No | MinProperties: 1 | |
observedGeneration integer | observedGeneration is the generation this status was computed for. | No | Minimum: 1 | |
activeJob ActiveJob | activeJob is the Job currently running for this object, if any. | No | ||
lastRun LastRun | lastRun is the result of the most recent completed Job. | No | ||
lastDriftCheck Time | lastDriftCheck is when the last drift check completed. | No | ||
lastRefresh Time | lastRefresh is when the last refresh or drift check completed, or, for a kind whose apply itself can give a definite health reading, when that apply finished instead (that reading stands in for the refresh after the apply). | No | ||
pendingRefreshes integer | pendingRefreshes counts the consecutive health samples (completed refresh or drift Jobs) that read pending since the last other reading or the last apply; unset otherwise. It spaces the refreshes while health is pending: 30s, then 1m, 2m, 4m and at most 5m. It lives in status only, so it restarts at 0 (30s) after clusterctl move. | No | Minimum: 1 | |
observedStateSerial integer | observedStateSerial is the Terraform state serial the outputs were read from. | No | Minimum: 1 | |
stateSecretSuffix string | stateSecretSuffix is the kubernetes backend secret_suffix of this object’s state. Informational: the controller derives it deterministically. | No | MaxLength: 63 MinLength: 1 | |
source SourceStatus | source records what the last Job actually ran. | No | MinProperties: 1 | |
stateBackups StateBackup array | stateBackups are the state backups the controller keeps (newest first), as of the last backup, prune or restore request. | No | MaxItems: 16 MinItems: 1 | |
ready boolean | ready is the v1beta1 compatibility field Cluster API v1.14 still reads to decide the pool is provisioned (external.IsReady, capi/core/reconcilers/machinepool/machinepool_controller_phases.go). It is latched together with initialization.provisioned: once true, it stays true for the object’s life. | No | ||
replicas integer | replicas is the group’s desired capacity as observed at the last refresh, from the module’s replicas output. Outside a scaling transition it equals len(providerIDList). | No | Minimum: 0 | |
instances MachinePoolInstance array | instances are the group’s members, from the module’s instances output. Provider-defined shape; not used by core Cluster API. | No | MaxItems: 1000 MinItems: 1 MinProperties: 1 |
TerraformMachinePoolTemplate
TerraformMachinePoolTemplate is the Schema for the terraformmachinepooltemplates API: a template for TerraformMachinePools, used by MachinePools. Unlike TerraformMachineTemplate it has no status: pools have no scale-from-zero, so there is no capacity or nodeInfo to resolve.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformMachinePoolTemplate | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformMachinePoolTemplateSpec | spec is the desired state of the TerraformMachinePoolTemplate. | Yes |
TerraformMachinePoolTemplateResource
TerraformMachinePoolTemplateResource describes the TerraformMachinePool created from a template.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformMachinePoolSpec | spec of the TerraformMachinePool created from this template. | Yes | MinProperties: 1 |
TerraformMachinePoolTemplateSpec
TerraformMachinePoolTemplateSpec is the desired state of a TerraformMachinePoolTemplate.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
template TerraformMachinePoolTemplateResource | template is the TerraformMachinePool created from this template. | Yes | MinProperties: 1 |
TerraformMachineSpec
TerraformMachineSpec is the desired state of a TerraformMachine: the machine-role module image and how to run it. source, identityRef, variables and variablesFrom define the machine and are immutable after creation, and providerID can only be set once, by the controller; the admission webhook enforces this. jobs, drift and remediation are operational policy and may change at any time.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
providerID string | providerID is the instance’s provider ID, set by the controller from the module’s provider_id output. It must equal the Node’s spec.providerID. | No | MaxLength: 512 MinLength: 1 | |
source Source | source is the role image: module code and runtime. | Yes | ||
identityRef IdentityReference | identityRef names the TerraformClusterIdentity whose credentials this object’s Jobs use. Whether it is required, and where it falls back to when unset, depends on the kind. | No | ||
jobs JobPolicy | jobs tunes the Jobs that run this object’s module. On a kind that inherits defaults it is merged field by field over them (see JobPolicy). | No | ||
variables RawExtension | variables are module variables, a JSON object: each key becomes a named argument of the role module, converted by the module’s declared type. Keys are Terraform identifiers; captf_ names and the role’s contract inputs are reserved. Inline variables win over variablesFrom. Whether a change re-applies or is rejected as immutable depends on the kind. A key the module does not declare fails the apply (“Unsupported argument”). | No | MaxProperties: 256 MinProperties: 1 Type: object | |
variablesFrom VariablesSource array | variablesFrom reads module variables from ConfigMaps and Secrets in this namespace labeled captf.io/variables=true, in list order: a later source wins on the same key, and inline variables win over all of them. Whether and when a change to a referenced source takes effect depends on the kind. | No | ExactlyOneOf: [configMapRef secretRef] MaxItems: 16 MinItems: 1 | |
drift MachineDriftPolicy | drift is merged field by field over the cluster’s defaults.drift. Drift on a machine is always reported, never remediated. | No | ||
remediation MachineRemediation | remediation configures how an unhealthy instance is signalled to Cluster API beyond the Ready condition. | No |
TerraformMachineStatus
TerraformMachineStatus is the observed state of a TerraformMachine. Nothing here is load-bearing: every value is rebuilt after clusterctl move.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
conditions Condition array | conditions of the TerraformMachine. Ready is mirrored by Cluster API into the Machine’s InfrastructureReady condition. | No | MaxItems: 32 | |
initialization Initialization | initialization is the v1beta2 contract’s initialization status. | No | MinProperties: 1 | |
observedGeneration integer | observedGeneration is the generation this status was computed for. | No | Minimum: 1 | |
activeJob ActiveJob | activeJob is the Job currently running for this object, if any. | No | ||
lastRun LastRun | lastRun is the result of the most recent completed Job. | No | ||
lastDriftCheck Time | lastDriftCheck is when the last drift check completed. | No | ||
lastRefresh Time | lastRefresh is when the last refresh or drift check completed, or, for a kind whose apply itself can give a definite health reading, when that apply finished instead (that reading stands in for the refresh after the apply). | No | ||
pendingRefreshes integer | pendingRefreshes counts the consecutive health samples (completed refresh or drift Jobs) that read pending since the last other reading or the last apply; unset otherwise. It spaces the refreshes while health is pending: 30s, then 1m, 2m, 4m and at most 5m. It lives in status only, so it restarts at 0 (30s) after clusterctl move. | No | Minimum: 1 | |
observedStateSerial integer | observedStateSerial is the Terraform state serial the outputs were read from. | No | Minimum: 1 | |
stateSecretSuffix string | stateSecretSuffix is the kubernetes backend secret_suffix of this object’s state. Informational: the controller derives it deterministically. | No | MaxLength: 63 MinLength: 1 | |
source SourceStatus | source records what the last Job actually ran. | No | MinProperties: 1 | |
stateBackups StateBackup array | stateBackups are the state backups the controller keeps (newest first), as of the last backup, prune or restore request. | No | MaxItems: 16 MinItems: 1 | |
addresses MachineAddress array | addresses of the instance, from the module’s addresses output, in the controller’s canonical order. | No | MaxItems: 256 MinItems: 1 | |
failureDomain string | failureDomain the instance actually runs in. | No | MaxLength: 256 MinLength: 1 | |
interruptible boolean | interruptible is true for spot/preemptible instances. Cluster API then labels the Node cluster.x-k8s.io/interruptible. | No | ||
unhealthySamples integer | unhealthySamples counts consecutive unhealthy health samples (one per completed refresh or drift Job after provisioning, or per apply whose own outputs stood in for the post-apply refresh); unset when the instance is healthy. It lives in status only, so it restarts at 0 after clusterctl move. | No | Minimum: 1 |
TerraformMachineTemplate
TerraformMachineTemplate is the Schema for the terraformmachinetemplates API: a template for TerraformMachines, used by MachineDeployments, MachineSets, KubeadmControlPlane and ClusterClass.
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
apiVersion string | infrastructure.cluster.x-k8s.io/v1alpha1 | Yes | ||
kind string | TerraformMachineTemplate | Yes | ||
kind string | Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds | No | ||
apiVersion string | APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources | No | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformMachineTemplateSpec | spec is the desired state of the TerraformMachineTemplate. | Yes | ||
status TerraformMachineTemplateStatus | status is the observed state of the TerraformMachineTemplate. | No | MinProperties: 1 |
TerraformMachineTemplateResource
TerraformMachineTemplateResource describes the TerraformMachine created from a template.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | No | ||
spec TerraformMachineSpec | spec of the TerraformMachine created from this template. | Yes | MinProperties: 1 |
TerraformMachineTemplateSpec
TerraformMachineTemplateSpec is the desired state of a TerraformMachineTemplate.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
template TerraformMachineTemplateResource | template is the TerraformMachine created from this template. | Yes | MinProperties: 1 |
TerraformMachineTemplateStatus
TerraformMachineTemplateStatus is the observed state of a TerraformMachineTemplate: the node size declared by its image, for Cluster Autoscaler scale from zero.
Validation:
- MinProperties: 1
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
conditions Condition array | conditions of the TerraformMachineTemplate (CapacityResolved). | No | MaxItems: 32 | |
nodeInfo NodeInfo | nodeInfo of the nodes the template creates, from the image label io.captf.node-info. | No | MinProperties: 1 | |
capacitySource CapacitySource | capacitySource is the spec image reference capacity and nodeInfo were resolved from; they are re-resolved when the spec image changes. | No |
VariablesFormat
Underlying type: string
VariablesFormat is how the data values of a variablesFrom source are passed to the module.
Validation:
- Enum: [String JSON]
Appears in:
| Field | Description |
|---|---|
String | VariablesFormatString passes each data value as a string. The module’s declared variable type converts it (“3” to a number, “true” to a bool). |
JSON | VariablesFormatJSON parses each data value as JSON, for lists, maps and objects. A value that is not valid JSON is VariablesInvalid. |
VariablesSource
VariablesSource reads module variables from the data of one ConfigMap or Secret in the object’s namespace: every data key becomes a variable of the same name. The source must carry the label captf.io/variables=true. Variables from a Secret are declared sensitive in the generated root, so Terraform redacts them in plan and apply output; they are still stored in the inputs Secrets and in state, like every input.
Validation:
- ExactlyOneOf: [configMapRef secretRef]
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
configMapRef VariablesSourceReference | configMapRef names a ConfigMap. Exactly one of configMapRef and secretRef is set. | No | ||
secretRef VariablesSourceReference | secretRef names a Secret. Exactly one of configMapRef and secretRef is set. | No | ||
optional boolean | optional makes a missing or unlabeled source contribute nothing instead of holding the object at DependenciesReady False (VariablesSourceNotFound). Defaults to false. | No | ||
format VariablesFormat | format of the data values: String passes each value as a string, JSON parses each as JSON. Defaults to String, applied at reconcile. | No | Enum: [String JSON] |
VariablesSourceReference
VariablesSourceReference names a ConfigMap or Secret in the object’s own namespace.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
name string | name of the ConfigMap or Secret. | Yes | MaxLength: 253 MinLength: 1 Pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*$ |
WorkspaceSpec
WorkspaceSpec is the part of a Job-running kind’s spec every such kind shares: the role module image, its identity, Job policy and module variables. TerraformClusterSpec and TerraformMachineSpec embed it.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
source Source | source is the role image: module code and runtime. | Yes | ||
identityRef IdentityReference | identityRef names the TerraformClusterIdentity whose credentials this object’s Jobs use. Whether it is required, and where it falls back to when unset, depends on the kind. | No | ||
jobs JobPolicy | jobs tunes the Jobs that run this object’s module. On a kind that inherits defaults it is merged field by field over them (see JobPolicy). | No | ||
variables RawExtension | variables are module variables, a JSON object: each key becomes a named argument of the role module, converted by the module’s declared type. Keys are Terraform identifiers; captf_ names and the role’s contract inputs are reserved. Inline variables win over variablesFrom. Whether a change re-applies or is rejected as immutable depends on the kind. A key the module does not declare fails the apply (“Unsupported argument”). | No | MaxProperties: 256 MinProperties: 1 Type: object | |
variablesFrom VariablesSource array | variablesFrom reads module variables from ConfigMaps and Secrets in this namespace labeled captf.io/variables=true, in list order: a later source wins on the same key, and inline variables win over all of them. Whether and when a change to a referenced source takes effect depends on the kind. | No | ExactlyOneOf: [configMapRef secretRef] MaxItems: 16 MinItems: 1 |
WorkspaceStatus
WorkspaceStatus is the part of a Job-running kind’s status every such kind shares. Nothing here is load-bearing: every value is rebuilt from spec, the state Secret, the durable inputs Secret or the Job list, because clusterctl move does not restore status. TerraformClusterStatus and TerraformMachineStatus embed it.
Appears in:
| Field | Description | Required | Default | Validation |
|---|---|---|---|---|
initialization Initialization | initialization is the v1beta2 contract’s initialization status. | No | MinProperties: 1 | |
observedGeneration integer | observedGeneration is the generation this status was computed for. | No | Minimum: 1 | |
activeJob ActiveJob | activeJob is the Job currently running for this object, if any. | No | ||
lastRun LastRun | lastRun is the result of the most recent completed Job. | No | ||
lastDriftCheck Time | lastDriftCheck is when the last drift check completed. | No | ||
lastRefresh Time | lastRefresh is when the last refresh or drift check completed, or, for a kind whose apply itself can give a definite health reading, when that apply finished instead (that reading stands in for the refresh after the apply). | No | ||
pendingRefreshes integer | pendingRefreshes counts the consecutive health samples (completed refresh or drift Jobs) that read pending since the last other reading or the last apply; unset otherwise. It spaces the refreshes while health is pending: 30s, then 1m, 2m, 4m and at most 5m. It lives in status only, so it restarts at 0 (30s) after clusterctl move. | No | Minimum: 1 | |
observedStateSerial integer | observedStateSerial is the Terraform state serial the outputs were read from. | No | Minimum: 1 | |
stateSecretSuffix string | stateSecretSuffix is the kubernetes backend secret_suffix of this object’s state. Informational: the controller derives it deterministically. | No | MaxLength: 63 MinLength: 1 | |
source SourceStatus | source records what the last Job actually ran. | No | MinProperties: 1 | |
stateBackups StateBackup array | stateBackups are the state backups the controller keeps (newest first), as of the last backup, prune or restore request. | No | MaxItems: 16 MinItems: 1 |
Conditions
Every condition CAPTF sets, per object kind, its polarity, and the meaning of each reason it can carry. See Observability for how conditions surface in kubectl describe and events.
Ready
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool, TerraformClusterIdentity.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | Ready | The Ready reason when every input is healthy. |
| True | SecretFound | The True reason when the identity’s credentials Secret exists. |
| False | NotReady | The Ready reason when an input is False. |
| False | SecretNotFound | The False reason when the identity’s Secret does not exist. |
| Unknown | ReadyUnknown | The Ready reason when an input is Unknown. |
DependenciesReady
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | DependenciesReady | The True reason. |
| False | ClusterNotTerraform | The False reason when the owning Cluster’s infrastructureRef is not a TerraformCluster; nothing is rendered or run. |
| False | OwnerMismatch | The False reason when an ownerRef of the expected kind resolves to an object that does not reference this one back (wrong or missing infrastructureRef, a UID mismatch, or a cluster-name label that disagrees with the owner): the ownerRef is forged or stale, so the object is treated as having no valid owner; no Job runs and nothing is written to the named owner or its Cluster. |
| False | OwnerNotFound | The False reason when the owner object is gone. |
| False | VariablesInvalid | The False reason when a variablesFrom source has a key that is not a Terraform identifier or is reserved, or a value that is not UTF-8 or (format JSON) not valid JSON. The message names the key, never the value; no Job starts. |
| False | VariablesSourceNotFound | The False reason when a ConfigMap or Secret named by spec.variablesFrom (not optional) is missing or does not carry captf.io/variables=true; no Job starts. |
| False | WaitingForOwnerMachine | The False reason when a fresh TerraformMachine has only a non-controller control-plane ownerRef and no Machine ownerRef yet. |
| False | WaitingForOwnerMachinePool | The False reason when a TerraformMachinePool has ownerRefs but no MachinePool ownerRef yet. |
| Unknown | WaitingForBootstrapData | The Unknown reason while the Machine’s or MachinePool’s bootstrap data Secret is not set or not found. |
| Unknown | WaitingForClusterExports | The Unknown reason while the cluster module’s exports output is not readable. |
| Unknown | WaitingForClusterInfrastructure | The Unknown reason while the TerraformCluster is not provisioned. |
| Unknown | WaitingForOwner | The Unknown reason while the owner is not set. |
IdentityAllowed
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | IdentityAllowed | The True reason. |
| False | IdentityNotFound | The False reason when the identity does not exist. |
| False | NamespaceNotAllowed | The False reason when allowedNamespaces excludes this namespace. |
| False | SecretNotFound | The False reason when the identity’s Secret does not exist. |
| Unknown | IdentityCheckFailed | The Unknown reason when the check could not be completed. |
CredentialsMirrored
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | Mirrored | The True reason. |
| False | MirrorFailed | The False reason when mirroring failed. |
| Unknown | MirrorPending | The Unknown reason before the first mirror. |
RunnerRBACReady
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | RBACReady | The True reason. |
| False | RBACFailed | The False reason when creating the binding failed. |
| False | ServiceAccountNotOptedIn | The False reason when an override ServiceAccount lacks the captf.io/runner=true label. |
ApplyJobSucceeded
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | ApplySucceeded | The True reason after an apply. |
| True | DestroySucceeded | The True reason after a destroy. |
| False | ApplyFailed | The False reason when an apply Job failed. |
| False | DestroyFailed | The False reason when a destroy Job failed. |
| False | DestructivePlanBlocked | The False reason when a TerraformCluster apply (including a drift remediation) stopped before a plan that deletes or replaces resources: the captf.io/approve-destructive-plan annotation does not name the inputs hash it renders. No apply of that hash runs until it does, or the inputs change. |
| False | IdentityNotAllowed | The False reason when no Job could be created because the identity is not allowed. |
| False | ImageInvalid | The False reason when the runner reported an image-layout error (missing /captf/module or a non-executable command). |
| False | ImagePullFailed | The False reason when the pod stayed in ErrImagePull/ImagePullBackOff past activeDeadlineSeconds. |
| False | InputsTooLarge | The False reason when the rendered root module and variables exceed the size a Secret can carry, so no Job starts. |
| False | JobDeadlineExceeded | The False reason when the Job hit activeDeadlineSeconds. |
| Unknown | NoApplyYet | The Unknown reason before the first apply completes. |
| Unknown | PlanAwaitingApproval | The Unknown reason while a TerraformCluster with applyPolicy Manual waits for the approval of the plan in status.plan (the captf.io/approve-plan annotation naming its hash). Unknown, so waiting never turns Ready False. |
| Unknown | PlanChanged | The Unknown reason when an approved apply planned other changes than the approved plan and stopped before applying them; the new plan in status.plan waits for its approval. |
| Unknown | WaitingForClusterOperation | The Unknown reason while a machine’s or pool’s apply or destroy waits for its TerraformCluster’s apply or destroy to finish. |
| Unknown | WaitingForMachineOperations | The Unknown reason while a TerraformCluster’s apply or destroy waits for its machines’ and pools’ applies and destroys in flight to finish; new ones wait for it meanwhile. |
| Unknown | WaitingForRunLease | The Unknown reason while the object’s run lease is held by another live Job, for example one another manager instance started: no Job starts until it finishes. It is also the DriftJobSucceeded reason for a refresh or drift that waits. |
StateReadable
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | StateRead | The True reason. |
| False | StateCorrupt | The False reason when the state cannot be decoded. |
| False | StateEncrypted | The False reason for OpenTofu-encrypted state, which v1 does not support. |
| False | StateInconsistent | The False reason when state chunks disagree. |
| False | StateLocked | The False reason while the state lock is held by something other than the object’s own runner, such as a workstation; every Job waits lockTimeoutSeconds for it and then fails. |
| False | StateLost | The False reason when a provisioned object’s state Secret is missing, or carries no inputs hash for an immutable kind: no Job runs until the state is restored. |
| Unknown | StateNotFound | The Unknown reason when no state exists yet. |
RestoreJobSucceeded
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | StateRestored | The True reason: the restore Job pushed the backup into the backend. |
| False | RestoreBackupNotFound | The False reason when the annotation names no existing backup (or is not a serial); no Job starts. |
| False | RestoreFailed | The False reason when the restore Job failed. It is not retried for the same serial until the annotation changes or the failed Job is deleted. |
| Unknown | WaitingForClusterOperation | The Unknown reason while a machine’s or pool’s apply or destroy waits for its TerraformCluster’s apply or destroy to finish. |
| Unknown | WaitingForMachineOperations | The Unknown reason while a TerraformCluster’s apply or destroy waits for its machines’ and pools’ applies and destroys in flight to finish; new ones wait for it meanwhile. |
| Unknown | WaitingForRunLease | The Unknown reason while the object’s run lease is held by another live Job, for example one another manager instance started: no Job starts until it finishes. It is also the DriftJobSucceeded reason for a refresh or drift that waits. |
OutputsValid
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | InstancesTruncated | A True reason: the machinepool role’s instances output had more entries than the controller keeps (internal/outputs.MaxInstances) and was shortened. The pool still provisions; nothing else about its outputs is invalid. |
| True | OutputsValid | The True reason. |
| False | FailureDomainMismatch | The False reason when a machine’s failure_domain output differs from the requested failure domain. |
| False | OutputsInvalid | The False reason when an output violates the contract or a Cluster API marker. |
| False | OutputsMissing | The False reason when a required output is not declared. |
| False | ProviderIDChanged | The False reason when provider_id changed after it was first written. |
| Unknown | OutputsPending | The Unknown reason while required outputs are null. |
InfrastructureHealthy
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | Healthy | The True reason (state running, healthy true). |
| False | InstanceDegraded | The False reason for health state degraded. |
| False | InstancePending | The False reason for health state pending. |
| False | InstanceStopped | The False reason for health state stopped. |
| False | InstanceTerminated | The False reason for health state terminated or a vanished instance. |
| False | InstanceUnhealthy | The False reason for state running, healthy false. |
| False | Provisioning | The False reason from the first apply start until provisioned. |
| Unknown | HealthUnknown | The Unknown reason for health state unknown. |
| Unknown | WaitingForProvisioning | The Unknown reason before the first apply. |
DriftJobSucceeded
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | DriftChecked | The True reason. |
| False | DriftJobDeadlineExceeded | The False reason when the drift Job hit activeDeadlineSeconds. |
| False | DriftJobFailed | The False reason when the drift Job failed. |
| Unknown | DriftJobRunning | The Unknown reason while a drift Job runs. |
| Unknown | DriftNotChecked | The Unknown reason before the first drift check (also used by DriftDetected). |
| Unknown | WaitingForRunLease | The Unknown reason while the object’s run lease is held by another live Job, for example one another manager instance started: no Job starts until it finishes. It is also the DriftJobSucceeded reason for a refresh or drift that waits. |
DriftDetected
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: negative (True is the problem).
| Status | Reason | Meaning |
|---|---|---|
| True | DriftPending | The True reason while remediation is pending, and after a failed cluster re-apply. |
| True | DriftRemediating | The True reason while a remediation apply runs. |
| True | DriftReported | The True reason with drift action Report. |
| False | NoDrift | The False reason. |
| Unknown | DriftNotChecked | The Unknown reason before the first drift check (also used by DriftDetected). |
DeletionBlocked
Carried by: TerraformCluster.
Polarity: negative (True is the problem).
| Status | Reason | Meaning |
|---|---|---|
| True | DependentsExist | The True reason. |
| False | NotBlocked | The False reason. |
EndpointAvailable
Carried by: TerraformCluster.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | EndpointAvailable | The True reason. |
| False | WaitingForEndpoint | The False reason when the cluster is provisioned and neither the module output nor Cluster.spec has a valid endpoint. |
AutoscalingActive
Carried by: TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | ReplicasManagedByModule | The True reason: both annotations are present and valid, and the controller writes the observed replicas back to MachinePool.spec.replicas. |
| False | AutoscalingAnnotationsInvalid | The False reason when an annotation is present but the pair is incomplete, unparsable, or min > max; the message names the problem. The pool still applies without autoscaling. |
| False | AutoscalingDisabled | The False reason when neither annotation is set: MachinePool.spec.replicas is the sole source of desired capacity. |
CapacityResolved
Carried by: TerraformMachineTemplate.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | CapacityNotDeclared | The True reason when the image carries neither label. |
| True | CapacityResolved | The True reason when both labels parsed. |
| False | CapacityLabelInvalid | The False reason when a label is present but invalid. |
| False | ImageInspectFailed | The False reason when the registry fetch or authentication failed. |
Paused
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: normal (True is healthy).
| Status | Reason | Meaning |
|---|---|---|
| True | Paused | When an object is paused. |
| False | NotPaused | When an object is not paused. |
Deleting
Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool.
Polarity: negative (True is the problem).
| Status | Reason | Meaning |
|---|---|---|
| True | Deleting | When an object is deleting because the DeletionTimestamp is set; used if none of the more specific reasons apply. |
| True | DeletionCompleted | When the deletion process has completed; set right after the corresponding finalizer is removed. |
| False | NotDeleting | When an object is not deleting because the DeletionTimestamp is not set. |
Ready summarization
Ready summarizes other conditions (sigs.k8s.io/cluster-api/util/conditions.SetSummaryCondition); it is the only condition Cluster API itself reads, mirrored into the Cluster’s or Machine’s InfrastructureReady. Which conditions feed it depends on the kind and on whether status.initialization.provisioned has latched true.
Before provisioning, every kind below summarizes the same inputs:
DependenciesReadyIdentityAllowedCredentialsMirroredRunnerRBACReadyApplyJobSucceededStateReadableOutputsValidInfrastructureHealthyDeleting
After provisioning, the inputs differ per kind:
- TerraformCluster:
InfrastructureHealthy,Deleting - TerraformMachine:
InfrastructureHealthy,Deleting - TerraformMachinePool:
InfrastructureHealthy,ApplyJobSucceeded,Deleting
TerraformMachinePool includes ApplyJobSucceeded after provisioning (unlike the cluster and the machine): a pool is mutable and re-applied on bootstrap rotation, so a failed re-apply must be visible in Ready.
TerraformClusterIdentity has a Ready condition of its own, set directly from whether its credentials Secret exists (SecretFound/SecretNotFound), not summarized from other conditions.
Events
Every reason CAPTF’s manager and runner record as a Kubernetes Event (kubectl get events or kubectl describe), its type, which object kind it is emitted on, and what it means. See Observability for how to watch them.
Manager events
Emitted by the manager’s reconcilers through the events.k8s.io/v1 recorder, at most once per transition or occurrence.
| Reason | Type | Emitted on | Meaning |
|---|---|---|---|
CapacityResolved | Normal | TerraformMachineTemplate | A TerraformMachineTemplate’s capacity or nodeInfo changed from its image labels. |
ConditionChanged | Normal or Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | Any other owned condition changed status or reason; Normal into its good or an informational state, Warning into its bad state. |
ControlPlaneEndpointSet | Normal | TerraformCluster | A TerraformCluster’s spec.controlPlaneEndpoint was written from the module output. |
DeletionStarted | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The first reconcile with a deletionTimestamp. |
Destroyed | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The destroy succeeded and cleanup ran. |
DestructivePlanApprovalConsumed | Normal | TerraformCluster | The approved destructive apply succeeded and its approval annotation was removed. |
DestructivePlanBlocked | Warning | TerraformCluster | A TerraformCluster apply stopped before a plan that deletes or replaces resources; once per blocked Job, in place of JobFailed. |
DigestPinned | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | An image digest was recorded on the durable inputs Secret, or re-pinned after an apply of a mutable kind. |
DigestUnknown | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | No digest could be pinned, or an operation runs the spec reference for lack of one. |
DriftDetected | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A drift check found a difference. |
DriftRemediationStarted | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | An apply remediating drift started. |
DriftResolved | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | DriftDetected went from True to False. |
FailureDomainsChanged | Normal | TerraformCluster | A TerraformCluster’s status.failureDomains changed. |
FinalizerRemoved | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The finalizer was removed; the object goes. |
ForceUnlocked | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A stale state lock was force-unlocked. |
IdentityNotAllowed | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | The identity does not allow the namespace. |
IdentitySecretFound | Normal | TerraformClusterIdentity | A TerraformClusterIdentity’s credentials Secret appeared. |
IdentitySecretNotFound | Warning | TerraformClusterIdentity | A TerraformClusterIdentity’s credentials Secret went missing. |
ImageInspectFailed | Warning | TerraformMachineTemplate | The registry could not be read for capacity. |
InputsChanged | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The inputs hash differs from the state’s and an apply of the new inputs starts. |
InstanceHealthy | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | InfrastructureHealthy became True. |
InstanceUnhealthy | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | InfrastructureHealthy became False for an unhealthy, degraded, stopped or terminated instance. |
JobCreated | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job was started (op, attempt, image, why). |
JobDeadlineExceeded | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job hit activeDeadlineSeconds. |
JobFailed | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job failed, or an apply or destroy could not start (ApplyJobSucceeded False without a Job). |
JobInterrupted | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job was stopped from outside (a drain, eviction or deletion); it is retried without backoff. |
JobSucceeded | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | An apply, destroy, refresh or drift Job succeeded. |
MirrorCreated | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The credential mirror of the namespace was created on behalf of this object. |
MirrorRemoved | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The credential mirror of the namespace was deleted on behalf of this object. |
OutputsInvalid | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | The module’s outputs broke the contract. |
Paused | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The Paused condition changed to True; reconciliation stops starting Jobs. |
PlanApplied | Normal | TerraformCluster | The approved plan was applied and its approval annotation removed. |
PlanApproved | Normal | TerraformCluster | The apply of an approved plan started. |
PlanChanged | Warning | TerraformCluster | An approved apply planned other changes and stopped before applying them; once per such Job. |
PlanReady | Normal | TerraformCluster | A plan Job planned a TerraformCluster’s change under applyPolicy Manual (counts, plan hash, the approve command); once per plan Job. |
ProviderIDSet | Normal | TerraformMachine, TerraformMachinePool | A TerraformMachine’s or TerraformMachinePool’s spec.providerID was written (a pool’s may change). |
Provisioned | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | Provisioned latched true. |
RemediationRequested | Warning | TerraformMachine | The owner Machine was annotated with cluster.x-k8s.io/remediate-machine. |
RemediationWithdrawn | Normal | TerraformMachine | The instance read Healthy again and the annotation CAPTF set was removed from the owner Machine. |
ReplicasWrittenBack | Normal | TerraformMachinePool | An autoscaled pool’s observed replicas output was patched onto MachinePool.spec.replicas (“X → Y”). |
Resumed | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The Paused condition changed to False after being True; reconciliation resumes. |
StateAdopted | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The state written by a successful apply was adopted with its new inputs hash. |
StateBackedUp | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | A new state serial was copied into a backup (once per backup). |
StateLocked | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | The state lock is held by something else (StateReadable False/StateLocked). |
StateLost | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A provisioned object’s state is gone or carries no inputs hash (StateReadable False/StateLost). |
StateRestoreFailed | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A restore Job failed; it is not retried for the same serial. |
StateRestored | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | A restore Job pushed a backup into the backend and the captf.io/restore-state annotation was removed. |
StateUnreadable | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | The state could not be read. |
StuckJobDeleted | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job that could never start (its per-run Secret is missing) was deleted to be started again. |
WaitingForClusterOperation | Normal | TerraformMachine, TerraformMachinePool | A machine’s apply or destroy waits for its TerraformCluster’s apply or destroy. |
WaitingForMachineOperations | Normal | TerraformCluster | A TerraformCluster’s apply or destroy waits for its machines’ applies and destroys in flight. |
WaitingForRunLease | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | An operation waits because another live Job holds the object’s run lease. |
Runner events
Emitted only when the manager runs with --runner-events (its default is true): each Job’s runner posts its own progress as Events on the object the Job is for, related to the Job itself. Emission is best effort and never fails or slows the run.
| Reason | Type | Emitted on | Meaning |
|---|---|---|---|
PlanSummary | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | A plan the runner parsed (drift, or a guarded cluster apply): counts only. |
ResourcesChanged | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | What an apply or destroy step changed, from the runtime’s summary line: counts only. |
RunFinished | Normal or Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | The run ended (result, total duration). |
RunStarted | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | The runtime is ready and the first step is about to run (op, image reference, runtime version). |
StepFailed | Warning | TerraformCluster, TerraformMachine, TerraformMachinePool | A runtime step failed; the note carries the runner’s curated failure summary, never raw stderr. |
StepStarted | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | A runtime step started. |
StepSucceeded | Normal | TerraformCluster, TerraformMachine, TerraformMachinePool | A runtime step finished successfully. |
Alerts
CAPTF ships a PrometheusRule (config/prometheus/rules.yaml) alerting on the metrics in Metrics; each alert below links to its runbook in Observability. Promtool unit tests for the rules live in config/prometheus/tests/rules_test.yaml.
CAPTFJobFailing
- Group: captf
- Severity: warning
sum by (kind, op) (increase(captf_jobs_total{result=~"failed|deadline"}[30m])) > 2
Summary: {{ $labels.kind }} {{ $labels.op }} Jobs keep failing
Description: More than two {{ $labels.op }} Jobs for {{ $labels.kind }} objects failed or hit their deadline in 30 minutes. Find the objects with ApplyJobSucceeded=False or DriftJobSucceeded=False and read status.lastRun and the Job logs.
Runbook: ../operator-guide/observability.md#captfjobfailing
CAPTFDestroyStuck
- Group: captf
- Severity: critical
- For: 30m
sum by (kind) (increase(captf_jobs_total{op="destroy",result!="succeeded"}[30m])) > 0
Summary: {{ $labels.kind }} destroy keeps failing
Description: Destroy Jobs for {{ $labels.kind }} objects have failed for 30 minutes. The objects keep their finalizer and their state; follow the stuck-destroy runbook.
Runbook: ../operator-guide/observability.md#captfdestroystuck
CAPTFClusterDrift
- Group: captf
- Severity: warning
- For: 1h
captf_drift_detected{kind="TerraformCluster"} == 1
Summary: TerraformCluster {{ $labels.namespace }}/{{ $labels.name }} has drifted
Description: The last drift check found changes for an hour. With drift.action Report this waits for an operator; with Remediate the remediation apply keeps failing (see ApplyJobSucceeded).
Runbook: ../operator-guide/observability.md#captfclusterdrift
CAPTFStateUnreadable
- Group: captf
- Severity: critical
sum by (kind, reason) (increase(captf_state_read_errors_total[15m])) > 0
Summary: {{ $labels.kind }} state became unreadable ({{ $labels.reason }})
Description: A {{ $labels.kind }} state Secret turned {{ $labels.reason }}. Nothing is applied until it reads again. Find the object with StateReadable=False.
Runbook: ../operator-guide/observability.md#captfstateunreadable
CAPTFForceUnlocks
- Group: captf
- Severity: warning
sum by (kind) (increase(captf_lock_force_unlocks_total[1h])) > 0
Summary: A {{ $labels.kind }} state lock was force-unlocked
Description: A Job force-unlocked a state lock whose holder pod was gone. Look for ForceUnlocked events and check why the previous Job died.
Runbook: ../operator-guide/observability.md#captfforceunlocks
CAPTFReconcileErrors
- Group: captf
- Severity: warning
- For: 10m
sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m])) > 0.1
Summary: The {{ $labels.controller }} controller keeps failing to reconcile
Description: Reconcile errors above 0.1/s for 10 minutes; read the manager logs.
Runbook: ../operator-guide/observability.md#captfreconcileerrors
CAPTFJobSlow
- Group: captf
- Severity: info
- For: 30m
histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1h]))) > 1800
Summary: {{ $labels.kind }} {{ $labels.op }} Jobs are slow
Description: The 90th percentile of {{ $labels.op }} Job duration has been above 30 minutes for 30 minutes.
Runbook: ../operator-guide/observability.md#captfjobslow
CAPTFJobQueueSlow
- Group: captf
- Severity: warning
- For: 10m
histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_queue_seconds_bucket[10m]))) > 300
Summary: {{ $labels.kind }} {{ $labels.op }} Jobs wait long to start
Description: The 90th percentile of the time {{ $labels.op }} Jobs for {{ $labels.kind }} objects take from creation to running the module has been above 5 minutes. Look for Pending runner pods: node capacity, quota, slow or failing image pulls.
Runbook: ../operator-guide/observability.md#captfjobqueueslow
CAPTFStateNearSecretLimit
- Group: captf
- Severity: warning
- For: 15m
captf_state_bytes > 900 * 1024
Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} state is near the Secret size limit
Description: The compressed state is {{ $value | humanize1024 }}B, above 900 KiB. An OpenTofu state Secret cannot grow past 1 MiB, and the next apply that crosses it fails to save state.
Runbook: ../operator-guide/observability.md#captfstatenearsecretlimit
CAPTFInputsNearLimit
- Group: captf
- Severity: warning
- For: 15m
captf_inputs_bytes > 900000
Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} inputs are near the size limit
Description: The rendered main.tf.json and terraform.tfvars.json are {{ $value | humanize }} bytes. Above 1000000 no Job starts (ApplyJobSucceeded InputsTooLarge).
Runbook: ../operator-guide/observability.md#captfinputsnearlimit
CAPTFNoRecentSuccess
- Group: captf
- Severity: warning
- For: 30m
time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 6 * 3600
Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} has had no successful {{ $labels.op }} for 6 hours
Description: Its {{ $labels.op }} Jobs keep failing or never start, so drift and health go unobserved. See DriftJobSucceeded and status.lastRun.
Runbook: ../operator-guide/observability.md#captfnorecentsuccess
CAPTF metrics
Declared with k8s.io/component-base/metrics (as Kubernetes components declare their own), all at StabilityLevel ALPHA; component-base prefixes every HELP string below with [ALPHA] and a space on the wire. Served on the diagnostics endpoint (--diagnostics-address) on controller-runtime’s registry, merged by internal/metrics.Bridge with controller-runtime’s own defaults (controller_runtime_*, workqueue_*, rest_client_*) and component-base’s legacyregistry (kubernetes_feature_enabled and any other component-base series; its go_* and process_* families are dropped, already served by controller-runtime’s own registry). Labels are bounded enums (kind, op, result, reason, step, action, error_kind); only the 8 per-object gauges carry namespace and name, and they are removed with their object.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
captf_jobs_total | counter | kind, op, result | Jobs completed, by result: succeeded, failed, deadline, interrupted (stopped from outside: a drain, eviction or deletion), blocked (a TerraformCluster apply stopped before a plan that deletes or replaces resources, awaiting approval) or plan_changed (an apply approved for one plan planned other changes and stopped, applyPolicy Manual). Op plan is a plan Job that applies nothing. |
captf_job_duration_seconds | histogram | kind, op, result | Wall time of a completed Job, from start to finish, by the result of captf_jobs_total. |
captf_job_step_duration_seconds | histogram | kind, op, step | Wall time of each runner step of a completed Job (init, force-unlock, validate, plan, show-json, apply, apply-refresh-only, destroy, state-push, state-list, prepare; anything else is other). |
captf_job_queue_seconds | histogram | kind, op | Time from a Job’s creation to its source container’s start: scheduling, image pulls and the runner copy. Not observed when the pod reports no start. |
captf_job_errors_total | counter | kind, op, error_kind, step | Jobs that did not succeed, by the runner’s error kind (step, image-layout, interrupted, blocked, plan-changed; deadline or unknown without a result) and failing step (none when no step failed). |
captf_jobs_active | gauge | kind, op | Jobs currently running, counted from the Job cache at scrape time. |
captf_job_attempts | histogram | kind, op | Retry number of a Job that succeeded: 1 plus the failed Jobs of its op since that op last succeeded (interrupted Jobs do not count). |
captf_resources_changed_total | counter | kind, op, action | Resources an apply or destroy Job changed, by action (add, change, destroy, import), from the runtime’s final summary line. |
captf_drift_resources_total | counter | kind, action | Resources a drift Job that detected drift found to add, change or destroy. |
captf_reconcile_op_decisions_total | counter | kind, op, reason | What the reconcile decided to run (op none: nothing) and why. |
captf_state_read_errors_total | counter | kind, reason | State reads that turned unreadable: inconsistent, encrypted or corrupt (an unsupported state version counts as corrupt), lost (a provisioned object’s state is gone) or locked (held by a holder that is not this object’s runner). |
captf_outputs_invalid_total | counter | kind, reason | Outputs that turned invalid against the module contract. |
captf_drift_detected | gauge | kind, namespace, name | 1 while DriftDetected is True, else 0. |
captf_ready | gauge | kind, namespace, name | 1, 0 or -1 for a True, False or Unknown Ready condition. |
captf_infrastructure_healthy | gauge | kind, namespace, name | 1, 0 or -1 for a True, False or Unknown InfrastructureHealthy condition. |
captf_state_resources | gauge | kind, namespace, name | Managed resources (not data sources) in the object’s state, as last read. |
captf_state_bytes | gauge | kind, namespace, name | Compressed size of the object’s state summed over its Secrets, as last read; the kubernetes backend holds at most 1 MiB per Secret. |
captf_inputs_bytes | gauge | kind, namespace, name | Size of the object’s rendered main.tf.json and terraform.tfvars.json (the durable inputs, or the last render); no Job starts above 1000000. |
captf_last_success_timestamp_seconds | gauge | kind, namespace, name, op | Unix time the newest successful Job of an op finished. Drift and refresh are exported only while that op is scheduled (not deleting or paused, with a drift interval or health checks). |
captf_unhealthy_samples | gauge | namespace, name | A TerraformMachine’s consecutive unhealthy health samples (status.unhealthySamples). |
captf_lock_force_unlocks_total | counter | kind | Stale state locks force-unlocked. |
captf_identity_denied_total | counter | reason | Identity refusals: notfound or namespace. |
captf_image_inspect_errors_total | counter | reason | Registry or image-label failures while resolving template capacity. |
captf_inputs_hash_changes_total | counter | kind | Applies started because the inputs of a mutable kind changed. |
captf_remediation_requests_total | counter | action | cluster.x-k8s.io/remediate-machine annotations set on (requested) or removed from (withdrawn) a Machine. |
captf_destructive_plan_approvals_consumed_total | counter | kind | Destructive-plan approvals removed after the approved apply succeeded. |
captf_lease_waits_total | counter | kind, reason | Operations that started waiting for a run lease, once per wait: run_lease (another live Job of the object holds it), cluster_operation (a machine’s apply or destroy waits for its TerraformCluster’s) or machine_operations (a TerraformCluster’s apply or destroy waits for its machines’). |
captf_state_backups_total | counter | kind, result | State backups: taken (a new state serial copied into captf-state-backup-* Secrets), pruned (a backup beyond –state-backups deleted) or skipped (a new serial not backed up: encrypted, unreadable or oversized state, or a failed copy). |
captf_state_restores_total | counter | kind, result | State restores requested with captf.io/restore-state: succeeded or failed (a restore Job finished), or not_found (the annotation names no backup). |
captf_plan_approvals_total | counter | kind, result | Plans approved with captf.io/approve-plan (applyPolicy Manual): approved (the apply of the approved plan succeeded and the annotation was removed) or changed (the approved apply planned other changes and stopped; the new plan waits for approval). |
captf_build_info | gauge | version, commit, contract | A metric with a constant ‘1’ value labeled by the version, commit and module contract of the manager. |
Manager Flags
The CAPTF controller manager (manager) is a Kubernetes controller-runtime binary, built on cobra and k8s.io/component-base, following the same flag conventions as kube-controller-manager and friends. Flags below use dashes; the command line also accepts underscores (cliflag.WordSepNormalizeFunc).
CAPTF
| Flag | Type | Default | Description |
|---|---|---|---|
--cluster-operation-gate | bool | true | Keep a TerraformCluster’s apply or destroy and its machines’ applies and destroys from running at once, through a per-Cluster write Lease. The per-object run Lease is always on. |
--drift-default-interval | duration | 30m0s | Drift check interval for objects that set none. |
--runner-events | bool | true | Have Job runners emit progress events (RunStarted, Step*, PlanSummary, ResourcesChanged, RunFinished) on the owning Terraform* object. Needs events create in the runner ClusterRole. |
--runner-image | string | Image of the init container that injects the runner binary into Jobs. Defaults to $CAPTF_MANAGER_IMAGE, the manager’s own image. | |
--state-backups | int | 5 | State backups to keep per object: every new state serial is copied into captf-state-backup-* Secrets and older copies are pruned. 0 takes no backups (existing ones stay and can still be restored). |
Controllers and concurrency
| Flag | Type | Default | Description |
|---|---|---|---|
--namespace | string | Namespace that the controller watches to reconcile objects. If unspecified, the controller watches all namespaces. | |
--sync-period | duration | 10m0s | The minimum interval at which watched resources are reconciled (e.g. 15m) |
--terraformcluster-concurrency | int | 10 | Number of TerraformClusters to process simultaneously |
--terraformmachine-concurrency | int | 10 | Number of TerraformMachines to process simultaneously |
--terraformmachinepool-concurrency | int | 10 | Number of TerraformMachinePools to process simultaneously |
--terraformmachinetemplate-concurrency | int | 10 | Number of TerraformMachineTemplates to process simultaneously |
--watch-filter | string | Label value that the controller watches to reconcile objects. Label key is always cluster.x-k8s.io/watch-filter. If unspecified, the controller watches all objects. |
Leader election
| Flag | Type | Default | Description |
|---|---|---|---|
--leader-elect | bool | false | Enable leader election for the controller manager, ensuring there is only one active manager. |
--leader-elect-lease-duration | duration | 15s | Interval at which non-leader candidates will wait to force acquire leadership (duration string) |
--leader-elect-renew-deadline | duration | 10s | Duration that the leading controller manager will retry refreshing leadership before giving up (duration string) |
--leader-elect-retry-period | duration | 2s | Duration the LeaderElector clients should wait between tries of actions (duration string) |
Webhooks
| Flag | Type | Default | Description |
|---|---|---|---|
--webhook-cert-dir | string | /tmp/k8s-webhook-server/serving-certs/ | Directory holding the webhook server’s serving certificate and key (mounted from the cert-manager Secret). |
--webhook-cert-name | string | tls.crt | File name of the serving certificate in –webhook-cert-dir. |
--webhook-key-name | string | tls.key | File name of the serving key in –webhook-cert-dir. |
--webhook-port | int | 9443 | Port the webhook server listens on. |
Diagnostics and TLS (CAPI)
| Flag | Type | Default | Description |
|---|---|---|---|
--diagnostics-address | string | :8443 | The address the diagnostics endpoint binds to. Per default metrics are served via https and withauthentication/authorization. To serve via http and without authentication/authorization set –insecure-diagnostics. If –insecure-diagnostics is not set the diagnostics endpoint also serves pprof endpoints and an endpoint to change the log level. |
--insecure-diagnostics | bool | false | Enable insecure diagnostics serving. For more details see the description of –diagnostics-address. |
--tls-cipher-suites | stringSlice | [] | Comma-separated list of cipher suites for the webhook server and metrics server (the latter only if –insecure-diagnostics is not set to true). If omitted, the default Go cipher suites will be used. Preferred values: TLS_AES_128_GCM_SHA256, TLS_AES_256_GCM_SHA384, TLS_CHACHA20_POLY1305_SHA256, TLS_ECDHE_ECDSA_WITH_AES_128_CBC_SHA, TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256, TLS_ECDHE_ECDSA_WITH_AES_256_CBC_SHA, TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384, TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305, TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305_SHA256, TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA, TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256, TLS_ECDHE_RSA_WITH_AES_256_CBC_SHA, TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384, TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305, TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305_SHA256. Insecure values: TLS_ECDHE_ECDSA_WITH_AES_128_CBC_SHA256, TLS_ECDHE_ECDSA_WITH_RC4_128_SHA, TLS_ECDHE_RSA_WITH_3DES_EDE_CBC_SHA, TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA256, TLS_ECDHE_RSA_WITH_RC4_128_SHA, TLS_RSA_WITH_3DES_EDE_CBC_SHA, TLS_RSA_WITH_AES_128_CBC_SHA, TLS_RSA_WITH_AES_128_CBC_SHA256, TLS_RSA_WITH_AES_128_GCM_SHA256, TLS_RSA_WITH_AES_256_CBC_SHA, TLS_RSA_WITH_AES_256_GCM_SHA384, TLS_RSA_WITH_RC4_128_SHA. |
--tls-curve-preferences | int32Slice | [] | Comma-separated list of numeric Go crypto/tls CurveID values, as the allowed key exchange mechanisms for the webhook server and metrics server (the latter only if –insecure-diagnostics is not set to true). The supported values depend on the Go version used. See https://pkg.go.dev/crypto/tls#CurveID for values supported for each Go version. The order of the list is ignored, and key exchange mechanisms are chosen by Go from this list using an internal preference order. If omitted, the default Go curves will be used. |
--tls-min-version | string | VersionTLS12 | The minimum TLS version in use by the webhook server and metrics server (the latter only if –insecure-diagnostics is not set to true). Possible values are VersionTLS10, VersionTLS11, VersionTLS12, VersionTLS13. |
Logging
| Flag | Type | Default | Description |
|---|---|---|---|
--feature-gates | mapStringBool | A set of key=value pairs that describe feature gates for alpha/experimental features. Options are: AllAlpha=true|false (ALPHA - default=false) AllBeta=true|false (BETA - default=false) ContextualLogging=true|false (BETA - default=true) LoggingAlphaOptions=true|false (ALPHA - default=false) LoggingBetaOptions=true|false (BETA - default=true) | |
--log-flush-frequency | duration | 5s | Maximum number of seconds between log flushes |
--log-json-info-buffer-size | quantity | 0 | [Alpha] In JSON format with split output streams, the info messages can be buffered for a while to increase performance. The default value of zero bytes disables buffering. The size can be specified as number of bytes (512), multiples of 1000 (1K), multiples of 1024 (2Ki), or powers of those (3M, 4G, 5Mi, 6Gi). Enable the LoggingAlphaOptions feature gate to use this. |
--log-json-split-stream | bool | false | [Alpha] In JSON format, write error messages to stderr and info messages to stdout. The default is to write a single stream to stdout. Enable the LoggingAlphaOptions feature gate to use this. |
--log-text-info-buffer-size | quantity | 0 | [Alpha] In text format with split output streams, the info messages can be buffered for a while to increase performance. The default value of zero bytes disables buffering. The size can be specified as number of bytes (512), multiples of 1000 (1K), multiples of 1024 (2Ki), or powers of those (3M, 4G, 5Mi, 6Gi). Enable the LoggingAlphaOptions feature gate to use this. |
--log-text-split-stream | bool | false | [Alpha] In text format, write error messages to stderr and info messages to stdout. The default is to write a single stream to stdout. Enable the LoggingAlphaOptions feature gate to use this. |
--logging-format | string | text | Sets the log format. Permitted formats: “json” (gated by LoggingBetaOptions), “text”. |
-v, --v | Level | 2 | number for the log level verbosity |
--vmodule | pattern=N,... | comma-separated list of pattern=N settings for file-filtered logging (only works for text log format) |
Other
| Flag | Type | Default | Description |
|---|---|---|---|
--health-addr | string | :9440 | The address the health endpoint binds to. |
-h, --help | bool | false | help for manager |
--kubeconfig | string | Paths to a kubeconfig. Only required if out-of-cluster. | |
--profiler-address | string | Bind address to expose the pprof profiler (e.g. localhost:6060) | |
--version | version | false | –version, –version=raw prints version information and quits; –version=vX.Y.Z… sets the reported version |
Feature gates
--feature-gates takes a comma-separated Key=value list. v1 registers no CAPTF feature of its own; the gates below are the component-base logging gates, which logsv1.ValidateAndApply reads.
| Name | Default | Maturity |
|---|---|---|
AllAlpha | false | ALPHA |
AllBeta | false | BETA |
ContextualLogging | true | BETA |
LoggingAlphaOptions | false | ALPHA |
LoggingBetaOptions | true | BETA |
Environment variables
This list is maintained by hand in internal/docsgen/manager.go, alongside the flags it documents.
| Variable | Description |
|---|---|
CAPTF_MANAGER_IMAGE | Default of –runner-image, the image of the init container that injects the runner binary into Jobs. Set by the shipped Deployment to the manager’s own image. |
KUBECONFIG | Read by controller-runtime’s –kubeconfig flag handling (ctrl.RegisterFlags) when –kubeconfig is not given and the manager runs out-of-cluster. |
version
manager version and manager --version (or --version=raw) both print the build stamp and exit.
Shipped manifest
config/manager/manager.yaml sets these arguments on the manager container:
| Flag | Value |
|---|---|
--leader-elect | |
--diagnostics-address | :8443 |
--insecure-diagnostics | false |
--webhook-port | 9443 |
Runner CLI
runner is the Job entrypoint the manager copies into every source image’s init container and runs as the main container’s command (copy, then run); it is not a tool operators invoke directly. See Image Contract for how the image and the runner fit together.
Global flags
Every subcommand inherits these (component-base logging and version flags, as every Kubernetes component registers them).
| Flag | Type | Default | Description |
|---|---|---|---|
--feature-gates | mapStringBool | A set of key=value pairs that describe feature gates for alpha/experimental features. Options are: AllAlpha=true|false (ALPHA - default=false) AllBeta=true|false (BETA - default=false) ContextualLogging=true|false (BETA - default=true) LoggingAlphaOptions=true|false (ALPHA - default=false) LoggingBetaOptions=true|false (BETA - default=true) | |
--log-flush-frequency | duration | 5s | Maximum number of seconds between log flushes |
--log-json-info-buffer-size | quantity | 0 | [Alpha] In JSON format with split output streams, the info messages can be buffered for a while to increase performance. The default value of zero bytes disables buffering. The size can be specified as number of bytes (512), multiples of 1000 (1K), multiples of 1024 (2Ki), or powers of those (3M, 4G, 5Mi, 6Gi). Enable the LoggingAlphaOptions feature gate to use this. |
--log-json-split-stream | bool | false | [Alpha] In JSON format, write error messages to stderr and info messages to stdout. The default is to write a single stream to stdout. Enable the LoggingAlphaOptions feature gate to use this. |
--log-text-info-buffer-size | quantity | 0 | [Alpha] In text format with split output streams, the info messages can be buffered for a while to increase performance. The default value of zero bytes disables buffering. The size can be specified as number of bytes (512), multiples of 1000 (1K), multiples of 1024 (2Ki), or powers of those (3M, 4G, 5Mi, 6Gi). Enable the LoggingAlphaOptions feature gate to use this. |
--log-text-split-stream | bool | false | [Alpha] In text format, write error messages to stderr and info messages to stdout. The default is to write a single stream to stdout. Enable the LoggingAlphaOptions feature gate to use this. |
--logging-format | string | text | Sets the log format. Permitted formats: “json” (gated by LoggingBetaOptions), “text”. |
-v, --v | Level | 0 | number for the log level verbosity |
--version | version | false | –version, –version=raw prints version information and quits; –version=vX.Y.Z… sets the reported version |
--vmodule | pattern=N,... | comma-separated list of pattern=N settings for file-filtered logging (only works for text log format) |
copy <dest>
Copy this binary to <dest> (init container).
run
Run one operation in the Job’s main container.
| Flag | Type | Default | Description |
|---|---|---|---|
--allow-deletes-hash | string | inputs hash approved for a destructive plan | |
--backend-config | stringArray | [] | init -backend-config value (repeatable) |
--bin | stringArray | [] | runtime command, one flag per element (default /captf/runtime) |
--config | string | /captf/config | directory holding the rendered root module and tfvars (the per-run Secret’s mount) |
--event-object | string | emit progress events about <apiVersion>/<kind>/<namespace>/<name>/<uid> (none when unset) | |
--expect-plan | string | apply: the approved plan hash; stop with error kind plan-changed unless the plan’s hash is this one | |
--force-unlock | string | stale lock to force-unlock after init | |
--guard-deletes | bool | false | apply: stop before a plan that deletes or replaces a resource unless –allow-deletes-hash is –inputs-hash |
--image | string | source image reference, echoed in the result | |
--inputs-hash | string | hash of the inputs the Job renders | |
--job-name | string | the Job the events relate to | |
--lock-timeout | duration | 5m0s | state lock timeout |
--module | string | /captf/module | module directory |
--op | string | operation: apply, destroy, refresh, drift, restore or plan | |
--providers | string | /captf/providers | provider mirror directory (optional) |
--restore-chunks | int | 0 | restore: the number of backup chunks under <config>/restore |
--restore-resources | int | 0 | restore: the backup’s managed resource count; state list must show one when it is not 0 |
--result | string | /dev/termination-log | where to write the result |
--stop-timeout | duration | 1m0s | time an interrupted step gets to stop before SIGKILL |
--workdir | string | /captf/work | writable work directory |
version
Print the version and exit.
Exit codes
| Code | Name | Meaning |
|---|---|---|
0 | ExitOK | Every step succeeded, or an apply’s plan had no changes to apply. |
1 | ExitFailure | A step failed, was interrupted, stopped before a destructive plan (blocked), or found its approved plan had changed; or preflight/prepare itself failed. max(runtime exit, 1) when the failing step’s own exit code is not already at least 1. |
2 | ExitUsage | Bad input: an –op Steps does not recognize, a bad flag, unexpected arguments, or invalid logging flags. |
Result error kinds
runner run always writes a result document (--result, the termination log by default); a failed run’s error.kind is one of:
| Kind | Meaning |
|---|---|
image-layout | The image does not follow the image contract (a Preflight check failed): the module or runtime is not where the contract puts it. |
step | A runtime step (init, plan, apply, …) failed, or the runner’s own setup (Prepare, a restore assembly) failed. |
interrupted | The run’s context was canceled (a drain, eviction or deletion) while a step was running; counts toward no retry backoff. |
blocked | A guarded apply stopped before a plan that deletes or replaces resources, without an approval for its inputs hash. It changed nothing and is not retried until the inputs or the approval change. |
plan-changed | An apply approved for one plan (–expect-plan) found its plan is now another. It changed nothing, carries the new plan, and waits for its approval. |
tfcapi-lint CLI
tfcapi-lint checks a Terraform/OpenTofu module, or the image built from it, against the CAPTF module contract. See tfcapi-lint for how to install and run it.
Global flags
| Flag | Type | Default | Description |
|---|---|---|---|
--version | version | false | –version, –version=raw prints version information and quits; –version=vX.Y.Z… sets the reported version |
image <image-ref>
Lint a built source image against the image contract.
| Flag | Type | Default | Description |
|---|---|---|---|
--all-platforms | bool | false | check every platform of a multi-platform image |
--allow-warning | stringSlice | [] | downgrade this check ID’s warnings to info (repeatable); errors cannot be allowed |
--contract | string | v1alpha1 | contract version |
--insecure | bool | false | allow a plain-HTTP registry |
--json | bool | false | print JSON instead of text |
--platform | string | linux/amd64 | the platform to check in a multi-platform image |
--role | string | module role: cluster, machine or machinepool (required) | |
--strict | bool | false | treat warnings as errors for the exit code |
module <module-dir>
Lint a module directory against the contract.
| Flag | Type | Default | Description |
|---|---|---|---|
--allow-warning | stringSlice | [] | downgrade this check ID’s warnings to info (repeatable); errors cannot be allowed |
--contract | string | v1alpha1 | contract version |
--json | bool | false | print JSON instead of text |
--role | string | module role: cluster, machine or machinepool (required) | |
--strict | bool | false | treat warnings as errors for the exit code |
version
Print the version and supported contract versions.
| Flag | Type | Default | Description |
|---|---|---|---|
--json | bool | false | print JSON instead of text |
Exit codes
| Code | Name | Meaning |
|---|---|---|
0 | ExitOK | no errors (and, with –strict, no warnings). |
1 | ExitFindings | at least one error, or a warning under –strict. |
2 | ExitUnparsable | the module could not be read or parsed, or the image could not be pulled or extracted. |
3 | ExitUsage | bad command line. |
Checks
| ID | Severity | Roles | Description |
|---|---|---|---|
input/default | warning | cluster, machine, machinepool | a contract input the controller always sets to a non-null value nonetheless carries a default, which would mask a controller mistake. |
input/required | error | cluster, machine, machinepool | a contract input is not declared as a variable. |
input/reserved | error | cluster, machine, machinepool | a variable uses the reserved captf_ prefix but is not itself a contract input. |
input/sensitive | warning | cluster, machine, machinepool | bootstrap_data is declared but not sensitive = true, though it carries the bootstrap payload. |
input/tags-declared | error | cluster, machine, machinepool | the module does not declare captf_tags, the mandatory common input every module must accept. |
input/tags-unused | warning | cluster, machine, machinepool | captf_tags is declared but never referenced, directly or forwarded into a nested local module that references it. |
input/type | error | cluster, machine, machinepool | a contract input’s declared type does not accept what the generated root passes, is missing, or could not be read. |
input/user-variable-default | warning | cluster, machine, machinepool | a variable outside the contract (a user variable) has no default, so an object that does not set it in spec.variables or variablesFrom fails to apply. |
module/backend | error | cluster, machine, machinepool | the root or a nested module declares a terraform { backend } block; the generated root owns the backend. |
module/cloud | error | cluster, machine, machinepool | the root or a nested module declares a terraform { cloud } block; the generated root owns it, the same as module/backend. |
module/provider-config | warning | cluster, machine, machinepool | the root or a nested module declares its own provider configuration; the generated root owns provider configuration. |
module/tofu-shadow | warning | cluster, machine, machinepool | a .tofu file shadows a .tf file, and the declarations OpenTofu loads from it differ from Terraform’s. |
module/version | info | cluster, machine, machinepool | informational; reports the module’s declared required_version constraint, or its absence. |
output/endpoint-never-set | warning | cluster | the control_plane_endpoint output is a literal null, so a KubeadmControlPlane cluster with no user-set endpoint would wait forever. |
output/health | error | cluster, machine, machinepool | the health output is declared with the wrong shape for the contract’s health check. |
output/provider-id-list-shape | error (the output is missing or not a list), warning (the expression is not sorted and deduplicated) | machinepool | the machinepool role’s provider_id_list output is missing, or its expression does not look like it forwards one ID per instance. |
output/required | error | cluster, machine, machinepool | a contract output is not declared. |
output/reserved | warning | cluster, machine, machinepool | an output uses a name reserved for a future contract output. |
pool/autoscaling-ignore-changes | warning | machinepool | the module uses var.autoscaling but no resource ignores changes to its desired capacity, so every apply resets the cloud autoscaler’s decision. |
Job Environment
Every Terraform/OpenTofu operation runs as a single-container Kubernetes Job the manager builds from internal/jobs. This page documents that Job’s environment, mounts and fixed fields; the runner’s own environment shaping; and the reserved variable names spec.jobs.env may not set.
Main container environment
| Name | Value | Meaning |
|---|---|---|
TF_IN_AUTOMATION | 1 | Tells Terraform/OpenTofu it is running unattended: it skips interactive follow-up hints in its output. |
TF_INPUT | 0 | Disables interactive prompts; the runner always answers non-interactively. |
HOME | /captf/work | The runner’s working directory (render.WorkDir), since the image’s real HOME may not be writable. |
TMPDIR | /tmp | The Job’s /tmp emptyDir, since the image’s root filesystem is read-only. |
KUBE_NAMESPACE | (the object's namespace) | The kubernetes backend’s namespace, so state Secrets land beside the owning object. |
CHECKPOINT_DISABLE | 1 | Stops Terraform from calling checkpoint-api.hashicorp.com on every command: a pod holding cloud credentials otherwise makes that call and can stall on its timeout when the namespace drops egress silently. |
| (envFrom) | (the identity credential mirror Secret) | Mirrors the resolved TerraformClusterIdentity’s provider credentials into the runner’s environment (internal/identity). |
spec.jobs.env entries that reuse a reserved name are dropped and reported as an event; a representative Spec with TF_LOG and KUBE_CONFIG_PATH set drops: TF_LOG, KUBE_CONFIG_PATH.
Volumes and mounts
| Mount path | Read-only | Meaning |
|---|---|---|
/captf/bin | true | The runner binary, copied in by the init container. |
/captf/work | false | Scratch: the generated root, the CLI configuration, plan files, and TF_DATA_DIR. |
/tmp | false | General temporary storage (TMPDIR). |
/captf/config | true | The per-run Secret: the generated inputs root, and for a restore Job, the backup’s state chunks (jobs.RestoreChunkDir). |
/var/run/captf/credentials | true | The identity Secret’s credential files, mode 0440: a non-root image user reads them through the pod’s fsGroup. |
Fixed Job fields
| Field | Value | Meaning |
|---|---|---|
| backoffLimit | 0 | The controller owns retries (the attempt number is in the Job name); the pod itself never retries. |
| terminationGracePeriodSeconds | 600 | SIGTERM lets the runner finish in-flight provider calls and write results before SIGKILL. |
| activeDeadlineSeconds (default) | 3600 | Overridable by spec.jobs.activeDeadlineSeconds. |
| lockTimeoutSeconds (default) | 300 | Overridable by spec.jobs.lockTimeoutSeconds. |
Default resources
| Container | Kind | Values |
|---|---|---|
| init (runner copy) | requests | cpu=10m, memory=32Mi |
| init (runner copy) | limits | cpu=100m, memory=64Mi |
| main (source), when spec.jobs.resources is unset | requests | cpu=250m, memory=512Mi |
| main (source), when spec.jobs.resources is unset | limits | memory=2Gi (no default CPU limit: throttling a slow apply is worse than a slow apply) |
Security contexts
| Scope | Field | Value |
|---|---|---|
| pod | seccompProfile | RuntimeDefault |
| pod | fsGroup | 65532 (so a non-root image user can read the 0440 credential files through the group) |
| init (runner copy) | allowPrivilegeEscalation | false |
| init (runner copy) | capabilities.drop | ALL |
| init (runner copy) | runAsNonRoot / runAsUser | true / 65532 |
| init (runner copy) | readOnlyRootFilesystem | true |
| main (source) | allowPrivilegeEscalation | false |
| main (source) | capabilities.drop | ALL |
| main (source) | readOnlyRootFilesystem | true |
| main (source) | runAsNonRoot | not defaulted: images built FROM hashicorp/terraform run as root |
Runner command and args
The main container’s command is /captf/bin/runner run. Its args, built by jobs.Build, are:
| Flag | Value | Meaning |
|---|---|---|
--op | apply, destroy, refresh, drift, restore or plan | The operation this Job runs. |
--bin | The runtime binary path (render.RuntimePath): the image’s tofu or terraform. | |
--image | (the source image reference) | Recorded in events and logs to identify which image ran. |
--module | The image’s role module (render.ModuleDir). | |
--providers | The image’s optional provider mirror (render.ProvidersDir); the runner checks whether it exists. | |
--workdir | The generated root’s parent (render.WorkDir). | |
--config | The per-run Secret’s mount (jobs.ConfigDir). | |
--lock-timeout | (jobs.DefaultLockTimeoutSeconds, or spec.jobs.lockTimeoutSeconds)s | How long the backend lock acquisition waits before failing. |
--stop-timeout | (terminationGracePeriodSeconds minus a margin)s | How long the runner has, after SIGTERM, to finish in-flight work and write results before the pod is killed. |
--backend-config=secret_suffix | (a per-attempt suffix) | The kubernetes backend’s Secret name suffix. |
--backend-config=namespace | (the object’s namespace) | The kubernetes backend’s namespace. |
--backend-config=in_cluster_config | Tells the kubernetes backend to use the pod’s in-cluster credentials. | |
--backend-config=labels | (the backend Secret labels, as an HCL object) | Labels the backend applies to the state Secrets it manages. |
--force-unlock | (a stale lock ID) | Set only when the controller detected a stale backend lock; the runner force-unlocks it after init. |
--guard-deletes | Set only for a TerraformCluster apply: the runner stops before a plan that deletes or replaces resources unless allowed. | |
--inputs-hash | (the rendered inputs hash) | Paired with –guard-deletes: the hash of the inputs this apply renders. |
--allow-deletes-hash | (an approved inputs hash) | The TerraformCluster’s captf.io/approve-destructive-plan hash, when set: allows a destructive plan for that exact input set. |
--expect-plan | (an approved plan hash) | Under applyPolicy Manual, the TerraformCluster’s captf.io/approve-plan hash the apply must re-plan and match before applying. |
--event-object | (apiVersion/kind/namespace/name/uid of the owner) | Set only when the manager runs with –runner-events: lets the runner report progress as Kubernetes events about the owning object. |
--job-name | (the Job’s name) | Paired with –event-object: the reporting Job’s own name. |
--restore-chunks | (chunk count) | Set only for a restore Job: how many backup state chunks are projected into the config volume. |
--restore-resources | (managed resource count) | Set only for a restore Job: the backup’s recorded managed-resource count; the runner fails the restore if the state list shows none when this is not 0. |
What the runner changes before running the runtime
internal/runner.Prepare calls the exported Environ helper to build the runtime’s step environment from the Job’s. Given an environment carrying every credential-bearing and backend-changing variable the identity Secret’s envFrom or the image’s own ENV might set:
TF_IN_AUTOMATION=1 TF_INPUT=0 KUBE_NAMESPACE=default CHECKPOINT_DISABLE=1 TF_LOG=DEBUG TF_VAR_region=us-east-1 TF_WORKSPACE=default TF_CLI_ARGS_apply=-parallelism=1 KUBE_CONFIG_PATH=/tmp/kubeconfig HOME=/root TMPDIR=/tmp AWS_ACCESS_KEY_ID=AKIA...
Environ produces (TMPDIR was set, so it is kept as-is):
AWS_ACCESS_KEY_ID=AKIA... CHECKPOINT_DISABLE=1 HOME=/captf/work KUBE_NAMESPACE=default TF_DATA_DIR=/captf/work/.terraform TF_INPUT=0 TF_IN_AUTOMATION=1 TMPDIR=/tmp
and reports the dropped names: KUBE_CONFIG_PATH, TF_CLI_ARGS_apply, TF_LOG, TF_VAR_region, TF_WORKSPACE.
Rules:
TF_DATA_DIRis always forced to<workDir>/.terraform, overriding any inherited value.HOMEis always forced to<workDir>.TMPDIRis kept when the Job’s environment already sets it (the Job sets it to/tmp, an emptyDir); otherwise it defaults to<workDir>/tmp, whichPreparecreates.TF_CLI_CONFIG_FILEis set only when the image ships a provider mirror (Preparedetects this by statting the image’s providers directory;Environitself never inspects the filesystem).- Every other
TF_*andKUBE_*variable is dropped, exceptTF_IN_AUTOMATION,TF_INPUTandKUBE_NAMESPACE, which the Job already sets and which never carry identity-Secret or image values because a container’s explicitenvwins overenvFrom.
spec.jobs.env rejected names
spec.jobs.env rejects (drops, with an event) any name starting with TF_ or KUBE_: the runner and the Job above own that namespace (internal/jobs.reservedEnv).
See the image contract for the image side of this: the fixed paths the runner expects, and what the image must not ship.
Annotations, Labels and Finalizers
Every annotation, label and finalizer key CAPTF sets or reads, generated from internal/docsgen’s registry, which references the real Go constants so a rename cannot leave this page stale.
User-facing keys
Keys the operator sets or reads.
| Key | Kind | On | Set by | Read by | Meaning |
|---|---|---|---|---|---|
captf.io/approve-destructive-plan | annotation | TerraformCluster, TerraformMachine, TerraformMachinePool | the operator | internal/controllers/shared (the reconcile loop’s destructive-plan guard) | Approves one destructive apply or drift remediation by naming the inputs hash it must render; the controller removes the annotation once an apply that used it succeeded. |
captf.io/approve-plan | annotation | TerraformCluster (applyPolicy Manual) | the operator | internal/controllers/shared (plan-preview approval check) | Approves one plan by naming its hash (status.plan.planHash); the apply runs only if it plans exactly the same changes again. The controller removes the annotation once the apply succeeded. |
captf.io/restore-state | annotation | TerraformCluster, TerraformMachine, TerraformMachinePool | the operator | internal/controllers/shared (the restore path) | Requests a state restore by naming the serial of a backup in status.stateBackups; the controller removes the annotation once the restore Job succeeded. Deletion wins over a pending restore. |
captf.io/variables | label | a ConfigMap or Secret named by spec.variablesFrom | the operator, on their own ConfigMap or Secret | internal/controllers/shared (variables resolution) and the manager’s cache/watch selector | Opts a ConfigMap or Secret in as a variablesFrom source; CAPTF only reads objects that carry it. It is a standing opt-in, never removed by the controller. |
captf.io/runner | label | a ServiceAccount named by spec.jobs.serviceAccountName | the operator, on their own ServiceAccount | internal/rbac (the opt-in check gating whether a Job is created with it) | Opts a custom runner ServiceAccount in, besides the default captf-runner. Without it the controller reports ServiceAccountNotOptedIn and creates no Job. |
Internal keys
Keys CAPTF sets and reads itself; do not edit them.
| Key | Kind | On | Set by | Read by | Meaning |
|---|---|---|---|---|---|
captf.io/managed | label | every object CAPTF owns (state Secrets, Leases, mirrors, run Jobs, the default runner ServiceAccount) | the controller | the manager’s cache and sweep selectors (internal/manager/cache.go) | Selects everything CAPTF owns. |
captf.infrastructure.cluster.x-k8s.io/owner-kind | label | state Secrets and Leases | internal/state.BackendLabels | internal/state.Selector and internal/runlease | The owning TerraformCluster/Machine/MachinePool’s kind, for backend and lease lookups. |
captf.infrastructure.cluster.x-k8s.io/owner-name | label | state Secrets and Leases | internal/state.BackendLabels | internal/state.Selector and internal/runlease | The owning object’s name, for backend and lease lookups. |
captf.io/inputs-hash | annotation | the base state Secret (and chunk 0 of a state backup) | internal/state (adopt) and internal/jobs.Build | internal/state and internal/controllers/shared (blocked-apply and approval messages) | The inputs hash of the state currently adopted. It is an annotation, not a label. |
captf.io/state-backup | label | a state backup chunk Secret | internal/state.TakeBackup | internal/state.BackupSelector | Marks a Secret as a state backup chunk. |
captf.io/state-backup-suffix | label | a state backup chunk Secret | internal/state.TakeBackup | internal/state.BackupSelector | The backend Secret suffix the backup was taken from. |
captf.io/state-backup-serial | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (ListBackups, FindBackup) | The backup’s state serial. |
captf.io/state-backup-lineage | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (ListBackups, FindBackup) | The backend state’s lineage ID at backup time. |
captf.io/state-backup-taken-at | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (ListBackups, FindBackup) | When the backup was taken. |
captf.io/state-backup-source-job | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (ListBackups, FindBackup) | The Job whose apply produced the backed-up state. |
captf.io/state-backup-digest | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (conflict detection, FindBackup) | A digest of the backed-up state, used to detect a concurrent conflicting backup. |
captf.io/state-backup-resources | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (FindBackup) | The backup’s managed resource count. |
captf.io/state-backup-set | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state.ListBackups | Groups a backup’s chunk Secrets into one backup. |
captf.io/state-backup-chunk | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (FindBackup) | The chunk’s index within its backup set. |
captf.io/state-backup-chunks | annotation | a state backup chunk Secret | internal/state.TakeBackup | internal/state (FindBackup) | The backup set’s total chunk count. |
captf.io/lease | label | a coordination.k8s.io Lease | internal/runlease.Acquire | internal/runlease | Distinguishes a CAPTF run or cluster lease from the backend’s own lock Lease. |
captf.io/lease-op | annotation | a coordination.k8s.io Lease | internal/runlease.Acquire | internal/runlease | The operation the lease holder is running, for diagnostics. |
captf.io/lease-acquired-at | annotation | a coordination.k8s.io Lease | internal/runlease.Acquire | internal/runlease.AcquiredAt | When the lease was acquired. |
captf.io/mirrored | label | the identity credential mirror Secret | internal/identity.EnsureMirror | internal/identity (conflict detection, Revoke) | Marks a Secret as an identity credential mirror. |
captf.io/source-hash | annotation | the identity credential mirror Secret | internal/identity.EnsureMirror | internal/identity.EnsureMirror | A hash of the source TerraformClusterIdentity’s credentials, so a change is detected and the mirror rewritten. |
captf.io/bookkept | annotation | a run Job | internal/controllers/shared.MarkBookkept | internal/controllers/shared (collectFinished) | Marks a finished Job as already accounted for in status, so it is not double-counted. |
captf.io/interrupted | annotation | a run Job | internal/controllers/shared.MarkBookkept | internal/controllers/shared (collectFinished, countFailures) | Marks a Job that stopped without a clean result (for example, evicted mid-run). |
captf.io/drift-remediation | annotation | a run Job | internal/controllers/shared (Job creation on the drift-remediation path) | internal/controllers/shared.applyDestroy | Marks a Job as a drift-remediation apply, distinct from an ordinary apply. |
captf.io/destructive-plan-blocked | annotation | a run Job | internal/controllers/shared.MarkBookkept | internal/controllers/shared (collectFinished, countFailures, DecideOp) | Marks an apply that stopped because its plan was destructive and unapproved. |
captf.io/plan-changed | annotation | a run Job | internal/controllers/shared.MarkBookkept | internal/controllers/shared (collectFinished, countFailures) | Marks an apply that stopped because a re-plan under Manual applyPolicy no longer matched the approved plan. |
captf.io/plan-unreadable | annotation | a run Job | internal/controllers/shared.MarkBookkept | internal/controllers/shared (collectFinished) | Marks a Job whose plan result could not be parsed. |
captf.io/approved-plan | annotation | a run Job | internal/controllers/shared (apply-Job creation under Manual applyPolicy) | internal/controllers/shared (plan-approval comparison) | Records the plan hash an apply Job was created to satisfy. |
captf.io/restore-serial | annotation | a restore Job | internal/jobs.Build | internal/controllers/shared (the restore path) | The state backup serial the restore Job pushes. |
captf.infrastructure.cluster.x-k8s.io/op | label | a run Job and its pod | internal/jobs.Labels | internal/jobs and internal/controllers/shared (listing/filtering Jobs by operation) | The operation the Job runs (apply, destroy, refresh, drift, restore or plan). |
captf.infrastructure.cluster.x-k8s.io/attempt | label | a run Job and its pod | internal/jobs.Labels | internal/jobs and internal/controllers/shared (retry/attempt tracking) | The Job’s attempt number. |
captf.io/endpoint-source | annotation | TerraformCluster | internal/controllers/terraformcluster (EndpointInput, ModuleEndpoint) | internal/controllers/terraformcluster | Records whether the control-plane endpoint came from the user or the module; written at most once. |
captf.io/remediation-requested | annotation | Machine (the CAPI object, not TerraformMachine) | internal/controllers/terraformmachine.patchRemediation | internal/controllers/terraformmachine.patchRemediation | Marks that CAPTF itself set clusterv1.RemediateMachineAnnotation, so it only ever clears an annotation it set. |
captf.io/image | annotation | the durable inputs Secret | internal/inputs | internal/inputs | Records spec.source.image as last written, for digest pinning. |
captf.io/image-digest | annotation | the durable inputs Secret | internal/inputs | internal/inputs | Records the resolved image digest, so a floating tag is pinned across reconciles. |
captf.io/identity | annotation | the durable inputs Secret and the identity credential mirror Secret | internal/inputs and internal/identity | internal/inputs and internal/identity | Names the TerraformClusterIdentity the credentials came from. |
Finalizers
TerraformClusterIdentity has no finalizer: nothing external depends on it directly, so its controller deletes cleanly without one.
| Key | Kind | On | Set by | Read by | Meaning |
|---|---|---|---|---|---|
terraformcluster.infrastructure.cluster.x-k8s.io | finalizer | TerraformCluster | internal/controllers/terraformcluster | Kubernetes garbage collection | Blocks deletion until the controller has torn down the cluster’s Terraform-managed resources. |
terraformmachine.infrastructure.cluster.x-k8s.io | finalizer | TerraformMachine | internal/controllers/terraformmachine | Kubernetes garbage collection | Blocks deletion until the controller has torn down the machine’s Terraform-managed resources. |
terraformmachinepool.infrastructure.cluster.x-k8s.io | finalizer | TerraformMachinePool | internal/controllers/terraformmachinepool | Kubernetes garbage collection | Blocks deletion until the controller has torn down the pool’s Terraform-managed resources. |
Cluster API and clusterctl keys
Keys owned by Cluster API or clusterctl that CAPTF reads or writes, imported as their real constants rather than retyped.
| Key | Kind | On | Set by | Read by | Meaning |
|---|---|---|---|---|---|
cluster.x-k8s.io/cluster-name | label | TerraformCluster, TerraformMachine, TerraformMachinePool, run Jobs, state Secrets, Leases | Cluster API | internal/state, internal/jobs, internal/runlease and internal/controllers/shared, to scope objects to their owning Cluster | The owning Cluster’s name. |
cluster.x-k8s.io/remediate-machine | annotation | Machine (the CAPI object) | internal/controllers/terraformmachine.patchRemediation (guarded by RequestedByAnnotation) | Cluster API’s machine health check / remediation | Requests Cluster API remediate (replace) the Machine. |
cluster.x-k8s.io/replicas-managed-by | annotation | TerraformMachinePool | internal/controllers/terraformmachinepool.replicas (value “captf”) | the cluster-autoscaler and Cluster API | Tells Cluster API and the autoscaler that CAPTF, not the autoscaler’s default path, owns spec.replicas. |
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size | annotation | TerraformMachinePool | the operator (autoscaler contract) | internal/controllers/terraformmachinepool.ParseAutoscaling | The pool’s minimum replica count under autoscaling. |
cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size | annotation | TerraformMachinePool | the operator (autoscaler contract) | internal/controllers/terraformmachinepool.ParseAutoscaling | The pool’s maximum replica count under autoscaling. |
clusterctl.cluster.x-k8s.io/move | label | state Secrets and Leases | internal/state.BackendLabels | clusterctl move | Includes state Secrets and Leases in a clusterctl move. |
clusterctl.cluster.x-k8s.io/delete-for-move | annotation | TerraformMachine | clusterctl move | internal/webhooks (ValidateDelete, recognizes a move-driven delete when the Cluster is paused) | Marks a delete that clusterctl move issues as part of moving the Cluster, not an ordinary delete. |
clusterctl.cluster.x-k8s.io/move-hierarchy | label | the TerraformClusterIdentity CRD itself | config/crd/patches/move-hierarchy_terraformclusteridentities.yaml (a static kustomize patch) | clusterctl move | Tells clusterctl move to bring along objects that reference a TerraformClusterIdentity, not just the identity object. |
captf_tags keys
The fixed keys internal/contract.Tags renders into every module’s captf_tags input variable (a Terraform tag map, not Kubernetes object metadata).
| Key | Kind | On | Set by | Read by | Meaning |
|---|---|---|---|---|---|
captf.io/cluster | tag | captf_tags (every module) | internal/contract.Tags | the module | The owning Cluster’s name. |
captf.io/namespace | tag | captf_tags (every module) | internal/contract.Tags | the module | The owning object’s namespace. |
captf.io/kind | tag | captf_tags (every module) | internal/contract.Tags | the module | The owning object’s kind (TerraformCluster, TerraformMachine or TerraformMachinePool). |
captf.io/name | tag | captf_tags (every module) | internal/contract.Tags | the module | The owning object’s name. |
captf.io/managed-by | tag | captf_tags (every module) | internal/contract.Tags | the module | Always captf: marks infrastructure CAPTF manages. |
captf.io/template | tag | captf_tags (every module) | internal/contract.Tags | the module | The object’s cluster.x-k8s.io/cloned-from-name annotation, or empty when absent. |
clusterctl Variables
clusterctl substitutes these shell-style variables when generating a cluster from templates/*.yaml (see Quick Start); the second table lists the ClusterClass topology variables of the clusterclass flavor.
| Variable | Required | Default | Used in | Description |
|---|---|---|---|---|
CLUSTER_NAME | yes | - | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Name of every object. |
CONTROL_PLANE_MACHINE_COUNT | yes | - | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Control-plane replicas: an odd number of at least 3 keeps a maxSurge: 0 rollout available. |
KUBERNETES_VERSION | yes | - | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | KubeadmControlPlane and MachineDeployment version. |
NAMESPACE | yes | - | identity-libvirt.yaml, identity.yaml | The namespace allowed to use the identity; used by identity.yaml and identity-libvirt.yaml. |
POD_CIDR | no | 192.168.0.0/16 | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Cluster.spec.clusterNetwork.pods. |
SERVICE_CIDR | no | 10.128.0.0/12 | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Cluster.spec.clusterNetwork.services. |
TERRAFORM_CLUSTER_IMAGE | yes | cluster-template-libvirt.yaml: ghcr.io/captf-io/cluster-api-provider-terraform/libvirt-cluster:v0.1.0-opentofu | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | TerraformCluster.spec.source.image. |
TERRAFORM_CP_NODE_STARTUP_TIMEOUT | no | 1800 | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Control-plane MachineHealthCheck nodeStartupTimeoutSeconds; longer than the worker default, since kubeadm init/join plus kube-vip does strictly more work than a worker join. |
TERRAFORM_CP_UNHEALTHY_TIMEOUT | no | 3600 | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Control-plane MachineHealthCheck timeout; longer, since replacing a control-plane machine is costlier. |
TERRAFORM_IDENTITY_NAME | yes | cluster-template-libvirt.yaml: libvirt | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml, identity.yaml | The identity the cluster, and by default its machines, use. |
TERRAFORM_MACHINE_IMAGE | yes | cluster-template-libvirt.yaml: ghcr.io/captf-io/cluster-api-provider-terraform/libvirt-machine:v0.1.0-opentofu | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | TerraformMachineTemplate.spec.template.spec.source.image (control plane and workers). |
TERRAFORM_NODE_STARTUP_TIMEOUT | no | 1200 | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Worker MachineHealthCheck nodeStartupTimeoutSeconds. |
TERRAFORM_UNHEALTHY_TIMEOUT | no | 1800 | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | Worker MachineHealthCheck: seconds of InfrastructureReady=False before remediation. It must exceed apply time plus one drift interval (30m by default), because the window starts when the apply starts. |
TERRAFORM_VIP | yes | - | cluster-template-libvirt.yaml | The control-plane VIP kube-vip announces; must be in 192.168.150.200-220 (cluster-template-libvirt.yaml). |
WORKER_MACHINE_COUNT | yes | - | cluster-template-clusterclass.yaml, cluster-template-libvirt.yaml, cluster-template.yaml | MachineDeployment replicas. |
ClusterClass topology variables
| Variable | Required | Type | Description |
|---|---|---|---|
identityName | yes | string | Name of the TerraformClusterIdentity the cluster and, by default, its machines use. |
clusterImage | yes | string | Source image of the cluster module (TerraformCluster spec.source.image). |
machineImage | yes | string | Source image of the machine module, for control-plane and worker machines. Changing it rolls the machines out. |
Make Targets
make help prints this same list, grouped and described from the Makefile’s own ##@ and ## comments.
General
| Target | Description |
|---|---|
help | Display this help. |
Development
| Target | Description |
|---|---|
generate | Generate deepcopy code (controller-gen object). |
manifests | Generate CRD, RBAC and webhook manifests into config/. |
fmt | Format Go code (gofmt -s, goimports -local). |
vet | Run go vet in every module. |
lint | Run golangci-lint (.golangci.yml) in every module, then kube-api-linter on api/. |
lint-modules | Run tfcapi-lint module –strict over the noop modules and variants and the libvirt modules. |
verify-noop-modules | Stage every no-op module and variant; init, validate (terraform, tofu) and lint each. |
lint-api | Run kube-api-linter (.golangci-kal.yml) on the api/ module. |
lint-fix | Run golangci-lint and kube-api-linter with auto-fixers. |
Test
| Target | Description |
|---|---|
test | Run unit tests in every module (-race -count=1). |
Build
| Target | Description |
|---|---|
build | Build every cmd/* binary into bin/. |
manager | Build the static manager into bin/manager. |
run | Run the manager out of cluster against the current kubeconfig (dev). |
runner | Build the static Job runner into bin/runner and check it is static. |
docker-build | Build the manager/runner image $(IMG) for the host platform. |
docker-buildx | Build a multi-arch manifest list $(IMG) for $(PLATFORMS) (podman). |
docker-push | Push $(IMG); a manifest list from docker-buildx is pushed with all its images. |
noop-images | Build the no-op source images captf-test/noop-<name>:{terraform,opentofu} (NOOP=, ROLE=, VARIANT=). |
libvirt-images | Build captf-test/libvirt-<role>:opentofu for each module under modules/libvirt. |
test-libvirt-modules | Run each libvirt module’s tests (tofu test, mocked provider). |
libvirt-images-push | Push the libvirt module images as $(NOOP_REGISTRY)/libvirt-<role>:$(VERSION)-opentofu. |
noop-images-push | Push the no-op images as $(NOOP_REGISTRY)/noop-<name>:$(VERSION)-<base>. |
lint-images | Build the negative image fixtures and tfcapi-lint them and the noop images (needs podman and jq). |
release-lint-snapshot | Build tfcapi-lint release assets and checksums into dist/ (goreleaser snapshot). |
release-lint-binaries | Alias of release-lint-snapshot. |
release-lint | Build tfcapi-lint release assets from the current git tag. |
manifests-release | Build out/infrastructure-components.yaml (RELEASE_IMG); copy metadata.yaml and templates/*.yaml. |
release-preflight | Check the tree is clean, HEAD carries tag $(VERSION), and metadata.yaml is append-only. |
release | Build and push the images, push the noop images, and build every asset into out/release (VERSION=vX.Y.Z). |
release-image-digest | Print the registry digest of $(RELEASE_REPO):$(VERSION). |
release-assets | Build every release asset for VERSION into out/release (needs the git tag). |
release-notes | Write out/release/notes.md from the commits since the previous tag. |
release-github | Create the GitHub release for VERSION from out/release (publishes: operator only). |
Documentation
| Target | Description |
|---|---|
require-docs-dir | Fail unless DOCS_DIR names a captf-io/docs checkout (a prerequisite of the DOCS_DIR targets). |
docs-gen | Regenerate the book’s generated pages (reference/, contract schemas, noop module copy) in DOCS_DIR. |
docs-api | Regenerate the book’s reference/api.md in DOCS_DIR from api/v1alpha1 (hack/crd-ref-docs). |
verify-docs-api | Check that the book’s reference/api.md in DOCS_DIR matches what docs-api would generate. |
verify-docs | Check the book in DOCS_DIR against this tree: generated pages, schemas, noop module copy, spec tables, quick start, example Containerfiles, captf.io/docs URLs. |
lint-docs-md | Lint Markdown docs with markdownlint-cli2 (part of verify). |
lint-docs-prose | Lint doc prose with vale against the house style (part of verify). |
Verify
| Target | Description |
|---|---|
verify | Run all verifications. |
verify-godoc | Check that every declaration, parameter, return value and package is documented (hack/godoccheck). |
verify-local-repository | Generate the provider and both flavors from a clusterctl local repository of the release assets (offline, no cluster). |
verify-tilt-provider | Check tilt-provider.yaml against the tree and config/default. |
docs-check | Check relative links in README.md/templates/README.md and doc URLs repo-wide (DOCS_DIR checks them against a captf-io/docs checkout). |
verify-templates | Render templates/ with the pinned clusterctl (hack/verify-templates.sh). |
verify-containerfiles | Check the noop and example Containerfiles (structure, hadolint; VERIFY_CONTAINERFILES_BUILD=1 also builds). |
check-licenses | Check that no MPL-2.0 dependency of tfcapi-lint applies Exhibit B (hack/check-licenses.sh). |
verify-components | Check the clusterctl components built from config/default. |
verify-metadata | Validate metadata.yaml and check releaseSeries is append-only against the previous tag. |
verify-modules | Check go.work/go.mod pins and the no-replace rule. |
verify-schemas | Validate the contract JSON Schemas and their examples (python3 + jsonschema). |
verify-gen | Check generated files are up to date (needs git). |
promtool-check | Check the alert rules with promtool and build the Prometheus component on config/default. |
promtool-test | Unit-test the alert rules (config/prometheus/tests/rules_test.yaml). |
Tools
| Target | Description |
|---|---|
tools | Install every pinned tool into hack/tools/bin. |
kubebuilder | Install kubebuilder (scaffolding only) into hack/tools/bin. |
Cleanup
| Target | Description |
|---|---|
clean | Remove build output (bin/, dist/). |
Glossary
Short definitions of CAPTF terms used across this book, and the Cluster API and Terraform or OpenTofu terms CAPTF’s own docs assume. Each entry links to the page that covers the term in full; this page never repeats what that page already says.
- apply — The Job operation that renders an object’s current inputs, records them as its durable inputs, then plans and applies them. The destructive-plan guard stops it after the plan step; see choosing the next operation.
- applyPolicy — A
TerraformClusterfield,Automatic(the default) orManual, that decides whether every apply first waits for a reviewed plan. See Plan Approval. - bookkeeping — The reconcile step that reads an object’s finished
Jobs, records each one’s result in
status.lastRun, pins a succeeded apply’s image digest, and releases the Job’s leases. See the reconcile lifecycle. - ClusterClass — Cluster API’s reusable cluster template mechanism.
CAPTF ships a
noopClusterClass, which expands aTerraformClusterTemplateand twoTerraformMachineTemplates. See Templates and ClusterClass. - cluster operation gate — The manager’s
--cluster-operation-gateflag (trueby default), which keeps aTerraformCluster‘s own apply or destroy from running at the same time as its machines’ and machine pools’, through a per-cluster write lease. See run leases and the cluster operation gate. - contract — The normative interface between the controller and a module: which inputs each role receives and which outputs it must produce. See Module contract.
- destroy — The Job operation that runs the module’s
destroystep against an object’s durable inputs, run only on deletion. See deletion order. - destructive plan — A plan that deletes or replaces at least one
resource. A
TerraformClusterstops before applying one until the exact inputs hash it belongs to is approved. See Plan Approval. - drift — A difference between a provisioned object’s Terraform or OpenTofu state and reality, found by a drift Job’s plan. See Drift and Health.
- durable inputs (Secret) —
captf-inputs-<kindshort>-<name>, the record of what the controller rendered when it last started an apply Job for an object. An immutable machine’s destroy, drift and refresh always use it; a mutable cluster or pool’s destroy prefers it and falls back to current inputs, while its drift and refresh prefer current inputs and fall back to it. See Job Inputs. - identity — A
TerraformClusterIdentity: names the Secret of cloud credentials aTerraformCluster,TerraformMachineorTerraformMachinePooluses, and the namespaces allowed to use it. See Identities and Credentials. - InfraCluster / InfraMachine / InfraMachinePool — Cluster API’s
generic roles for an infrastructure provider’s objects. CAPTF’s are
TerraformCluster,TerraformMachineandTerraformMachinePool. See The Kinds. - inputs hash —
captf.io/inputs-hash: covers the contract version, the role,spec.source.imageas written, the rendered inputs, and any user variables. A change to it re-applies a mutable object. See Job Inputs. - kindshort — The short kind code in a Secret, Lease or Job name:
cforTerraformCluster,mforTerraformMachine,mpforTerraformMachinePool. See Secret names and the suffix. - mirror Secret —
captf-creds-<identity>, the copy of an identity’s credentials Secret the controller mirrors into each allowed namespace where an object uses the identity, and keeps in sync, so a Job never reads the source Secret directly. See Identities and Credentials. - module — The Terraform or OpenTofu code that implements one contract role, shipped in an OCI image together with the runtime that runs it. See Image Contract.
- per-run Secret —
captf-run-<job>, a private copy of a Job’s inputs, owned by the Job, mounted at/captf/config, and deleted when the Job finishes. See Job Inputs. - plan hash — Under
applyPolicy: Manual, the hash over a plan’s sorted, non-no-op changes that acaptf.io/approve-planannotation names to approve it. See Plan Approval. - provider mirror — The optional
/captf/providersfilesystem mirror an image can ship, soinitneeds no registry egress at run time. See Runtime Environment. - refresh — The Job operation that runs
apply -refresh-only: it updates state from reality and produces a health reading, without planning or finding drift. See Drift and Health. - restore — The Job operation, requested by the
captf.io/restore-stateannotation, that pushes a listed state backup back into the backend withstate push -force. See Restore. - role — Which of
cluster,machineormachinepoola module implements, and which ofTerraformCluster,TerraformMachineorTerraformMachinePoolruns it. See Module contract. - run lease — A
coordination.k8s.io/v1Lease the controller takes on an object before starting a Job, so two reconciles never start two Jobs for the same object at once. See run leases and the cluster operation gate. - runner — The binary that is a Job’s entrypoint: it drives init and
the requested operation (plan, apply, destroy, refresh, drift or
restore), enforces the destructive-plan guard and, under
Manual, a plan’s approval, and emits step events on the object. See Runner CLI. - source image —
spec.source.image, the OCI image naming aTerraform*object’s module and runtime. It is CAPTF’s trust boundary: whoever may set it controls what the Job’s Pod does. See Security Model. - state backup —
captf-state-backup-<suffix>-<serial>, a versioned copy of an object’s state Secrets, taken whenever the controller observes a state serial it has not backed up before. See Terraform State. - workload kinds —
TerraformCluster,TerraformMachineandTerraformMachinePool: the three kinds that each run one Job-backed Terraform or OpenTofu module role and share a common spec and status shape. See The Kinds.
See also
- API Reference for every field these terms name.
- Conditions for the reasons the kinds above report.
- Module contract for the normative meaning of a role and its inputs and outputs.
Third-party licenses
CAPTF is Apache-2.0. The modules below are linked into tfcapi-lint, and
this file lists those whose license is not Apache-2.0. The full dependency
set is in go.mod; a license check over the full set is not yet
automated.
| Module | Version | License | Used for |
|---|---|---|---|
github.com/hashicorp/terraform-config-inspect | v0.0.0-20260904064934-75d64de68c31 | MPL-2.0 | tfcapi-lint reads a module’s variables and outputs without running Terraform |
github.com/hashicorp/hcl/v2 | v2.20.1 | MPL-2.0 | the second lint pass: backend/cloud blocks, provider attributes, .tofu files, expressions |
github.com/hashicorp/hcl | v0.0.0-20170504190234-a4b07c25de5f | MPL-2.0 | legacy HCL parser required by terraform-config-inspect |
github.com/zclconf/go-cty | v1.14.4 | MIT | HCL’s type and value system (the lint pass reads literal values with it) |
github.com/mitchellh/go-wordwrap | v1.0.0 | MIT | diagnostic formatting in HCL |
MPL-2.0
MPL-2.0 is file-level copyleft. An Apache-2.0 binary may link unmodified MPL-2.0 code (MPL-2.0 section 3.3, “Distribution of a Larger Work”), provided the MPL-covered source stays available: these modules are published, and CAPTF does not modify them.
The combination would not be permitted for a file carrying Exhibit B, the
“Incompatible With Secondary Licenses” notice. hack/check-licenses.sh
(part of make verify) checks the three MPL modules at their pinned
versions: the phrase must appear only in their LICENSE files, as the
license’s own definitions (sections 1.5, 3.3 and 10.4), and in no source
file. It last passed on 2026-09-26. This remains subject to license review
before a release.
Contributing
This page is for anyone changing CAPTF itself: the repository layout, the prerequisites, the build/lint/test/verify loop, the conventions the checks enforce, and running the manager under Tilt.
Repository layout
| Path | Holds |
|---|---|
api/ | The v1alpha1 Go types: its own module in go.work, so the CRD types carry no dependency on controller-runtime or the manager. |
cmd/manager, cmd/runner, cmd/tfcapi-lint | The three binaries. Each has an app package with its wiring and an app/options or equivalent for flags, so main.go stays a thin entry point. |
internal/ | Everything the binaries share, one package per concern: controllers (one subpackage per reconciled kind — TerraformClusterTemplate and TerraformMachinePoolTemplate have no controller of their own — plus shared for the common reconcile flow and sweep for the orphan RBAC sweep), jobs, runner, state, identity, rbac, runlease, inputs, render, outputs, locks, conditions, contract, hash, ownership, webhooks, lint, docsgen, feature, imageinspect, manager, metrics. |
config/ | Kustomize bases: crd, rbac, webhook, manager, certmanager, assembled by default; network-policy and prometheus are separate optional overlays, and samples holds example custom resources. |
templates/ | The clusterctl generate templates and flavors. |
modules/ | Reference Terraform/OpenTofu modules: noop (with its variants) and libvirt. |
hack/ | Build and verify tooling: pinned tool installers under hack/tools, the verify-*.sh/check-*.sh scripts make verify runs, hack/godoccheck, and the vale house style under hack/vale. |
test/fixtures | Frozen fixtures unit tests read, such as real captured Terraform/OpenTofu state. |
The book itself lives in a separate repository, captf-io/docs, published at https://captf.io/docs/; see Writing Documentation.
Each Go module (. and api) is listed in go.work. hack/verify-modules.sh
(make verify) checks that neither carries a replace directive and that
both agree on the Kubernetes and controller-runtime versions they share.
Because api only exists inside the workspace, go mod tidy run from the
repository root does not update its go.mod; add or bump one of its
dependencies by editing api/go.mod’s require block directly, then run
go mod tidy inside api/ with GOWORK=off.
Before you begin
- Go, matching the version
go.workdeclares,make,git,jqandcurl. podman(the defaultCONTAINER_TOOL) ordocker, to build images.- Node.js 22 or later with
npm, formake lint-docs-mdand so formake verify: it installsmarkdownlint-cli2from the lockfile underhack/tools/markdownlint. python3, with thejsonschemaandPyYAMLpackages, formake verify-schemas,make verify-metadata,make promtool-check,make promtool-testandmake release-assets.
Everything else — controller-gen, kustomize, golangci-lint (plus its
kube-api-linter build), clusterctl, promtool, goreleaser, goimports,
crd-ref-docs, mdbook, mdbook-mermaid, lychee and vale — is a pinned
binary make downloads for you.
The build, lint, test and verify loop
- Install the pinned tools once:
make tools. Each lands inhack/tools/binas<name>-<version>, plus an unversioned symlink; bumping a version in theMakefilere-downloads it. - After changing
api/v1alpha1or a controller’s markers, regenerate the deepcopy code and the CRD/RBAC/webhook manifests:make generate manifests.make verify-gen(part ofmake verify) fails if either is stale. - Format and lint:
make fmt lint.fmtrunsgofmt -sandgoimports -local github.com/captf-io/cluster-api-provider-terraform;lintrunsgolangci-lint(.golangci.yml) in every module, then kube-api-linter (.golangci-kal.yml) onapi/andtfcapi-lint module --stricton the reference modules undermodules/.make lint-fixreruns both linters with their auto-fixers. - Run the unit tests:
make test. See Testing for what it covers and how to run less than everything. - Before sending a change, run
make verify: every check listed in Make Targets, including that generated code and the API reference are current and thatmake lint-docs-mdandmake lint-docs-proseare clean. A change that touches a generated page needs acaptf-io/docscheckout too: see Writing Documentation. make buildcompiles everycmd/*binary tobin/; never to the repository root.make runbuilds and runs the manager out of cluster against your currentkubeconfig, for a quick check against a real API server.
The full target list, grouped the same way, is in Make Targets.
Conventions
- Commits: an imperative subject in
<subsystem>: <summary>form, such asdocsgen: render event reasons as proseorrunner: name the exit codes— checkgit logfor the subsystem names already in use. - License header: every hand-written Go file starts with the Apache 2.0
header the other files in its package carry;
hack/boilerplate.go.txtis the headercontroller-genwrites onapi/v1alpha1/zz_generated.deepcopy.go. - Documentation comments:
hack/godoccheck(make verify-godoc, part ofmake verify) enforces one rule per declaration and one per package, everywhere except generated files andhack/tools:- Every function, method and type has a doc comment starting with its
name; every interface method and top-level
constorvargroup needs only a doc comment, not one starting with its name. - Every named parameter is mentioned by name in that comment. Receivers,
parameters named
_, and aTest/Benchmark/Fuzzfunction’s conventional*testing.T/*testing.B/*testing.Fparameter are exempt. - A function or method that returns anything says what it returns, using “return”, “returns”, “returned” or “reports”.
- Every non-generated, non-external-test package has a
doc.gowhose package comment is a real overview of at least 400 characters.
- Every function, method and type has a doc comment starting with its
name; every interface method and top-level
- Generated code:
zz_generated.deepcopy.goand the CRD/RBAC/webhook manifests come frommake generate manifests; the book’s reference pages come frommake docs-gen DOCS_DIR=<path to a captf-io/docs checkout>(see Writing Documentation). Never hand-edit a generated file:make verify-genfails when the deepcopy code or the manifests no longer match their source, andmake verify-docs DOCS_DIR=<path>fails when a generated reference page does.
Common changes
Each make docs-gen below needs DOCS_DIR=<path to a captf-io/docs checkout> (see Writing Documentation).
-
Adding a manager flag: also add it to the
managerFlagGroupsmap ininternal/docsgen/manager.go, naming the reference-page heading it belongs under.TestManagerFlagGroupsfails when a flag the manager registers has no entry, or an entry names a flag that no longer exists. Runmake docs-gento regeneratereference/manager-flags.md. -
Adding a condition reason: condition types, their reasons and each reason’s doc comment live in
api/v1alpha1/conditions_consts.go;ConditionReasonsreturns the full type-to-reason table, and a reason’s doc comment becomes its Meaning in the generated reference page. A reason is set frominternal/conditionswhen it applies across controllers, or frominternal/controllers/sharedwhen it belongs to one reconcile flow. Afterward, runmake generate manifests docs-gento regenerate the CRD schema, the deepcopy code andreference/conditions.md. Until you do,internal/docsgen’sTestConditionKindsComplete,TestConditionTypeOrderComplete,TestReasonMeaningandTestGoldenConditions, plusmake verify-genandmake test, fail.An event reason follows the same shape: it lives in
internal/controllers/shared/events.go, orinternal/runner’s equivalent for a runner event, with a doc comment that becomes its Meaning;TestManagerEventsCompleteorTestRunnerEventsComplete, andTestGoldenEvents, fail untilmake docs-genregeneratesreference/events.md. Atfcapi-lintcheck ID ininternal/lintis the same again: its doc comment becomes its description, andTestCheckDescriptionsCompletefails untilmake docs-genregeneratesreference/tfcapi-lint-cli.md.
Tilt
tilt-provider.yaml wires this repository into a
Cluster API checkout’s Tilt setup: add
../cluster-api-provider-terraform to that checkout’s tilt-settings.yaml
provider_repos, and terraform to enable_providers. Tilt then builds and
live-reloads only the manager binary (from cmd, api, internal,
go.mod and go.sum) and substitutes its image for
ghcr.io/captf-io/cluster-api-provider-terraform in config/default.
Tilt never rebuilds the runner image. A Job always takes its runner from
CAPTF_MANAGER_IMAGE, which Tilt does not rewrite, so it stays whatever
:dev image is already loaded in the Tilt cluster. After a change under
internal/runner or cmd/runner, run make docker-build yourself and load
the result into the cluster before a Job needs it. make verify-tilt-provider
checks tilt-provider.yaml against the tree and config/default.
See also
Testing
This page describes CAPTF’s test suite: what make test runs, what its
tiers mean here, golden files, and running less than the whole suite.
Running the tests
make test runs go test -race -count=1 ./... in every module (. and
api). That is the entire suite, and it is all unit tests: nothing it runs
creates a Kubernetes cluster, calls a real cloud API, or invokes terraform
or tofu.
- Controllers and webhooks run against the controller-runtime fake client, never a real API server; there is no envtest in this repository.
- Job execution is faked the same way: a reconciler test asserts on the
batch/v1.JobCAPTF would create, and a runner test drives the runner’s own logic directly, without a pod ever starting. internal/state,internal/outputsandinternal/locksinstead read realkubernetes-backend state:test/fixtures/stateholds state Secrets and lock Leases captured once from real Terraform and OpenTofu runs, checked in as frozen data. Regenerating them needs a capture setup that does not exist in this repository; treat the files undertest/fixtures/stateas read-only.go test ./templates/decodes every shippedtemplates/*.yamlobject strictly into its API type and checks the valueshack/verify-templates.shrenders; it also runs CAPI’s ownClusterClassadmission webhook and topology generator, applying the class patches, againstclusterclass-noop.yaml.
Test tiers, and what does not exist yet
CAPTF does not have a separate “integration” build tag or make target.
Everything that would fall under that name — a reconciler driven through
several packages at once, still against the fake client — runs as an
ordinary unit test under make test, alongside the narrower ones; for
example, internal/controllers/shared’s TestBringUpJobCount reconciles a
whole cluster bring-up against the fake client and asserts the resulting
Job count.
End-to-end — provisioning a real cluster on real infrastructure — does
not exist yet. The closest things to it are checks that run outside go test, still without a real cloud:
make test-libvirt-modulesruns eachmodules/libvirt/*module’s owntofu testagainst a mocked provider.make verify-noop-modulesstages every no-op module and variant, runsterraform/tofu initandvalidateagainst it, and lints it withtfcapi-lint module --strict; it needs those binaries onPATHand is not part ofmake verify.make verify-local-repository(part ofmake verify) runsclusterctl generate providerandgenerate clusteragainst a local repository of the release assets, offline.
A release’s smoke test against a real management cluster is done by hand; see Releasing.
Two shell scripts, hack/docs-check_test.sh and hack/check-metadata_test.sh,
test the hack/docs-check.sh and hack/check-metadata.sh checks themselves
against fixtures under hack/testdata. They are not part of make test;
make docs-check and make verify-metadata run them automatically, before
the check itself, as part of make verify.
Golden files
Several packages pin an exact rendered output as a checked-in file and compare against it on every run, rather than asserting field by field:
| Package | What it pins |
|---|---|
internal/contract | The rendered JSON of fully populated contract inputs. |
internal/jobs | The batch/v1.Job of every operation, as YAML. |
internal/outputs | Rendered output fixtures. |
internal/render | The generated root module. |
internal/docsgen, internal/metrics | The generated reference pages in a captf-io/docs checkout (see Writing Documentation). |
cmd/tfcapi-lint | Its --json output. |
Run the affected test with UPDATE_SNAPSHOTS=1 to rewrite its golden files,
then read the diff before committing it — a passing rewrite is not the same
as a correct one. internal/docsgen’s and internal/metrics’s golden
tests need DOCS_DIR (an absolute path to a captf-io/docs checkout) and
skip without it; make docs-gen DOCS_DIR=<path> does exactly this for the
generated reference pages: DOCS_DIR=<path> UPDATE_SNAPSHOTS=1 go test ./internal/docsgen/... ./internal/metrics/... -run 'Docs|Golden'.
Running one package, or one test
go test -race ./internal/jobs/...
go test -race -run TestGoldenJobs ./internal/jobs/...
UPDATE_SNAPSHOTS=1 go test ./internal/jobs/... -run TestGoldenJobs
go.work makes the api module’s tests reachable from the repository
root too: go test ./api/... needs no cd.
Coverage
There is no make target for it and no enforced threshold. Build a profile
and view it with the standard toolchain:
mkdir -p bin
go test -race -coverprofile=bin/cover.out ./...
go tool cover -html=bin/cover.out -o bin/cover.html
See also
Writing Documentation
This page is the house style for the CAPTF book, which lives in this
repository under src and is published at
https://captf.io/docs/. Every page follows it,
and make verify enforces the parts a tool can check.
Where things go
The book is split by reader and by page type. Pick the section by who reads the page, then the page type by what they need.
| Section | Reader | Page types |
|---|---|---|
| Getting Started | Anyone trying CAPTF for the first time | Tutorials: a guided path that ends in a working result |
| Concepts | Anyone who needs to understand how CAPTF works | Explanation: how and why, no step lists |
| User Guide | People creating clusters with CAPTF | How-to: one task per page, in steps |
| Module Authors | People writing Terraform or OpenTofu modules for CAPTF | The normative contract, plus how-to pages |
| Operator Guide | People installing and running the CAPTF manager | How-to pages and runbooks |
| Reference | Everyone | Lookup tables, mostly generated from code |
| Developer Guide | People changing CAPTF itself | How-to and conventions |
Every fact lives on exactly one page. Other pages link to it instead of restating it. The page that owns a topic is the one whose title names it; when two pages could own a fact, the more specific one does.
Generated pages
These pages under reference/ are generated from code and must not be edited
by hand:
api.md, bycrd-ref-docsfrom the Go types inapi/v1alpha1;conditions.md,events.md,alerts.md,manager-flags.md,runner-cli.md,tfcapi-lint-cli.md,environment.md,annotations-labels.md,clusterctl-variables.mdandmake-targets.md, byinternal/docsgen;metrics.md, byinternal/metrics.
Each starts with a “Generated by” comment. Generated prose comes from
single-line Go doc comments and usage strings, where wrapping would mean
rewrapping source text, so each generated page also carries a
<!-- markdownlint-disable MD013 --> directive, ahead of the “Generated
by” comment so it also covers that comment’s own line, that scopes the
line-length rule off for that page only; every other Markdown rule still
applies to generated pages.
These pages, and the contract schemas under
module-author/contract/v1alpha1/schemas, are generated from a checkout of
the provider repository
(cluster-api-provider-terraform),
not from anything in this repository. To change one, change the code or
the Go doc comment it comes from there, then run make docs-gen DOCS_DIR=<path to this checkout> from the provider repository to rewrite
the pages here. make verify-docs DOCS_DIR=<path to this checkout>, also
run from the provider repository, fails when a generated page, schema copy
or example is stale; this repository’s own make verify does not check
staleness, since it has no access to the provider repository’s source.
Hand-written pages link to the generated ones for field lists, flags, reasons, events, metrics and keys, and never copy those tables.
Page shape
- One H1, the page title, matching its
SUMMARY.mdentry. - A first paragraph that says what the page covers and who it is for.
- For how-to pages and runbooks: a “Before you begin” list of prerequisites, then numbered steps, then how to confirm it worked.
- A closing “See also” list when related pages exist.
- Headings in sentence case: “Rotate the credentials”, not
“Rotate The Credentials”.
SUMMARY.mdentries use title case. - Heading text stays stable once published, since links and alert
runbook_urls point at the anchors derived from it.
Voice and wording
- Address the reader as “you”. Use the present tense and the active voice.
- Say what happens, not what “should” happen. Reserve MUST, MUST NOT, SHOULD and MAY (RFC 2119, in capitals) for the normative module contract.
- American English: behavior, labeled, license, canceled.
- “CAPTF” is the project. Spell out “Cluster API Provider Terraform (CAPTF)” on the introduction page only.
- “Terraform or OpenTofu” in prose;
terraformandtofuin code for the command-line tools. “OpenTofu” is always written with a capital O and T. - Kinds use their exact names in code spans on first mention in a section:
TerraformCluster,TerraformMachine,TerraformMachinePool, their*Templatekinds, andTerraformClusterIdentity. In running prose, “machine pool” is fine. - Expand an abbreviation on first use on each page: KubeadmControlPlane (KCP).
- No references to source files, functions or line numbers on Getting Started, User Guide or Operator Guide pages. Describe the behavior. Module Author and Developer Guide pages may name source files when the reader needs them.
- No
§section references. Link to the heading instead. - No dates, review notes, “TODO”, “pending” or “planned” statements. State what is true now. If something does not exist yet, say that it does not exist.
Markdown
- Wrap prose at 80 columns. Tables, code blocks and headings are exempt.
- Fence every code block with a language:
sh,yaml,hcl,json,text. - Shell examples show commands only, without a
$prompt, and use<angle-bracket>placeholders the reader replaces. Explain each placeholder below the block. - Outside fenced code blocks, a placeholder always goes in a code span:
`<namespace>`. A bare<name>in prose or a table cell is read as an HTML tag, andmake verify-bookfails on it. - Tables use the compact style:
| a | b |with a| --- |separator row. - Link to other book pages with relative links to the
.mdfile, including the anchor when you mean a section:[drift](../concepts/drift-and-health.md#drift). - Link to files in the provider repository with a full
https://github.com/captf-io/cluster-api-provider-terraform/blob/main/...URL. Relative links must not leave this repository’ssrc. - Include real files instead of pasting them, with a path relative to the
page:
{{#include examples/Containerfile.opentofu}}on a page next to theexamples/directory. - Diagrams are Mermaid code blocks, rendered by
mdbook-mermaid.
Checks
In this repository, make verify runs:
| Command | Checks |
|---|---|
make verify-book | The book builds with no mdBook warnings |
make links | Every link and anchor in the rendered book |
make lint-md | Markdown structure and line length |
make lint-prose | The house style above, with Vale |
The provider repository checks this book’s generated content against its
own source, from a checkout of this repository passed as DOCS_DIR: make verify-docs DOCS_DIR=<path> (generated pages, contract schemas and
examples) and make docs-check (that every https://captf.io/docs/ URL
named in the provider repository’s code and Markdown resolves to a page
and heading here). Neither is part of this repository’s own make verify.
Preview the book with make serve and open http://localhost:3001.
Releasing
This page is for whoever cuts a CAPTF release: what a release consists of, the checklist, and installing the assets before or instead of publishing them.
A release is a tag vX.Y.Z (or vX.Y.Z-rc.N) on a clean main, the
manager image ghcr.io/captf-io/cluster-api-provider-terraform:vX.Y.Z, the
noop example images, and a GitHub release with the clusterctl assets and the
tfcapi-lint binaries. Tags never move and nothing is force-pushed: a bad
release candidate gets a new -rc.N.
Before you begin
- Write access to push a signed tag and create a GitHub release.
- Registry push access for
ghcr.io/captf-io/cluster-api-provider-terraform. ghauthenticated against this repository, formake release-github.skopeo, formake releaseto read the pushed image’s registry digest.
Assets
| Asset | Built by |
|---|---|
infrastructure-components.yaml | make manifests-release: config/default with the release image, and CAPTF_MANAGER_IMAGE set to the same image. |
metadata.yaml | The repository root file; hack/check-metadata.sh enforces an append-only releaseSeries. |
cluster-template.yaml, cluster-template-clusterclass.yaml, clusterclass-noop.yaml, cluster-template-libvirt.yaml, identity.yaml, identity-libvirt.yaml | templates/. |
tfcapi-lint-<os>-<arch>, tfcapi-lint-checksums.txt | GoReleaser; see tfcapi-lint. |
The libvirt flavor’s own module images (below) are not a release asset: they
are a development-host target, never published as part of make release.
libvirt module images
The libvirt flavor (templates/cluster-template-libvirt.yaml) defaults
TERRAFORM_CLUSTER_IMAGE and TERRAFORM_MACHINE_IMAGE to
ghcr.io/captf-io/cluster-api-provider-terraform/libvirt-{cluster,machine},
tagged v0.1.0-opentofu. Those images are built with make libvirt-images
and pushed with make libvirt-images-push VERSION=v0.1.0 (tags and pushes
$(NOOP_REGISTRY)/libvirt-<role>:$(VERSION)-opentofu for each module under
modules/libvirt, mirroring noop-images-push). This is a manual step, run
from the development host with registry access. It is not part of make release and has no fixed cadence tied to a provider release: repeat it
whenever modules/libvirt/* changes, using a VERSION that matches the tag
the template defaults reference (or override
TERRAFORM_CLUSTER_IMAGE/TERRAFORM_MACHINE_IMAGE at clusterctl generate
time to point at a different tag).
Checklist
-
If this release starts a new minor series, append it to
metadata.yamlreleaseSeries. Never remove or change an existing series. -
Update the contract changelog for anything that changes what a module sees or must implement, and Upgrades for anything an operator needs to do when moving to this release. Run
make docs-gen DOCS_DIR=<path to a captf-io/docs checkout>so the generated reference pages (Writing Documentation) match the code going into the release, then commit and push the result in that checkout too. -
make lint test verifyis green.verifyincludesverify-local-repository, which generates the provider and the default, clusterclass and libvirt flavors from a clusterctl local repository of the release assets, offline. -
Tag and push the tag:
git tag -s vX.Y.Z && git push origin vX.Y.Z. -
Build and push the images and build the assets:
make release VERSION=vX.Y.Zrelease-preflightrefuses a dirty tree, an untagged HEAD, a malformed version, or a metadata change that is not append-only. The assets land inout/release/. -
Smoke-test from a local repository against a real cluster. There is no automated end-to-end test; do this step by hand.
-
Publish:
make release-github VERSION=vX.Y.Zwritesout/release/notes.mdfrom the commits since the previous tag and runsgh release createwith every asset.
Installing from a local repository
To try the assets before publishing, or offline, copy out/release/* to
~/local-repository/infrastructure-terraform/vX.Y.Z/: the directory name
is the provider label infrastructure-terraform. Point a clusterctl
config’s url at the local path instead of a release URL (see
Register the provider
for the rest of the config entry):
url: file:///home/<you>/local-repository/infrastructure-terraform/vX.Y.Z/infrastructure-components.yaml
<you> is your username on this machine; a file:// URL needs the
absolute path, so ~ does not work here. Pin the version at install time:
clusterctl init --config clusterctl.yaml --infrastructure terraform:vX.Y.Z
hack/verify-local-repository.sh does the same layout in a temp directory
and runs clusterctl generate provider and generate cluster (the
default, clusterclass and libvirt flavors) against it.
See also
- Writing Documentation for the book’s own build and verification commands.
- Contributing for the general build/lint/test/verify loop.
- libvirt Development Host for building and pushing the libvirt module images from the development host.
libvirt Development Host
This page sets up a Linux host with libvirt/KVM so that a kind management
cluster running on it can drive real VMs with libvirt: the Jobs run inside
kind, reaching libvirtd over qemu+ssh:// from the kind network. It is
one-time host setup, run by hand by the operator of that host. Everything
lives under a scratch directory of your choosing, nothing on /; the
commands below use:
export CAPTF_LIBVIRT_DIR=~/captf-libvirt
mkdir -p "$CAPTF_LIBVIRT_DIR"
The modules/libvirt/* modules have not been applied against a real
hypervisor (see What the modules
create below for their only exercise today).
There is no automated acceptance check; step 6 below lists the manual
checks that stand in for one.
Before you begin
- A Linux host, separate from or the same as the one running kind, with a
user who can run commands with
sudo. virsh,podmanandcurlon that host.- The kind management cluster already up, so you can read its network’s gateway address in step 4.
1. Packages, daemon, group
sudo dnf install -y qemu-kvm libvirt virt-install guestfs-tools
sudo systemctl enable --now virtqemud.socket virtnetworkd.socket virtstoraged.socket
sudo usermod -aG libvirt "$USER"
Log out and in (or newgrp libvirt) so the group applies, then check:
virsh -c qemu:///system list --all.
2. The captf network
A NAT bridge virbr-captf on 192.168.150.0/24. DHCP serves .10–.199;
.200–.220 stay free for control-plane VIPs (kube-vip in ARP mode needs the
VIP on the same L2 bridge as the VMs).
cat > "$CAPTF_LIBVIRT_DIR/captf-net.xml" <<'EOF'
<network>
<name>captf</name>
<forward mode='nat'/>
<bridge name='virbr-captf' stp='on' delay='0'/>
<ip address='192.168.150.1' netmask='255.255.255.0'>
<dhcp>
<range start='192.168.150.10' end='192.168.150.199'/>
</dhcp>
</ip>
</network>
EOF
virsh -c qemu:///system net-define "$CAPTF_LIBVIRT_DIR/captf-net.xml"
virsh -c qemu:///system net-autostart captf
virsh -c qemu:///system net-start captf
3. Storage pool and base image
mkdir -p "$CAPTF_LIBVIRT_DIR/pool"
virsh -c qemu:///system pool-define-as captf dir --target "$CAPTF_LIBVIRT_DIR/pool"
virsh -c qemu:///system pool-autostart captf
virsh -c qemu:///system pool-start captf
curl -fL -o "$CAPTF_LIBVIRT_DIR/pool/ubuntu-24.04-cloudimg-amd64.img" \
https://cloud-images.ubuntu.com/releases/24.04/release/ubuntu-24.04-server-cloudimg-amd64.img
virsh -c qemu:///system pool-refresh captf
The qemu user must be able to traverse the path; if VMs fail to open the
image and $CAPTF_LIBVIRT_DIR is under your home directory, give it
access:
setfacl -m u:qemu:x ~ "$(dirname "$CAPTF_LIBVIRT_DIR")" "$CAPTF_LIBVIRT_DIR"
setfacl -R -m u:qemu:rwX "$CAPTF_LIBVIRT_DIR/pool"
4. The identity credential: an ssh key for qemu+ssh
A dedicated key, authorized for the user in the libvirt group:
ssh-keygen -t ed25519 -N '' -C captf-libvirt -f "$CAPTF_LIBVIRT_DIR/id_ed25519"
cat "$CAPTF_LIBVIRT_DIR/id_ed25519.pub" >> ~/.ssh/authorized_keys
The identity Secret carries these keys. Each reaches the Job as an
environment variable and as a file under /var/run/captf/credentials/ (see
how credentials reach a
Job):
| Key | Module | Value on this host | Default (unset) |
|---|---|---|---|
LIBVIRT_URI | both | qemu+ssh://$USER@<host-ip>/system?keyfile=/var/run/captf/credentials/id_ed25519&no_verify=1 (keyfile and no_verify are libvirt’s own qemu+ssh transport query parameters) | qemu:///system |
id_ed25519 | both | the private key file | (required) |
LIBVIRT_BASE_IMAGE | machine | $CAPTF_LIBVIRT_DIR/pool/ubuntu-24.04-cloudimg-amd64.img (absolute path, $CAPTF_LIBVIRT_DIR expanded), since the pool lives outside libvirt’s default image directory once you follow step 3 above | /var/lib/libvirt/images/ubuntu-24.04-cloudimg-amd64.img |
LIBVIRT_FAILURE_DOMAIN | cluster | <host>: any name identifying this hypervisor host | libvirt |
Set a key only if you diverge from its default.
<host-ip> is the host’s address as seen from the kind network: the IPv4
gateway of the podman network kind uses (the network is dual-stack; take
the IPv4 subnet’s gateway). With the cluster up:
podman network inspect kind | jq -r '.[0].subnets[] | select(.gateway | contains(":") | not) | .gateway'
5. Firewall
Allow all traffic from the kind network’s IPv4 subnet to the host:
sudo firewall-cmd --permanent --zone=trusted --add-source=<kind-subnet>
sudo firewall-cmd --reload
<kind-subnet> is the IPv4 subnet from podman network inspect kind. If
you prefer a narrower rule, open only port 22 for that source in the active
zone instead of trusting the whole subnet.
6. Checks
virsh -c qemu:///system net-list --all # captf active
virsh -c qemu:///system pool-list --all # captf active
podman run --rm --network kind docker.io/library/alpine:3.22 nc -zv <host-ip> 22
There is no automated acceptance check. The manual equivalent is a
throwaway VM that boots the cloud image with cloud-init and gets a DHCP
lease on virbr-captf, together with a kind pod running virsh -c "$LIBVIRT_URI" list.
7. Module images
templates/cluster-template-libvirt.yaml defaults TERRAFORM_CLUSTER_IMAGE
and TERRAFORM_MACHINE_IMAGE to images built from modules/libvirt/* and
pushed with make libvirt-images libvirt-images-push VERSION=v0.1.0 (see
Releasing). Rebuild and push after
any change under modules/libvirt/; override either variable at
clusterctl generate time to use a different image instead. The images
pin dmacvicar/libvirt 0.9.9 in their provider mirror.
What the modules create
modules/libvirt/cluster/ leases a control-plane VIP; modules/libvirt/machine/
boots one VM (2 vCPU, 4 GiB, a 20 GiB qcow2 overlay on the Ubuntu 24.04
cloud image) from the bootstrap payload, on the captf network.
There is no load balancer; KubeadmControlPlane runs kube-vip as a static
pod on the control-plane machines, announcing the VIP over ARP on
virbr-captf. The cluster module only chooses and leases the VIP:
- The
libvirtflavor supplies the VIP (TERRAFORM_VIPbecomesTerraformCluster.spec.controlPlaneEndpoint); the module passes it through and rejects one outside.200-.220. - Without one, the module derives
192.168.150.(200 + sha256("<namespace>/<cluster>")[0:8] mod 21): a pure function of the Cluster, known at the first plan and never changed, with no resource backing it, so no plan can replace it. - Either way, the module leases the VIP as a 1-byte volume
captf-vip-<vip>in thecaptfpool; a second cluster that hashes to the same VIP fails its apply rather than sharing the address (21 addresses, so collisions are expected past a handful of clusters — this is a development module, not a VIP allocator). Destroying the cluster frees it.
The domain is named <namespace>-<machine_name> (domain names are
host-wide, Machine names only namespace-wide), and provider_id is
libvirt:///<namespace>-<machine_name>: the KubeadmConfig templates set
the same string as the kubelet’s provider-id, so no cloud controller
manager is needed. kubernetes_version also rides in the cloud-init
NoCloud meta-data, so a KCP/MachineDeployment upgrade installs the
version the machine was actually created for rather than one clusterctl
baked in at generate time. Addresses and health come from the captf
network’s DHCP leases, re-read on every refresh (the domain create waits
for a lease with wait_for_ip): no lease reads running/unhealthy with
message no DHCP lease on network captf (a remediation candidate),
never pending. The root disk is an overlay, so destroy removes it, the
seed ISO and the domain, leaving the base image untouched.
make test-libvirt-modules runs each module’s tests/*.tftest.hcl with
tofu test against a mocked dmacvicar/libvirt provider — the only
exercise these modules get; there is no automated end-to-end test against
a real hypervisor. The modules target OpenTofu; Terraform 1.16.4 validates
both and passes the cluster tests, but its mock cannot supply the machine
tests’ lease data (a nested list attribute).
See also
- Releasing, for the
libvirtmodule images this host builds and pushes. modules/libvirt/README.mdfor the module source itself.