# Cluster API Terraform > Cluster API Provider Terraform. Your modules are the provider. CAPTF (Cluster API Provider Terraform) is a Cluster API infrastructure provider. Instead of cloud-specific Go controllers, it provisions cluster infrastructure by running Terraform or OpenTofu modules, packaged as OCI images, as Kubernetes Jobs: "your modules are the provider". How it works: Cluster API's `Cluster`, `Machine` and `MachinePool` reference CAPTF's `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool`. The CAPTF manager renders each object's inputs, runs the module image in a runner Job, keeps Terraform state in a Kubernetes Secret next to the object, and reads the module's outputs back from that state. A module is written to the `v1alpha1` module contract and linted with `tfcapi-lint`. Status: pre-release. Every kind and the module contract are `v1alpha1`; the contract is frozen for implementation but may change. There are no end-to-end tests yet, and the reference cloud modules (AWS, Google Cloud, Azure, OCI, OpenStack) have not yet been applied to a real cloud. Using this index: the docs are organized by reader. Overview and Start here for evaluators; User Guide and Cloud Modules for people running clusters; Module Authors for people writing modules (the contract is normative); Operations and Troubleshooting for people running the manager; How It Works for internals; Reference for generated API, conditions, events, metrics, flags and environment. Every page is also available as Markdown at its URL followed by `index.md`, and https://captf.io/llms-full.txt is the whole site in one file. News: https://captf.io/feed_rss_created.xml. # Start here # Your modules are the provider Cluster API Provider Terraform (CAPTF) is a Cluster API infrastructure provider that provisions cluster infrastructure by running Terraform or OpenTofu modules as Kubernetes Jobs, instead of implementing cloud-specific logic in Go. [Quick Start]() [Choose a cloud]() [Write a module]() ## Why CAPTF exists Cluster API needs one infrastructure provider per cloud, and most teams already have Terraform or OpenTofu modules that provision that cloud’s infrastructure. CAPTF turns a module written to its contract into a Cluster API infrastructure provider directly: you package the module as an OCI image, and CAPTF runs it, reads its outputs back from state, and reconciles `Cluster`, `Machine` and `MachinePool` objects against them. There is no Go controller to write for a new cloud, and no second copy of infrastructure logic to keep in sync with the module that already exists. ## Find your path - **Try it** --- Install the provider and bring up a cluster with the no-op modules, on any Kubernetes cluster and with no cloud account. - **Run clusters** --- Credentials, module variables, ClusterClass templates, machine pools, drift, plan approval and teardown. - **Use a cloud module** --- Reference modules for AWS, Google Cloud, Azure, OCI and OpenStack, and what each one creates. - **Write a module** --- Write, lint and package a module to the `v1alpha1` contract, then integrate it with a control plane. - **Operate the manager** --- Install, configure, secure, observe, upgrade and recover the CAPTF manager. - **Fix something** --- Start from a condition, an event, an alert or a symptom, and follow the runbook. ## How it fits together ``` flowchart LR subgraph capi["CAPI objects"] Cluster["Cluster, Machine, MachinePool"] end subgraph tf["Terraform* objects"] TFObj["TerraformCluster, TerraformMachine(Pool)"] end Manager["manager"] Job["runner Job"] Module["module image"] Cloud["cloud APIs"] State[("state Secret")] Cluster --> TFObj TFObj --> Manager Manager -- creates --> Job Job --> Module Module --> Cloud Module --> State State --> Manager ``` Cluster API’s core objects (`Cluster`, `Machine`, `MachinePool`) reference CAPTF’s own objects (`TerraformCluster`, `TerraformMachine`, `TerraformMachinePool`) as their infrastructure. The CAPTF manager reconciles those objects, renders their inputs, and runs a Kubernetes Job for each operation. The Job’s runner executes a Terraform or OpenTofu module image against those inputs, calling the cloud’s own APIs, and its state — including the outputs the manager reads back — lives in a Kubernetes Secret. See [Architecture]() for the components behind this diagram and how one apply flows through them. ## Project status > [!WARNING] > > **Pre-release: `v1alpha1`** > > Every kind — `TerraformCluster`, `TerraformClusterTemplate`, `TerraformMachine`, `TerraformMachineTemplate`, `TerraformMachinePool`, `TerraformMachinePoolTemplate` and `TerraformClusterIdentity` — is at API version `v1alpha1`. The module contract they implement is also `v1alpha1` and provisional: it is frozen for implementation, but may still change before the first real module has provisioned a cluster with it (see the contract’s [changelog]()). > > CAPTF has no end-to-end tests yet. Its test suite is unit tests that mock the Kubernetes API and Job execution; nothing in it creates a real cluster. See [Testing]() for how the test suite is organized, and [Known Limitations]() for everything CAPTF does not do yet. # Quick Start This tutorial takes CAPTF from an empty management cluster to a `TerraformCluster` and a control-plane `TerraformMachine` that both report `Ready`, using the no-op modules that ship with CAPTF: they create no real infrastructure, so the tutorial needs no cloud account and no credentials. > [!WARNING] > > **This flow has not been run end to end** > > This tutorial has not been run end to end against a live management cluster. Because the no-op modules provision nothing real, nothing ever boots a kubelet: no `Node` ever joins, so `KubeadmControlPlane` never initializes and the `Machine` never reaches its `Running` phase. What you watch come up in this tutorial is CAPTF’s own objects finishing their applies, not a usable Kubernetes cluster. See [Your First Module]() for writing a module that does create something, and the [module contract]() for turning one into a real cloud provider. > [!NOTE] > > **Before you begin** > > - A Kubernetes cluster to use as the management cluster, and `kubectl` pointed at it. > - `clusterctl`, `make`, Go and `podman` or `docker`: CAPTF has no release yet, so this tutorial builds the provider’s manager image from a clone of this repository instead of fetching it. > - A container registry you can push to, and that the management cluster can pull from, for the manager image. > - A management cluster that can pull from `ghcr.io`, where the no-op module images are published. Run every command below from the root of that clone: the `make` targets and the `templates/...` paths are relative to it. ## 1\. Install the provider CAPTF has not published a release, so `clusterctl` cannot fetch its manifest from a URL yet; build one into a local repository instead. Build and push the manager image, then render the manifest against it: ```sh export IMG=registry.example.com/you/cluster-api-provider-terraform:v0.1.0 make docker-build docker-push IMG="${IMG}" make manifests-release RELEASE_DIR="${HOME}/local-repository/infrastructure-terraform/v0.1.0" \ RELEASE_IMG="${IMG}" VERSION=v0.1.0 ``` Point a `clusterctl` config at that directory; the config entry’s `name` is `terraform`, CAPTF’s registered provider name: clusterctl.yaml ```yaml providers: - name: terraform type: InfrastructureProvider url: file:///home//local-repository/infrastructure-terraform/v0.1.0/infrastructure-components.yaml ``` `` is your home directory’s user name, so the `url` is the absolute path of the directory `manifests-release` just wrote. ```sh clusterctl init --config clusterctl.yaml --infrastructure terraform:v0.1.0 ``` This also installs Cluster API’s core, bootstrap and control-plane providers, and `cert-manager` itself if a compatible version is not already present, since CAPTF’s webhooks need it. See [Installation]() for what this creates and how to confirm it, and [Installing from a local repository]() for the general form of the local-repository steps above. ## 2\. Apply an identity Cloud credentials come from a cluster-scoped `TerraformClusterIdentity`. An admin applies one per set of credentials, naming the namespaces allowed to use it: ```sh export TERRAFORM_IDENTITY_NAME=aws-prod NAMESPACE=team-a clusterctl generate yaml --from templates/identity.yaml | kubectl apply -f - ``` The no-op modules read no credentials, so the generated Secret’s placeholder keys can stay as they are; a module that calls a real cloud provider reads its credentials from the same Secret. `kubectl get terraformclusteridentity aws-prod` shows `Ready=True` once the Secret exists and whoever applied the identity was allowed to `get` it — the admission webhook checks. See [Identities and Credentials]() for creating, rotating and revoking credentials, and how they reach a Job. ## 3\. Choose the no-op module images The no-op modules are published as images by [`captf-io/noop-modules`](), so there is nothing to build. The cluster and machine roles are: ```sh export NOOP_CLUSTER_IMAGE=ghcr.io/captf-io/noop-cluster:terraform export NOOP_MACHINE_IMAGE=ghcr.io/captf-io/noop-machine:terraform ``` Each image comes in two tags, one per base image: `-terraform` and `-opentofu`, such as `vX.Y.Z-terraform`. Either satisfies the [image contract](). The bare `terraform` and `opentofu` tags move to the newest release; pin a release tag, or a digest, in anything you keep. This tutorial uses the `terraform` tag. See [No-op]() for what the modules return. ## 4\. Generate and apply a cluster `Cluster` objects and everything they own live in a namespace `clusterctl` does not create: ```sh kubectl create namespace team-a ``` Generate the default flavor and apply it. This tutorial asks for one control-plane machine and no workers, since a worker never gets bootstrap data until a real control plane initializes, which the no-op modules never do: ```sh export TERRAFORM_CLUSTER_IMAGE="${NOOP_CLUSTER_IMAGE}" export TERRAFORM_MACHINE_IMAGE="${NOOP_MACHINE_IMAGE}" export TERRAFORM_IDENTITY_NAME=aws-prod clusterctl generate cluster my-cluster --from templates/cluster-template.yaml \ --target-namespace team-a \ --kubernetes-version v1.36.3 \ --control-plane-machine-count 1 --worker-machine-count 0 \ | kubectl apply -f - ``` `--from` renders the template file directly, so this command needs no provider registration and works the same before or after a release exists. A ClusterClass-based flavor is also available; see [Templates and ClusterClass](), which also covers every variable this template accepts, and [clusterctl variables]() for their defaults and built-in safeguards. ## 5\. Watch it come up ```sh kubectl get clusters,machines,machinepools,terraformclusters,terraformmachines,terraformmachinepools -n team-a ``` `machinepools` returns nothing: the default flavor creates none. See [Add an autoscaled pool to a generated cluster]() to add one to this cluster. Each `Terraform*` kind reports a `Ready` condition, the only one Cluster API reads (it is mirrored into the owning `Cluster`’s or `Machine`’s `InfrastructureReady`). `Ready` is `Unknown` while an object waits on dependencies or on its apply to finish, and `True` once the module’s apply succeeds: ```sh kubectl get terraformcluster -n team-a my-cluster \ -o jsonpath='{.status.conditions[?(@.type=="Ready")]}' ``` Expect `TerraformCluster my-cluster` and the control-plane `TerraformMachine` to both reach `Ready=True`, and their `PHASE` columns to reach `Provisioned` on `Cluster my-cluster` and its control-plane `Machine` too, since the no-op modules do return a control-plane endpoint and a provider ID. What you will not see, because nothing real ever boots: the `Machine` reaching phase `Running` (no `Node` ever registers), and `KubeadmControlPlane` reporting itself initialized. That gap is expected here and is exactly what a module that creates real infrastructure closes. > [!WARNING] > > **Watch within about 30 minutes** > > Past that, the control-plane `MachineHealthCheck`’s node-startup timeout fires because no `Node` ever registers, and `KubeadmControlPlane` starts remediating the `Machine` (see [Remediation]()). For what each condition type and reason means, see [Conditions](); for the reconcile flow behind these states, see [The Reconcile Lifecycle](). ## 6\. Clean up > [!CAUTION] > > **Delete the Cluster first, not the namespace** > > Delete the `Cluster` first, and wait for it to be gone, rather than deleting the namespace outright: once a namespace starts terminating, the API server refuses to create the destroy Jobs each `Terraform*` object still needs to run before its own finalizer clears. ```sh kubectl delete cluster my-cluster -n team-a kubectl wait --for=delete cluster/my-cluster -n team-a --timeout=10m ``` That wait only returns once every descendant — `KubeadmControlPlane`, the `MachineDeployment`, and both `TerraformCluster` and the control-plane `TerraformMachine` — is gone too, the latter two only after their own destroy Job finished. An identity cannot be deleted while a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` still references it, so delete it only once the wait above returns, then the now-empty namespace: ```sh clusterctl generate yaml --from templates/identity.yaml | kubectl delete -f - kubectl delete namespace team-a ``` To remove the provider itself as well: ```sh clusterctl delete --config clusterctl.yaml --infrastructure terraform ``` > [!NOTE] > > **See also** > > - [Your First Module]() to write a module that provisions something real. > - [Installation]() for what `clusterctl init` installs and how to verify it. > - [Security Model]() for the trust boundary a `Terraform*` object’s Job operates inside. > - [The Kinds]() for how CAPTF’s objects relate to Cluster API’s. # Overview # Architecture This page describes CAPTF’s components, what runs where, and how one apply flows through them. It is for anyone who needs to understand how CAPTF works before reading a more specific page. ## Components **The manager Deployment** is CAPTF’s one Deployment and one container image. It runs the controllers that reconcile every kind except `TerraformClusterTemplate` and `TerraformMachinePoolTemplate`, which have no reconciler of their own, and it serves the validating webhooks for all seven kinds from the same process, on a separate port. It watches every namespace by default: `clusterctl init` never sets a namespace restriction. See [Configuration]() for how to scope or tune it. **The validating webhooks** reject an invalid or disallowed change to a `Terraform*` object before it is persisted: an immutable field, a malformed image reference, a `spec.variables` name that collides with a contract input, or (for `TerraformClusterIdentity`) a Secret the requester cannot read. There are no mutating webhooks — CAPTF resolves every default at reconcile time instead of writing it back onto the object. See [Security Model]() for the trust boundary a webhook decision sits inside, and [RBAC]() for who can change what. **A runner Job** is a `batch/v1` Job the manager creates for one operation (apply, destroy, drift, refresh, restore or plan) on one `Terraform*` object, in that object’s own namespace. Kubernetes never retries a Job pod; the manager owns retries itself, one attempt per Job name. See [The Reconcile Lifecycle]() for how the manager decides which operation runs next. **The module image** is the OCI image a `Terraform*` object’s `spec.source.image` names. It bundles a Terraform or OpenTofu module written to CAPTF’s contract for one role (`cluster`, `machine` or `machinepool`) together with the runtime binary that runs it, at fixed paths, and optionally a provider filesystem mirror. It never bundles the runner. See [Image Contract](). **The runner binary** is copied into each Job’s pod by an init container, rather than shipped in the module image. The manager’s own image bundles both `/manager` and a static `/runner` binary; the init container runs `/runner copy /captf/bin/runner` into an `emptyDir` the module container also mounts, so by default the runner every Job executes is exactly the build of the manager that created the Job; an operator can point it at a different image instead (see [Configuration]()). The module container then runs that copy as its command, against the module and runtime the image provides. The init container’s resources are fixed and not configurable, since it only copies one static binary and never varies with the module. **The Kubernetes state backend** is the Terraform/OpenTofu `kubernetes` backend, which the manager configures for every generated root module. State lives in Secrets in the `Terraform*` object’s own namespace, chunked when it is large, and labeled by the object’s immutable identity — namespace, kind and name, never its UID — so state survives a `clusterctl move` and a changed identity is never mistaken for existing state. See [Terraform State](). ## What runs where | Component | Runs as | Namespace | | --- | --- | --- | | manager, webhooks | One Deployment, one container | The provider’s own namespace (`captf-system` by default) | | runner Job | One `batch/v1` Job per operation | The `Terraform*` object’s own namespace | | runner binary, module | Init and main containers of the Job’s pod | Same pod as the Job | | state | Secrets | The `Terraform*` object’s own namespace | ## The data flow of one apply The manager reconciles a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` through a shared core (externally-managed check, owner lookup, finalizer, pause) and decides that an apply is the next operation. It renders that object’s inputs into a generated root module — `main.tf.json`, which configures the `kubernetes` backend and calls the module image’s role module with exactly the contract inputs, and `terraform.tfvars.json`, the input values — and writes them to a durable inputs Secret and a per-run Secret, hashing the inputs as it does. See [Job Inputs]() for exactly where each value comes from. The manager then builds and creates the Job: the init container that copies the runner binary, and the module container that mounts the per-run Secret read-only, the identity’s mirrored credentials Secret, and empty scratch volumes. Once the pod starts, the runner initializes the backend, plans and applies against the module’s runtime, and — for a `TerraformCluster` apply — stops before a plan that deletes or replaces a resource unless that plan was already approved (see [Plan Approval]()). The generated root re-exports the module’s contract outputs, so they land in state alongside everything else the runtime tracks. The runner reports how the Job finished as its pod’s termination message. The manager reads that message, and — once the Job has finished — the state Secrets themselves, to set the object’s status and conditions and, after a successful apply, to adopt (re-label) the new state Secrets as its own. ## What a module author provides versus what an operator provides A **module author** writes and builds the OCI image: the role module (and any local modules it calls), pinned to the [module contract]() and the [image contract](), optionally with a provider mirror baked in. They check it with [tfcapi-lint]() before publishing it. An **operator** installs the manager ([Installation]()), creates the `TerraformClusterIdentity` credentials CAPTF’s objects reference ([Identities and Credentials]()), and creates the `Terraform*` objects — directly or through [templates and ClusterClass]() — that name a module author’s image. They tune Job resources, deadlines and security contexts per object ([Tuning Jobs]()), and watch the conditions, events and alerts the manager and runner produce ([Observability]()). # The Kinds If you are learning how CAPTF’s objects relate to Cluster API’s, start here. CAPTF adds seven kinds to one API group and version, `infrastructure.cluster.x-k8s.io/v1alpha1`. Three of them are the objects that carry the desired state Cluster API’s core controllers drive as a cluster comes up; three are templates that stamp those out; the seventh holds the cloud credentials the others use. Every field, default and validation rule is in [Custom Resources](); this page explains how the seven fit together, and what stays fixed once an object exists. ## The seven kinds | Kind | Cluster API role | Scope | Referenced by | | --- | --- | --- | --- | | `TerraformCluster` | InfraCluster | Namespaced | `Cluster.spec.infrastructureRef`, always directly | | `TerraformMachine` | InfraMachine | Namespaced | `Machine.spec.infrastructureRef`, always directly | | `TerraformMachinePool` | InfraMachinePool | Namespaced | `MachinePool.spec.template.spec.infrastructureRef`, always directly | | `TerraformClusterTemplate` | Template | Namespaced | A ClusterClass | | `TerraformMachineTemplate` | Template | Namespaced | A MachineDeployment, MachineSet, KubeadmControlPlane, another control-plane provider, or a ClusterClass | | `TerraformMachinePoolTemplate` | Template | Namespaced | A ClusterClass | | `TerraformClusterIdentity` | None (CAPTF-only) | Cluster | `spec.identityRef` of a cluster, machine or pool in an allowed namespace | ## The three workload kinds A `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` runs one Terraform or OpenTofu module role — cluster, machine or machinepool — as a Job. What goes into that Job and how the reconcile loop drives it are covered in [Job inputs]() and [the reconcile lifecycle](); this section is about the object itself. Each of the three is owned by its Cluster API counterpart, not by CAPTF: the counterpart’s own core controller sets the owner reference. A `TerraformCluster` or `TerraformMachinePool` can be an object you create and point the `Cluster` or `MachinePool` at directly, or one a ClusterClass topology expands from a `TerraformClusterTemplate` or `TerraformMachinePoolTemplate`; either way, the resulting object is owned the same way. A `TerraformMachine` is normally expanded from a `TerraformMachineTemplate`, one per `Machine`, by a MachineDeployment, a MachineSet or a control-plane provider, whether or not a ClusterClass is involved; a `Machine` can also point at a directly created `TerraformMachine`, the same way a `Cluster` or `MachinePool` can point at a directly created `TerraformCluster` or `TerraformMachinePool`. - **`TerraformCluster`** is owned by a `Cluster`. `spec.controlPlaneEndpoint` is mutable until it has a host — set by you, or once by the controller from the module’s output — and immutable after: every Machine and kubeconfig of the cluster points at it. Its `host` and `port` must be set together; the webhook rejects one without the other. Everything else, including the module image, can change on a live object; a new image or a changed input re-applies the module against the existing state. `spec.identityRef` is required. `spec.applyPolicy` and the destructive-plan guard are covered in [plan preview and approval](). - **`TerraformMachine`** is owned by a `Machine`. `spec.source`, `spec.identityRef`, `spec.variables` and `spec.variablesFrom` define the machine and are immutable after creation: change them through a MachineDeployment or control-plane rollout, not in place. `spec.providerID` can only move from empty to non-empty, and only by the manager’s ServiceAccount; a create may still carry one, so `clusterctl move` can restore it. A `TerraformMachineTemplate` rejects `providerID` outright. `spec.jobs`, `spec.drift` and `spec.remediation` are operational policy and stay mutable at any time, so a stuck machine’s deadline or drift interval can be changed without rolling it. A direct delete is refused while a Machine that is not itself being deleted references the object through `spec.infrastructureRef`, except a `clusterctl move` delete (`clusterctl.cluster.x-k8s.io/delete-for-move`) while the Cluster is paused; see [the reconcile lifecycle]() for deletion and finalizers. - **`TerraformMachinePool`** is owned by a `MachinePool`. Unlike a `TerraformMachine`, every field is mutable: a change to `spec.source`, `spec.variables`, `spec.variablesFrom` or any other field re-applies the module on the next reconcile. A pool has no `spec.applyPolicy`. Its apply is guarded only when it renders a change of the cluster’s exports (see [the guard]()). With autoscaling enabled, it also writes the group’s observed replica count back to `MachinePool.spec.replicas`; see [machine pools](). ## Templates `TerraformClusterTemplate`, `TerraformMachineTemplate` and `TerraformMachinePoolTemplate` each hold `spec.template`, the metadata and spec a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` is created with. A `TerraformClusterTemplate` and a `TerraformMachinePoolTemplate` are each expanded by a ClusterClass topology into one `TerraformCluster` or `TerraformMachinePool`, which the `Cluster` or `MachinePool` then references directly. A `TerraformMachineTemplate` is expanded, with or without a ClusterClass, by a MachineDeployment, a MachineSet, a KubeadmControlPlane or another control-plane provider, into one `TerraformMachine` per `Machine`. See [ClusterClass]() for how templates and flavors fit together. > [!WARNING] > > **A template’s spec is immutable** > > `spec.template.spec` is immutable once a template exists, like every Cluster API template: create a new template and point the owner at it instead of editing one in place. `spec.template.metadata` stays mutable. The one exception is a ClusterClass topology dry-run, which the webhook recognizes and exempts from the immutability check, so a topology patch that only touches `spec.template.spec` can still be validated without tripping it. A `TerraformMachineTemplate` additionally reports `status.capacity` and `status.nodeInfo`, resolved from its image’s labels, for Cluster Autoscaler scale-from-zero. `TerraformMachinePoolTemplate` has no status: a pool has no scale-from-zero, so there is no capacity to resolve. ## `TerraformClusterIdentity` `TerraformClusterIdentity` is cluster-scoped and has no Cluster API counterpart: it exists only to hold `spec.secretRef`, the credentials Secret a `TerraformCluster`’s, `TerraformMachine`’s or `TerraformMachinePool`’s module run is allowed to use, and `spec.allowedNamespaces`, which namespaces may reference it. Every field is mutable, including `secretRef`; changing it re-checks that the requester may read the new Secret. Deleting an identity is refused while any object still uses it or still has its credentials mirrored into a namespace. See [identities]() for creating, rotating and revoking credentials, and how they reach a Job. ## What a cluster passes to its machines and pools A `TerraformCluster`’s `spec.defaults` apply only to its `TerraformMachine`s and `TerraformMachinePool`s, never to the cluster itself, and are merged field by field: a field the machine or pool sets wins, an unset one comes from `spec.defaults`, and a field neither sets gets the built-in default. A machine or pool finds its `Cluster` by its own `cluster.x-k8s.io/cluster-name` label, then the `TerraformCluster` through the Cluster’s `spec.infrastructureRef`. - **`identityRef`**: a machine or pool uses its own `identityRef` when it sets one, else the cluster’s `spec.defaults.identityRef`, else the cluster’s own `spec.identityRef`. - **`jobs`**: merged field by field with `spec.defaults.jobs`, own over defaults; see [Tuning Jobs]() for the full merge algorithm. - **`drift`**: only `intervalSeconds` is inherited — the object’s own, else `spec.defaults.drift.intervalSeconds`, else the manager’s built-in default. A pool’s own `drift.action` is never inherited; unset, it defaults to `Report` regardless of the cluster’s defaults, which carry no `action` to inherit. > [!NOTE] > > **The pool exception** > > A `TerraformMachine`’s drift can be turned off with `intervalSeconds: 0`, and an inherited `spec.defaults.drift` of 0 turns it off the same way, but a `TerraformMachinePool`’s drift can never be turned off; see [Drift and health]() for why. > [!NOTE] > > **See also** > > - [Custom Resources]() for every field, default and validation rule. > - [The reconcile lifecycle]() for what runs when, retries, deletion and finalizers. > - [Job inputs]() for what each role’s Job receives. > - [ClusterClass]() for templates in use. > - [Machine pools]() for autoscaling and replica write-back. > - [Identities]() for `TerraformClusterIdentity` in depth. # Glossary Short definitions of CAPTF terms used across this book, and the Cluster API and Terraform or OpenTofu terms CAPTF’s own docs assume. Each entry links to the page that covers the term in full; this page never repeats what that page already says. ****apply**** The Job operation that renders an object’s current inputs, records them as its durable inputs, then plans and applies them. The destructive-plan guard stops it after the plan step; see [choosing the next operation](). ****applyPolicy**** A `TerraformCluster` field, `Automatic` (the default) or `Manual`, that decides whether every apply first waits for a reviewed plan. See [Plan Approval](). ****bookkeeping**** The reconcile step that reads an object’s finished Jobs, records each one’s result in `status.lastRun`, pins a succeeded apply’s image digest, and releases the Job’s leases. See [the reconcile lifecycle](). ****ClusterClass**** Cluster API’s reusable cluster template mechanism. CAPTF ships a `noop` `ClusterClass`, which expands a `TerraformClusterTemplate` and two `TerraformMachineTemplate`s. See [Templates and ClusterClass](). ****cluster operation gate**** The manager’s `--cluster-operation-gate` flag (`true` by default), which keeps a `TerraformCluster`’s own apply or destroy from running at the same time as its machines’ and machine pools’, through a per-cluster write lease. See [run leases and the cluster operation gate](). ****contract**** The normative interface between the controller and a module: which inputs each role receives and which outputs it must produce. See [Module contract](). ****destroy**** The Job operation that runs the module’s `destroy` step against an object’s durable inputs, run only on deletion. See [deletion order](). ****destructive plan**** A plan that deletes or replaces at least one resource. A `TerraformCluster` stops before applying one until the exact inputs hash it belongs to is approved. See [Plan Approval](). ****drift**** A difference between a provisioned object’s Terraform or OpenTofu state and reality, found by a drift Job’s plan. See [Drift and Health](). ****durable inputs (Secret)**** `captf-inputs--`, the record of what the controller rendered when it last started an apply Job for an object. An immutable machine’s destroy, drift and refresh always use it; a mutable cluster or pool’s destroy prefers it and falls back to current inputs, while its drift and refresh prefer current inputs and fall back to it. See [Job Inputs](). ****identity**** A `TerraformClusterIdentity`: names the Secret of cloud credentials a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` uses, and the namespaces allowed to use it. See [Identities and Credentials](). ****InfraCluster / InfraMachine / InfraMachinePool**** Cluster API’s generic roles for an infrastructure provider’s objects. CAPTF’s are `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool`. See [The Kinds](). ****inputs hash**** `captf.io/inputs-hash`: covers the contract version, the role, `spec.source.image` as written, the rendered inputs, and any user variables. A change to it re-applies a mutable object. See [Job Inputs](). ****kindshort**** The short kind code in a Secret, Lease or Job name: `c` for `TerraformCluster`, `m` for `TerraformMachine`, `mp` for `TerraformMachinePool`. See [Secret names and the suffix](). ****mirror Secret**** `captf-creds-`, the copy of an identity’s credentials Secret the controller mirrors into each allowed namespace where an object uses the identity, and keeps in sync, so a Job never reads the source Secret directly. See [Identities and Credentials](). ****module**** The Terraform or OpenTofu code that implements one contract role, shipped in an OCI image together with the runtime that runs it. See [Image Contract](). ****per-run Secret**** `captf-run-`, a private copy of a Job’s inputs, owned by the Job, mounted at `/captf/config`, and deleted when the Job finishes. See [Job Inputs](). ****plan hash**** Under `applyPolicy: Manual`, the hash over a plan’s sorted, non-no-op changes that a `captf.io/approve-plan` annotation names to approve it. See [Plan Approval](). ****provider mirror**** The optional `/captf/providers` filesystem mirror an image can ship, so `init` needs no registry egress at run time. See [Runtime Environment](). ****refresh**** The Job operation that runs `apply -refresh-only`: it updates state from reality and produces a health reading, without planning or finding drift. See [Drift and Health](). ****restore**** The Job operation, requested by the `captf.io/restore-state` annotation, that pushes a listed state backup back into the backend with `state push -force`. See [Restore](). ****role**** Which of `cluster`, `machine` or `machinepool` a module implements, and which of `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` runs it. See [Module contract](). ****run lease**** A `coordination.k8s.io/v1` Lease the controller takes on an object before starting a Job, so two reconciles never start two Jobs for the same object at once. See [run leases and the cluster operation gate](). ****runner**** The binary that is a Job’s entrypoint: it drives init and the requested operation (plan, apply, destroy, refresh, drift or restore), enforces the destructive-plan guard and, under `Manual`, a plan’s approval, and emits step events on the object. See [Runner CLI](). ****source image**** `spec.source.image`, the OCI image naming a `Terraform*` object’s module and runtime. It is CAPTF’s trust boundary: whoever may set it controls what the Job’s Pod does. See [Security Model](). ****state backup**** `captf-state-backup--`, a versioned copy of an object’s state Secrets, taken whenever the controller observes a state serial it has not backed up before. See [Terraform State](). ****workload kinds**** `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool`: the three kinds that each run one Job-backed Terraform or OpenTofu module role and share a common spec and status shape. See [The Kinds](). > [!NOTE] > > **See also** > > - [Custom Resources]() for every field these terms name. > - [Conditions]() for the reasons the kinds above report. > - [Module contract]() for the normative meaning of a role and its inputs and outputs. # Security Model This page states CAPTF’s trust boundary: what creating a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` grants, what the Job that runs it can read, and what CAPTF keeps out of status, events and logs. Read it before deciding who may create or update a `Terraform*` object, and before deciding whether two tenants can share a namespace. For the mechanics behind each control, see [Identities and Credentials](), [RBAC]() and [Secrets](). ## A `Terraform*` object is a Pod Once its Cluster API owner references it, a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` makes CAPTF run Jobs whose main container runs `spec.source.image`, with `jobs.env`, as the namespace’s `captf-runner` ServiceAccount (or an opted-in override named by `jobs.serviceAccountName`), with the resolved identity’s credentials injected as environment variables and mounted as read-only files ([Identities and Credentials]()). An object no owner references runs nothing (see [Owner references are checked](<#owner-references-are-checked-against-the-owner>)). > [!CAUTION] > > **Whoever controls the image has everything the Pod runs with** > > Running a Job is equivalent to granting whoever controls the image everything the Pod runs with: the runner’s access to Secrets in the namespace (below) and the identity’s cloud credentials. Module code, and so any image a `Terraform*` object names, runs with that access; a `provider "kubernetes" {}` block or a `local-exec` provisioner can use it directly. Consequences: - **The source image is the trust boundary.** Whoever may set `spec.source.image` on an owned `Terraform*` object in a namespace, or on the template it is cloned from, already has, in effect, the Secret access described below and the cloud credentials of every identity allowed in that namespace. `TerraformCluster` and `TerraformMachinePool` are mutable, so `update` on them is enough. Restrict `create`/`update` on `terraform*` kinds and their templates to principals who already hold that. The controller checks only that `spec.source.image` is a syntactically valid image reference; it does not police which registry or repository it names. An admission policy on `spec.source.image` by registry prefix is the recommended control for restricting which images a namespace may run. - **One identity and one workload cluster per tenant namespace.** Every `Terraform*` object in a namespace, and every image any of them names, can read the Secrets described below, including the credential mirrors of every other identity in use in that namespace and the state and inputs of every other object there. A namespace is the boundary between tenants; sharing one namespace between two tenants’ clusters or identities gives each tenant everything described in this page for the other’s cluster too. - **Pod Security Admission applies to Jobs like any workload.** A namespace that enforces it gets the defaults described in [Pod security](<#pod-security>). ## Owner references are checked against the owner CAPTF treats a Cluster API object as the owner of a `Terraform*` object only when the owner references it back. For a `TerraformMachine`, the `Machine` named by its `Machine`-kind `ownerReferences` entry must pass three checks: - The Machine’s `spec.infrastructureRef` names this `TerraformMachine`: matching `apiGroup`, `kind: TerraformMachine` and `name`. Cluster API sets it before it adds the owner reference, so every `TerraformMachine` Cluster API creates passes. - When the `ownerReferences` entry carries a UID, it equals the Machine’s UID. - When the `TerraformMachine` has a `cluster.x-k8s.io/cluster-name` label, it equals the Machine’s `spec.clusterName`. `TerraformMachinePool` applies the same checks to its `MachinePool`, using `spec.template.spec.infrastructureRef`, and `TerraformCluster` to its `Cluster`, using `spec.infrastructureRef`. An owner that fails any check is not an owner: `DependenciesReady=False`/`OwnerMismatch` names the failed check, no Job runs, and CAPTF writes nothing to that object or its Cluster: no remediation annotation on a Machine, no replica count on a MachinePool. Deleting the `Terraform*` object works as it does when its owner is gone: destroy runs from the durable inputs. The `TerraformMachine` delete webhook uses the same checks. It refuses a direct delete only while an owner that passes them exists and is not being deleted, so an object whose `ownerReferences` name a Machine that does not own it can always be deleted. The owner must reference the object through `spec.infrastructureRef`. A delete for `clusterctl move` is allowed while the Cluster is paused. `create` on a `terraform*` kind alone therefore runs nothing: a Job starts only once a Cluster API object that references the new object exists. Who may create those, and who may set `spec.source.image` on an owned object, is what the section above describes. ## What the runner can read, and why The runner ServiceAccount’s permissions come from a single, static ClusterRole bound namespace-by-namespace; see [RBAC]() for the exact rules and how a custom ServiceAccount opts in. On Secrets, it holds `get`, `list`, `create`, `update` and `delete`, and none of that is scoped by name or label: the Terraform and OpenTofu Kubernetes state backend needs `list` to enumerate its state chunks and workspaces on every read and write, and `list` (like `create`) cannot be restricted to named Secrets at all. `get`, `update` and `delete` could in principle be scoped with `resourceNames`, but the backend names each state chunk itself, per apply and per workspace, so no fixed rule can list them in advance. The runner, and therefore any module image, can as a result read, replace or delete every Secret in its namespace, which includes: - the state and inputs Secrets of every `Terraform*` object in the namespace, not only the one the running Job belongs to; - the credential mirrors (`captf-creds-*`) of every identity in use in the namespace, not only the one the running Job was given ([Secrets]() lists every Secret CAPTF reads or writes and its sensitivity); - any CAPI core Secret in the namespace, such as a cluster’s kubeconfig, certificate authority (CA) or a Machine’s bootstrap data. Write access here means a hostile module can substitute a cluster’s CA or kubeconfig, not only read them. > [!WARNING] > > **One identity and one workload cluster per namespace is the only tenant isolation** > > This is why one identity and one workload cluster per namespace, above, is the only real isolation CAPTF offers between tenants sharing a management cluster. The manager itself can also read every Secret in the cluster, as any CAPI infrastructure provider that runs the Kubernetes state backend effectively can. ## Image pinning by digest The first time an apply Job of a `Terraform*` object **succeeds**, the controller records the image digest the kubelet actually ran (read from the pod’s container status, not from `spec.source.image`) as `captf.io/image-digest` on the object’s durable inputs Secret; a failed apply pins nothing. See [Annotations, Labels and Finalizers]() for the annotation. - `TerraformMachine` is immutable ([the kinds]()): once pinned, every later drift check and destroy Job for that machine runs `repo@sha256:…`, never the tag in `spec.source.image`. A tag that moves after the successful apply can therefore never change the code that destroys an existing machine. - `TerraformCluster` and `TerraformMachinePool` are mutable: each spec-driven apply re-resolves the tag and re-pins the digest it ran. Pinning guards against an accident — a tag moved out from under a running cluster changing what a later destroy runs — not against a hostile image: the pinned digest is whichever image `spec.source.image` named when the apply succeeded, and that image’s own runtime computed the plan and ran the providers. Approving a blocked destructive plan or a Manual-policy plan preview is exactly as privileged as setting `spec.source.image`, since both need `update` on the object; see [Plan Approval]() and [who can approve](). > [!WARNING] > > **Module image signatures are not verified** > > CAPTF’s own manager image is not signed, and nothing in CAPTF verifies a module image’s signature before running it: digest pinning fixes which image ran after the fact, it does not check who published it. ## Pod security The Job’s pod defaults satisfy the Pod Security `baseline` profile without any configuration: seccomp defaults to `RuntimeDefault`, and `fsGroup` defaults to `65532` so a non-root image user can read the credential files, mounted at mode `0440`, through that supplementary group regardless of the image’s own user or group. The main container additionally defaults to `allowPrivilegeEscalation: false`, `capabilities.drop: [ALL]` and `readOnlyRootFilesystem: true`, and the runner’s own init container is fixed non-root, fully locked down, and never configurable. `runAsNonRoot` is **not** defaulted at the pod level, because an image built `FROM hashicorp/terraform` runs as root unless it sets `USER`; reaching the `restricted` profile needs an image that tolerates `runAsNonRoot: true`, set through `jobs.podSecurityContext` or `jobs.securityContext`. Regardless of profile, the admission webhook rejects `privileged: true`, `allowPrivilegeEscalation: true`, any `capabilities.add`, `readOnlyRootFilesystem: false`, a `seccompProfile` of `Unconfined`, `procMount: Unmasked`, `windowsOptions.hostProcess` and an explicit `runAsUser: 0` or `runAsNonRoot: false` in `jobs.securityContext`, on every object and template that carries a jobs policy (`spec.jobs`, `spec.defaults.jobs` on a `TerraformCluster`, and the same field on a `TerraformMachine`, `TerraformMachinePool` and all three `*Template` kinds), because that container holds the resolved identity’s cloud credentials. Pod Security Admission remains the namespace-wide control for everything else a jobs policy does not set, such as host namespaces and volume types. These stricter rules apply on create and whenever the jobs policy changes. An existing object with an older, weaker policy still accepts unrelated updates and can always be deleted. For what the approval gates promise and what they do not, see [Approvals and Gates](). For the Secrets themselves (the state, its backups, the inputs and the credential mirrors) and what the runner’s access to them means, see [Security considerations](). ## What CAPTF keeps out of status, events and logs `Terraform*` object status is readable by anyone who can `get` the object — far more people, in general, than can read Secrets in the namespace — so it never carries raw process output. `status.lastRun.error.summary` is the runner’s own short description of a failure, at most 512 bytes, never the failing step’s stderr; the full output stays in the Job’s own logs, which need `pods/log` access to read. The events the runner emits on the object (`RunStarted`, `StepStarted`, …, `RunFinished`; see [Observability]()) carry step names, exit codes, durations and resource-change counts, and, on failure, the same curated summary as status — never tfvars, plan output or resource values. Runner failure summaries, events and logs also replace known secrets with `(sensitive)`. That covers environment values whose name looks like a credential (it contains `KEY`, `SECRET`, `TOKEN`, `PASS`, `CREDENTIAL`, `PRIVATE`, `AUTH`, `CERT` or `SESSION`) or whose value is at least 16 bytes, the values of sensitive variables, `bootstrap_data` both encoded and decoded (each decoded line of 16 bytes or more individually), and the sensitive resource attributes from the plan. Plan outputs are not redacted. > [!WARNING] > > **Redaction is best effort** > > Redaction is best effort, not a guarantee: do not rely on it to make a secret safe to print. A module variable sourced from a Secret ([Module Variables]()) is automatically declared `sensitive = true` in the generated root — one sourced from a ConfigMap or given inline never is — so Terraform redacts it from the Job’s own plan and apply output and the controller redacts it from its own trace-level logs. That redaction stops there: like every other input, the value is written in clear into the object’s durable and per-run inputs Secrets and into the Terraform state Secret, so anyone who can read Secrets in the namespace can read it in either place ([Secrets](), [Job Inputs]()). Cloud credentials never go through this path at all: the resolved identity’s credentials are mounted into the Job and never rendered into a variable, so they never reach the inputs Secrets or the state. ## Network exposure CAPTF ships no `NetworkPolicy` by default; applying one is an opt-in step covered in [Installation](). Without it, nothing restricts which pods the manager or a Job can reach, or which pods can reach them, beyond whatever the cluster otherwise enforces. The manager needs only inbound traffic to its webhook, metrics and health-probe ports, and outbound traffic to DNS, the API server and the registries it reads module image metadata from. A Job’s pod is the more sensitive workload: it holds the resolved identity’s cloud credentials and a ServiceAccount token that can write every Secret in its namespace (above), so its egress is worth restricting to DNS, the API server, and the specific provider and registry endpoints the module it runs needs — nothing reaches it inbound. A `NetworkPolicy` only has an effect on a CNI that enforces one. > [!NOTE] > > **See also** > > - [Identities and Credentials]() > - [RBAC]() > - [Secrets]() > - [The Kinds]() # Known Limitations This page lists what CAPTF does not do, or does with a catch, as the code stands today. Each item says what the limit is, what follows from it, and where the detail is. A limit that has a workaround says so. For versions and what has and has not been tested, see [Compatibility](). ## Maturity - **Pre-alpha.** No release is published. Every API kind is `v1alpha1`, and the module contract is `v1alpha1` and provisional: it may change before a real module has provisioned a cluster with it. See [Project status](). - **No end-to-end run against a live management cluster has happened.** The tests are unit tests that mock the Kubernetes API and Job execution; CI runs unit tests, lint and the offline verifications. Anything that needs a real cluster, a real cloud or a real provider is untested: see [Compatibility](). ## Approvals and gates - **Only the `TerraformCluster` is fully gated.** `applyPolicy: Manual` and the destructive-plan guard exist on the cluster. A machine’s apply never waits for approval. A pool’s apply is guarded in one case only: when it renders changed cluster exports. See [What is guarded](). - **A pool’s exports guard has limits.** A destructive pool apply of changed exports is **held**: the pool keeps applying with the exports of its last successful apply until the change is approved. The guard needs the record of the last applied exports, which shares the durable Secret’s budget with the rendered inputs, and `clusterctl move` re-runs a blocked plan once. Keep shared and destructive infrastructure in the cluster module, keep exports stable, and use `prevent_destroy` on what must not go. See [Machine pools]() and [Cluster outputs reach pools and machines](). - **An approval binds a plan, not the apply.** It names the plan’s hash and the apply must plan the same again; it does not promise what the provider does. See [An approval binds a plan](). ## Security > [!CAUTION] > > **A module image can read every Secret in its namespace and forge a plan hash** > > - **The module image can read the plan key.** The per-object key that makes plan hashes unforgeable by accident is mounted read-only in the module’s container for plan and approved-apply Jobs. A hostile module can read it and forge a plan hash. It protects against drift, not against the module. See [The plan key is readable by the module](). > - **The runner can read and write every Secret in its namespace.** RBAC cannot scope the state backend’s access. Namespaces are the only tenant boundary. See [Multi-Tenancy](). - **No encryption at rest of its own.** State, inputs and credential mirrors are Kubernetes Secrets. See [No encryption at rest](). ## State - **OpenTofu state encryption is unsupported.** An encrypted state reads as `StateReadable=False`/`StateEncrypted`, and no backup of it is ever taken. See [Unreadable State](). - **The state file version must be 4.** Another version reads as `StateCorrupt`. - **Backups are not disaster recovery.** They are in the same namespace, owned by the object, and rotated out after `--state-backups`. See [Disaster Recovery](). - **Owner references are repaired on the next reconcile, not at once.** The controller owns a state chunk, backup, durable inputs Secret and plan key again on the next reconcile that finds no Job running. The gaps are narrow: nothing is re-owned while the object is paused (so `clusterctl move` is not raced), state chunks wait while a Job holds the run lease, and a restored Secret that still names an old UID can be garbage-collected before the first reconcile. See [Disaster Recovery](). ## Jobs and leases - **The lease grace gap.** If a `TerraformCluster`’s inputs change while it waits for machine operations, the old leases are held by a Job that will never exist, and the new Job waits out the one-minute grace. It resolves itself. See [The known gap](). - **No flag caps running Jobs.** The concurrency flags cap reconciles per kind, not Jobs. The bounds are one Job per object and your quotas. - **No quota-specific condition.** A ResourceQuota refusal shows as a reconcile error or a Job that never starts. See [Production Readiness](). - **A single manager watches one namespace or all of them.** The leader-election lease name is fixed, so per-namespace manager instances do not work; use `--watch-filter`. See [Namespace scoping](). ## Deletion and move - **A paused object’s deletion waits.** A paused object, or one under a paused Cluster, never runs a destroy. `Deleting` says so. This is by design, so `clusterctl move` can delete source objects. See [Pause stops a deletion](). - **A moved object in a deleted namespace looks never-applied.** Status does not move, and the `captf.io/applied` marker is on a Secret the namespace deletion removes. Its finalizer then comes off with nothing destroyed. See [The known limit](). - **The credential source Secret does not move.** Copy it to the target yourself. See [clusterctl move](). ## Machine pools - **The Kubernetes Cluster Autoscaler is unsupported on pools.** The controller supports **cloud-native** autoscaling in the module, driven by the pool’s min and max annotations, and writes the observed replicas back to `MachinePool.spec.replicas`. The Cluster Autoscaler’s `clusterapi` provider needs MachinePool Machines and drains nodes before it scales down, and neither exists. Its changes to `spec.replicas` are overwritten by the write-back. See [Machine Pools](). - **MachinePool Machines are unsupported.** Pool instances have no `Machine` objects, so a `MachineHealthCheck` never selects them. - **The bootstrap Secret is not watched.** A rotated token reaches the pool at its next reconcile, at the latest one membership-refresh interval later. > [!NOTE] > > **See also** > > - [Compatibility](). > - [Production Readiness](). > - [Approvals and Gates: Limits](). # Compatibility This page lists the versions CAPTF is built against and what it needs from its environment. It separates what is **tested** from what is **assumed**, because CAPTF is pre-alpha and has never run against a live management cluster. The numbers come from the repository at the time of writing (`go.mod`, `metadata.yaml`, the Makefile and the noop module images), so check them against the tag you install. ## Versions | Component | Version | Source | | --- | --- | --- | | Go | 1.26 (toolchain go1.26.8) | `go.mod`, `Makefile` | | Cluster API | v1.14.2 (the Go module the controllers build against) | `go.mod` | | Cluster API contract | `v1beta2` | `metadata.yaml`, the CRD label | | controller-runtime | v0.24.1 | `go.mod` | | Kubernetes client libraries | v0.36.3 (`k8s.io/api`, `apimachinery`, `client-go`) | `go.mod` | | CAPTF release series | 0.1 (none published) | `metadata.yaml` | | CAPTF API and module contract | `v1alpha1`, provisional | the CRDs, [contract]() | | cert-manager | `cert-manager.io/v1` API required; no minimum release stated | [Installation]() | | Terraform | Any 1.x with the CLI surface below; modules require `>= 1.5` | [image contract]() | | OpenTofu | Any 1.x with the same surface; modules require `>= 1.5` | the same | | Terraform in the noop images | 1.16.4, pinned by digest | [`noop-modules`]() `Dockerfile.terraform` | | OpenTofu in the noop images | 1.12.6, pinned by digest | [`noop-modules`]() `Dockerfile.opentofu` | | State file format | Version 4 only | the state reader | | Architectures | `linux/amd64` and `linux/arm64` image builds | the Makefile’s `PLATFORMS` | ### Kubernetes CAPTF links the Kubernetes client libraries at v0.36, which corresponds to Kubernetes 1.36. It does not state a supported server range. The other bounds are Cluster API’s own: the management cluster must run a Cluster API release that implements contract `v1beta2`. Treat Kubernetes 1.36 as the version it is built for and anything else as unverified; the features it uses are standard (Jobs, Leases, Secrets, validating webhooks, `SubjectAccessReview`). ### The runtime CLI The module image supplies the `terraform` or `tofu` binary at `/captf/runtime`. CAPTF needs the 1.x CLI surface `version`, `init`, `validate`, `plan`, `apply`, `destroy`, `force-unlock`, `show` and `state push`/`state list`. It does not check a minimum version. The reference modules declare `required_version = ">= 1.5"`, because they use `terraform_data` and `plantimestamp()`. CAPTF reads the state through the Kubernetes backend and only accepts state file version 4. OpenTofu client-side state encryption is unsupported. See [Image Contract]() and [Runtime Environment](). ### Cluster API providers CAPTF is an infrastructure provider. Pairing it with a control-plane and bootstrap provider is covered by [Control-Plane Integration](): KubeadmControlPlane and RKE2ControlPlane are the documented ones. That documentation is derived from those providers’ contracts, not from a live run. ## What is tested The continuous-integration workflow runs these on every push: | Check | Covers | | --- | --- | | `make test-cover` and `make cover-check` | Unit tests of the controllers, runner, webhooks, linter and libraries, against fake clients, with per-package coverage floors | | `make lint` and `make vet` | Go lint, API lint, and `go vet`, including the e2e-tagged test code | | The `verify` targets | Generated files are current, component manifests, templates, JSON schemas, `metadata.yaml` append-only, the local clusterctl repository layout, licenses and the Prometheus rules (`promtool` check and tests) | The workflow also builds every binary and takes a `tfcapi-lint` release snapshot. It does not run the end-to-end suites: `make e2e-foundation` and `make e2e-noop` run them on a local kind cluster that pulls the published no-op images, and you run them yourself. The docs checks and the release flow are also outside the workflow. ## What is assumed > [!WARNING] > > **Nothing below has been exercised against a live system** > > Treat every item in this list as unverified. - **A real management cluster.** No `clusterctl init`, `upgrade` or `move` has run end to end; the move behavior is derived from the code and `clusterctl`’s documented rules. - **Real Terraform or OpenTofu execution under CAPTF.** The runner is unit tested against recorded plan and state JSON. Other versions of the runtime than the pinned noop ones are assumed to behave. - **Real infrastructure providers.** The noop modules create no cloud resources, and no module that does has been run under CAPTF in CI. - **Kubernetes server versions other than 1.36, cert-manager releases, and Cluster API releases other than the one built against.** - **The `arm64` image.** It is built by the Makefile; no CI job runs it. - **KubeadmControlPlane and RKE2ControlPlane integration.** Documented against their contracts; not run. If you find a version that works or does not, that is the information this page needs. See [Known Limitations]() and the [project status](). > [!NOTE] > > **See also** > > - [Installation]() and [Upgrades](). > - [Testing]() for how the suite is organized. > - [Production Readiness](). # User Guide # Identities and Credentials A `TerraformClusterIdentity` is a cluster-scoped object that names a Secret of cloud credentials and the namespaces allowed to use it. Every `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` needs one, directly or by inheriting its cluster’s, before it can run a Job. This page covers creating an identity, choosing which namespaces it allows, referencing it from each kind, rotating and revoking its credentials, and deleting it. For the trust an identity’s credentials carry once mounted into a Job, see [Security model](); for the mirror’s lifecycle, see [Credentials](). > [!NOTE] > > **Before you begin** > > - Cluster-scoped access to create `TerraformClusterIdentity` objects, and namespace access to create the Secret it names. > - `get` access to that Secret: creating an identity, or pointing an existing one at a different Secret, checks it. > - A `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` to reference the identity from, or a namespace to create one in. ## How credentials reach a Job An identity’s credentials never reach a Job directly. Into each allowed namespace where an object uses the identity, the controller mirrors the source Secret as `captf-creds-` and keeps it in sync. Each Job that resolves to this identity gets that mirror both ways: - `envFrom`, so every key becomes an environment variable — the convention most provider SDKs (`AWS_*`, `ARM_*`, `GOOGLE_*`, …) read. A key starting with `TF_` or `KUBE_` never reaches the module’s environment, except `TF_IN_AUTOMATION`, `TF_INPUT` and `KUBE_NAMESPACE`, which the Job sets itself and which a Secret key cannot override; see [Runtime environment]() for why. - A read-only file mount at `/var/run/captf/credentials/`, one file per key, for file-based authentication (`GOOGLE_APPLICATION_CREDENTIALS`, a kubeconfig, a PEM key). The `TF_` and `KUBE_` filtering above applies only to the environment: a matching key still appears as a file. The runner ServiceAccount that a Job runs as can itself read every Secret in its namespace; the identity mechanism controls which credentials a Job carries, not what its ServiceAccount could otherwise reach — see [RBAC](). The mirror Secret itself, its lifecycle and its sensitivity are cataloged in [Secrets](); its labels and annotations are listed in [Annotations and labels](), and the `MirrorCreated`/`MirrorRemoved` events in [Events](). ## Create the credentials Secret Create a plain `Opaque` Secret with one key per credential your module’s providers read: ```yaml apiVersion: v1 kind: Secret metadata: name: namespace: type: Opaque stringData: AWS_ACCESS_KEY_ID: AWS_SECRET_ACCESS_KEY: ``` - `` — reused below as the identity’s own name; the two do not have to match, but it keeps the pair easy to find. - `` — any namespace the identity’s creator can read Secrets in. It is commonly a provider or platform namespace, separate from the tenant namespaces that use the identity. The keys are exactly what reaches every Job that resolves to this identity: name them for what your module’s providers expect, not for CAPTF. ## Create the identity ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: spec: secretRef: name: namespace: allowedNamespaces: list: - ``` `allowedNamespaces` decides which namespaces may use the identity: | Value | Meaning | | --- | --- | | unset | No namespace may use it | | `list` | Exactly the named namespaces | | `selector` | Namespaces whose labels match | | `selector: {}` (empty selector) | Every namespace, including one that does not exist yet | | `list` and `selector` both set | The union of the two | | `{}` (empty object) | Rejected: write `selector: {}` for every namespace instead | An empty selector reads as “every namespace” rather than “no namespace”, so `allowedNamespaces: {}` is rejected outright rather than treated as one or the other. `list` accepts up to 100 namespace names; `selector` follows the usual label selector rules. Full field details, including size limits, are in [TerraformClusterIdentity](). Creating the identity, and any later change to `spec.secretRef` or to `spec.allowedNamespaces`, widening or narrowing, runs a `SubjectAccessReview` for the requesting user to `get` the named Secret, and the request is rejected if they may not. This stops a role that may manage identities but not read Secrets from pointing one at a Secret it cannot read, or from granting more namespaces access to one, and having the controller mirror it somewhere that role can. Updates that change neither field, such as to labels or annotations, are not re-checked. ## Reference it An identity is used through `identityRef`, which names it by `name`: - **`TerraformCluster`**: `spec.identityRef` is required and mutable — changing it re-resolves the cluster’s credentials on its next reconcile. - **`TerraformCluster` defaults for machines and pools**: `spec.defaults.identityRef`, mutable, is the identity a machine or pool without its own `identityRef` uses; unset, they fall back further to the cluster’s own `spec.identityRef`. See [Kinds]() for how `spec.defaults` inheritance works, and [Templates and ClusterClass]() for a `TerraformClusterTemplate` that leaves `identityRef` for a ClusterClass patch. - **`TerraformMachine`**: `spec.identityRef` is optional, and immutable once the object is created — including from unset to a value. Change it by replacing the machine (a `MachineDeployment` or control-plane rollout), not by editing it in place. A machine created with no `identityRef` keeps falling back to its cluster’s current default or `spec.identityRef`, so changing those still changes which credentials such a machine uses. - **`TerraformMachinePool`**: `spec.identityRef` is optional and mutable, with the same fallback as a machine. `identityRef.name` must resolve to an identity that exists and allows the object’s namespace; the [Conditions]() reference has every reason `IdentityAllowed` and `CredentialsMirrored` can carry. ## Confirm it worked > [!TIP] > > ```sh > kubectl get terraformclusteridentity > kubectl describe terraformcluster -n > ``` > > The identity’s own `Ready` condition is `True`/`SecretFound` once its Secret exists, `False`/`SecretNotFound` otherwise, and `status.namespaces` lists every namespace currently holding a mirror of it. On the object that references it, `IdentityAllowed` and `CredentialsMirrored` turn `True` once the namespace is allowed and the mirror is in place; until then no Job for that object starts. ## Rotate credentials - **Edit the credentials Secret’s data in place**, keeping its name and namespace. The controller cannot watch that Secret, so this is not picked up at once: the mirror in each allowed namespace is rewritten the next time an object that uses the identity reconciles for any reason, and at the latest within one [`--sync-period`]() (10 minutes by default). A Job created after that reconcile gets the new values through its `envFrom`. A Job already running gets them too, but only in its file mount: the kubelet resyncs that volume from the mirror on its own schedule, even for a pod that started before the rotation; whether a long-running provider process re-reads a changed file is up to the provider. Only the environment is fixed for the life of the pod. A Job already running when you rotate keeps the old credentials in its environment for its whole run: the old credential must stay valid until every mirror has picked up the rotation and the longest `activeDeadlineSeconds` any in-flight Job could still run for has passed (default 3600 seconds; see [Deadlines and lock waits]()), not just until you edited the Secret. Every mirror carries a `captf.io/source-hash` annotation of the source Secret’s data at the time it was last written; it changes whenever the mirror is rewritten, so watching it move off the value it held before the edit confirms that namespace’s mirror has picked up the rotation, without comparing credential values directly: ```sh kubectl get secret captf-creds- -n \ -o jsonpath='{.metadata.annotations.captf\.io/source-hash}' ``` Repeat for every namespace `status.namespaces` lists, or diff the mirror’s `data` against the source Secret’s `data` directly if you want to confirm the values themselves rather than only that a rewrite happened. Only once every mirror has moved past its pre-rotation hash is the old credential safe to revoke. - **Point the identity at a different Secret**, by editing `spec.secretRef`. This is a change to the identity object itself, so it is watched: every object that uses the identity reconciles at once and refreshes its mirror. It also re-runs the `SubjectAccessReview`, so the user making the change needs `get` access to the new Secret. ## Revoke access Editing `spec.allowedNamespaces` to drop a namespace is a change to the identity object, so every object of that namespace using the identity reconciles at once: the controller starts no new Job for it, leaves any infrastructure it already created alone, and deletes the mirror Secret in that namespace, whatever objects still reference it. > [!WARNING] > > **Revoking a namespace also blocks destroy** > > Deleting such an object waits, reporting that the destroy is waiting for the identity to allow the namespace again, until access is restored. ## Delete an identity Deletion is refused while: - a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`, in any namespace, currently resolves to this identity through its own `identityRef` or a cluster’s fallback — an object being deleted still counts, since its destroy still needs the credentials; or - `status.namespaces` still lists a namespace holding a mirror of it. The rejection names one object still using it, or every namespace still holding a mirror. Switching an object’s `identityRef` away from this identity does not by itself release its mirror in that namespace: the mirror keeps that object as an owner until the object is deleted, so `status.namespaces` keeps listing the namespace, and the identity cannot be deleted, until it is. If nothing resolves to the identity there anymore, deleting the mirror Secret directly also releases the namespace: it is only a copy, so this does not touch the source Secret, and the identity’s status catches up as soon as the deletion is seen. > [!NOTE] > > **Deleting an identity keeps its Secret** > > Deleting an identity never deletes the Secret it names: that Secret belongs to you, and stays behind for you to remove or reuse. ## Move considerations `TerraformClusterIdentity` moves with `clusterctl move`, but its credentials Secret deliberately does not: see the [move runbook]() for why, and the procedure for copying it to the target cluster yourself. > [!NOTE] > > **See also** > > - [Security model]() — the trust boundary a mounted identity’s credentials sit inside. > - [Secrets]() — every Secret CAPTF reads or writes, including the credential mirror. > - [RBAC]() — the runner ServiceAccount that Jobs run as. > - [Kinds]() — how `spec.defaults` inheritance works. > - [Conditions]() — every `IdentityAllowed` and `CredentialsMirrored` reason. > - [Module Variables]() — a `variablesFrom` Secret is a different mechanism: it supplies module-specific values (sizes, CIDRs, passwords), never cloud credentials, and CAPTF only ever reads it, in place, rather than mirroring it like an identity’s Secret. # Module Variables This page shows you how to pass your own variables to a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`’s module: instance sizes, CIDRs, SSH keys, database passwords, anything besides the [contract inputs]() CAPTF sets itself. > [!NOTE] > > **Before you begin** > > - The module declares each variable you set, with a default so it still plans when nobody sets the variable ([`contract/v1alpha1/common.md`]()). A variable the module does not declare fails the apply. > - To pass a variable from a ConfigMap or Secret, you need permission to create and label one in the object’s namespace. ## Set an inline variable Add the variable to `spec.variables` (or `spec.template.spec.variables` on a `*Template` kind), a JSON object: ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: workers-large spec: template: spec: source: image: ghcr.io/example/aws-machine:v1.4.0 variables: instance_type: m6i.xlarge root_volume_gib: 100 subnet_ids: [subnet-0a1, subnet-0b2] ``` Each key becomes a named argument of the module: the generated root declares it without a type and passes it through, so the **module’s own declared type** converts the value (a ConfigMap string `"100"` becomes the number `100` for a `number` variable). ## Set a variable from a ConfigMap or Secret 1. Create the ConfigMap or Secret in the object’s namespace, and label it so CAPTF is allowed to read it: ```sh kubectl create configmap aws-sizing -n \ --from-literal=instance_type=m6i.xlarge kubectl label configmap aws-sizing -n captf.io/variables=true ``` `` is the object’s own namespace: `variablesFrom` never reaches across namespaces. This Secret is unrelated to the credentials Secret a `TerraformClusterIdentity` names: the two are separate mechanisms with separate labels and separate readers (see [Identities and Credentials]()). The manager watches only labeled ConfigMaps and Secrets, and reads their data straight from the API server when it renders a Job; it never caches it. The label is an explicit opt-in, part of the [trust boundary](): an object cannot read an arbitrary Secret of its namespace just by naming it. 2. Reference it from `spec.variablesFrom`: ```yaml spec: variablesFrom: - configMapRef: name: aws-sizing - secretRef: name: db-credentials format: JSON - configMapRef: name: team-overrides optional: true ``` Every data key of a labeled source becomes a variable of that name. Set `optional: true` on a source that may not exist yet: a missing or unlabeled required source blocks the object at [`DependenciesReady=False/VariablesSourceNotFound`]() and no Job starts; an optional one just contributes nothing. Every value, including a Secret’s, must be UTF-8; a ConfigMap’s `binaryData` is read the same as `data`. 3. Confirm: `DependenciesReady` turns `True` once every required source resolves, and the object’s next Job carries the variable. Inspect the rendered `terraform.tfvars.json` of a live object as described in [Job inputs](). ## Merge order and formats 1. `spec.variablesFrom`, in list order: every source’s keys become variables, and a later source wins on a key it shares with an earlier one. 2. `spec.variables` (inline) wins over every source. A source’s `format` decides how its values are read: - **`String`** (the default): each value is passed as a string. Use it for scalars; the module’s declared type converts `"3"`, `"true"` and `"0.5"`. - **`JSON`**: each value is parsed as JSON, for lists, maps and objects (`["a","b"]`, `{"team":"a"}`). A value that is not valid JSON holds the object at `DependenciesReady=False/VariablesInvalid`, naming the source and the key, never the value. ## Names and limits A variable name is a Terraform identifier: `^[a-zA-Z_][a-zA-Z0-9_-]*$`. Rejected, whatever the format: - a name starting with `captf_`; - a contract input of the object’s role — a machine cannot set `machine_name`, but a cluster can, since `machine_name` is not one of its inputs ([contract reference]()); - a module meta-argument: `source`, `version`, `providers`, `count`, `for_each`, `depends_on`, `lifecycle`, `locals`. `spec.variables` holds at most 256 keys, and when set, at least one; the admission webhook rejects a bad inline key or too many of them at once, and also rejects a `variablesFrom` entry that names both a ConfigMap and a Secret, or neither. A bad key in a referenced source is caught only when the controller reads it, as `VariablesInvalid` naming the key. `spec.variablesFrom` holds at most 16 sources. Variables count toward the rendered-inputs limit shared with every other input (1,000,000 bytes): past it the apply is refused with `ApplyJobSucceeded=False/InputsTooLarge` ([Size limits]()). They are also part of the [inputs hash]() that decides whether a mutable object’s next reconcile re-applies. ## Sensitivity A variable’s **winning** value (after the merge above) is declared `sensitive = true` in the generated root when it came from a Secret, so Terraform and OpenTofu redact it in plan and apply output. A value that ends up inline or from a ConfigMap is not sensitive, even if an earlier, losing source was a Secret, and even if the module’s own declaration marks its variable sensitive. > [!WARNING] > > **Sensitive variables are still stored** > > Sensitive or not, every variable is still stored like any other input: in the object’s inputs Secrets and in the state ([Secrets](), [Terraform state]()). ## What a change does per kind - `TerraformCluster`: the variables are part of the inputs hash. Editing `spec.variables`, `spec.variablesFrom` or the data of a referenced source re-applies the module, guarded like any other input change ([Plan approval]()). The controller re-reads every referenced source on each reconcile, and a change to a labeled source it references wakes it. - `TerraformMachine`: `spec.variables` and `spec.variablesFrom` are immutable. The controller reads the sources until the machine is provisioned; after that it never reads them again, and the machine’s destroy, refresh and drift runs use the variables already pinned in its durable inputs Secret. Editing a referenced ConfigMap or Secret therefore affects only machines created afterward: roll the MachineDeployment, or bump the template, to replace existing ones. - `TerraformMachinePool`: `spec.variables` and `spec.variablesFrom` are mutable, and the sources are re-read on every reconcile for the pool’s whole life, the same as a TerraformCluster; a change to a labeled source it references wakes it immediately, the same way. The variables are part of the inputs hash, but the resulting re-apply is never guarded: a TerraformMachinePool has no `applyPolicy` ([Machine pools]()). - Templates: `spec.template.spec` of a `*Template` kind, so its `variables` and `variablesFrom` too, is immutable like the rest of the template ([The Kinds]()); a ClusterClass topology patch on `spec.template.spec.variables` rolls out through a new template ([Templates and ClusterClass]()). - `spec.defaults` on a `TerraformCluster` does not cover variables: a machine or pool without its own `identityRef`, `jobs` or `drift` inherits the cluster’s, but variables are never inherited, so set them on the machine, pool or their templates directly. - `clusterctl move` never carries a `variablesFrom` source: CAPTF puts no owner reference on a ConfigMap or Secret it reads. A `TerraformCluster` or `TerraformMachinePool`, or a machine not yet provisioned, waits at `VariablesSourceNotFound` on the target until you recreate the source there ([clusterctl move]()). ## Troubleshooting | Symptom | Reason | Fix | | --- | --- | --- | | `DependenciesReady=False/VariablesSourceNotFound` | A required `variablesFrom` source is missing, or not labeled `captf.io/variables=true` | Create or label the ConfigMap or Secret, or set `optional: true` | | `DependenciesReady=False/VariablesInvalid` | A source’s key is not a Terraform identifier, is reserved, or (format `JSON`) its value is not valid JSON | Rename or remove the key, or fix the source’s `format` | | Admission rejects `spec.variables` | Not a JSON object, an inline key is reserved or not a Terraform identifier, or there are more than 256 keys | Fix or trim the key named in the error | | `ApplyJobSucceeded=False/ApplyFailed`, `status.lastRun.error.summary` has `Unsupported argument` or `Extraneous JSON object property` | The module does not declare a variable you set | Add the variable to the module, or stop setting it | See [`conditions.md`]() for every `DependenciesReady` reason. > [!NOTE] > > **See also** > > - [Identities and Credentials]() — the separate Secret mechanism for cloud credentials. > - [`common.md`]() — how a module declares a user variable. > - [Job inputs]() — the full render pipeline and the inputs hash. > - [Templates and ClusterClass]() — patching a variable per Cluster. > - [Plan approval]() — the destructive-plan guard a TerraformCluster variable change is subject to. > - [Common Fields]() for every `variables` and `variablesFrom` field. # Templates and ClusterClass This page covers the clusterctl templates CAPTF ships, how to generate a cluster from them, and the `noop` `ClusterClass`: what it creates, its topology variables, and how to patch a module variable onto it. It is for anyone creating clusters with CAPTF, especially with `ClusterClass`. > [!NOTE] > > **Before you begin** > > - The provider is installed and registered with `clusterctl` ([Installation]()). > - A `TerraformClusterIdentity` is applied for the namespace you are generating into ([Identities and credentials]()). > - Your cluster and machine module images are built, linted and pushed ([tfcapi-lint](), [image contract]()). > - You know the values for the `clusterctl` variables you need ([clusterctl variables]()). ## The shipped flavors CAPTF ships two flavors as `templates/cluster-template*.yaml` files; once the provider is registered, `clusterctl generate cluster --infrastructure terraform` renders one of them from the registered repository ([Installation]()): | Flavor | Template file | What it creates | | --- | --- | --- | | Default (no `--flavor`) | `cluster-template.yaml` | A `Cluster`, `TerraformCluster`, `KubeadmControlPlane`, two `TerraformMachineTemplate`s (control plane and workers), a `MachineDeployment` with its `KubeadmConfigTemplate`, and a `MachineHealthCheck` for each of the control plane and the workers | | `clusterclass` | `cluster-template-clusterclass.yaml` | A `Cluster` with `spec.topology.classRef.name: noop`, referencing the `noop` `ClusterClass` | Every flavor’s `KubeadmControlPlane` sets `spec.remediation.maxRetry: 3` and `retryPeriodSeconds: 300`, so a module that fails deterministically stops being retried instead of churning cloud resources, and its `localAPIEndpoint.bindPort` (init and join) equals `Cluster.spec.clusterNetwork.apiServerPort` (6443); kubeadm never reads the Cluster field, so keep the two equal if you change one. ## Generate a cluster Set the identity, image and sizing variables, then generate and apply. Pick the flavor: ```sh export TERRAFORM_IDENTITY_NAME= \ TERRAFORM_CLUSTER_IMAGE= \ TERRAFORM_MACHINE_IMAGE= clusterctl generate cluster --infrastructure terraform \ --target-namespace \ --kubernetes-version \ --control-plane-machine-count --worker-machine-count \ | kubectl apply -f - ``` Enable the `ClusterTopology` feature gate when you register the provider ([Installation]()); it is alpha and off by default. Apply the class to the namespace once, then generate with `--flavor clusterclass`: ```sh kubectl apply -n -f templates/clusterclass-noop.yaml clusterctl generate cluster --infrastructure terraform \ --flavor clusterclass --target-namespace \ --kubernetes-version \ --control-plane-machine-count --worker-machine-count \ | kubectl apply -f - ``` `clusterclass-noop.yaml` carries no `clusterctl` variables of its own and no namespace, so apply it once to every namespace that generates clusters from it. Neither flavor defaults `CLUSTER_NAME`, `KUBERNETES_VERSION`, `CONTROL_PLANE_MACHINE_COUNT`, `WORKER_MACHINE_COUNT`, the two image variables or the identity name: a missing one fails `clusterctl generate` instead of deploying something unintended. See [clusterctl variables]() for the full list. ## The noop ClusterClass `templates/clusterclass-noop.yaml` defines `ClusterClass noop` and the templates it references: - `TerraformClusterTemplate/noop`, the infrastructure template. - `KubeadmControlPlaneTemplate/noop-control-plane` and `TerraformMachineTemplate/noop-control-plane`, the control plane and its infrastructure. - `TerraformMachineTemplate/noop-worker` and `KubeadmConfigTemplate/noop-worker`, the infrastructure and bootstrap for the `default-worker` machine deployment class. Each of the three Terraform templates carries the placeholder image `example.invalid/captf/set-by-clusterclass:unset`. > [!NOTE] > > **An unpatched image fails the pull** > > A Cluster that fails to patch a real image gets a failed image pull instead of running an unintended module. The class also carries default `MachineHealthCheck` timeouts for the control plane and the workers; the `clusterclass` flavor overrides them through `Cluster.spec.topology` with the same variables as `cluster-template.yaml` (see [clusterctl variables]()). The class declares three required string topology variables — `identityName`, `clusterImage` and `machineImage` — described in [the ClusterClass topology variables table](). Three patches apply them to the templates above: - `identity` sets `spec.template.spec.identityRef` on `TerraformClusterTemplate/noop` to `identityName`, and also sets `spec.template.spec.defaults.identityRef` to the same value, so machines without their own `identityRef` inherit it. - `clusterImage` replaces `spec.template.spec.source.image` on `TerraformClusterTemplate/noop` with `clusterImage`. - `machineImage` replaces `spec.template.spec.source.image` on both `TerraformMachineTemplate/noop-control-plane` and `TerraformMachineTemplate/noop-worker` with `machineImage`, matched by `matchResources.controlPlane` and `matchResources.machineDeploymentClass.names: [default-worker]`. `cluster-template-clusterclass.yaml` sets the three variables from `TERRAFORM_IDENTITY_NAME`, `TERRAFORM_CLUSTER_IMAGE` and `TERRAFORM_MACHINE_IMAGE`. ## Patch a module variable through ClusterClass A `ClusterClass` patch can also turn a topology variable into a module variable (`spec.template.spec.variables`; [module variables]()). Declare the topology variable, then add a patch whose `jsonPatches` write `/spec/template/spec/variables` (or one key under it, if the template already sets others): ```yaml spec: variables: - name: workerInstanceType required: true schema: openAPIV3Schema: type: string patches: - name: workerInstanceType definitions: - selector: apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate matchResources: machineDeploymentClass: names: [default-worker] jsonPatches: - op: add path: /spec/template/spec/variables valueFrom: template: | instance_type: {{ .workerInstanceType }} ``` The same patch on `TerraformClusterTemplate` reaches the `TerraformCluster` instead, whose `spec.variables` is mutable: changing the variable re-applies the cluster module, and a destructive plan still waits for approval ([plan approval]()). A `ConfigMap` or `Secret` named in `variablesFrom` is not part of the `ClusterClass`: create it, labeled `captf.io/variables=true`, in each namespace that uses the class ([module variables]()). ## Template immutability and rolling out a change > [!WARNING] > > **A direct edit to a template is rejected** > > A `*Template` kind’s `spec.template.spec` cannot be edited in place once created; only its metadata can change (see [mutability per kind]()). `ClusterClass` topology reconciliation is exempt from that rule for its own dry-runs, which is how patching works at all, but a direct edit to a `*Template` object is rejected. Two situations follow from this: - **A field the class exposes as a topology variable.** Changing the variable’s value on a Cluster’s `spec.topology.variables` is enough: since the resulting `TerraformMachineTemplate` would differ from the one already referenced, and templates are immutable, the topology controller creates a new `TerraformMachineTemplate` with the patched spec and rolls the affected `MachineDeployment` or control plane onto it. Changing `machineImage` on a `clusterclass`-flavor Cluster works this way. - **A field the class does not expose as a variable**, such as a `KubeadmControlPlaneTemplate` field or a new patch. Create a new template object under a new name with the desired `spec.template.spec`, then update the `ClusterClass`’s reference to it — `spec.infrastructure.templateRef`, `spec.controlPlane.templateRef`, `spec.controlPlane.machineInfrastructure.templateRef`, or the matching `workers.machineDeployments[].infrastructure.templateRef` / `bootstrap.templateRef` — and apply the class. Topology reconciliation then rolls every Cluster on the class onto the new template, subject to each machine deployment’s or control plane’s own rollout strategy (the `noop` class’s control plane sets `maxSurge: 0`; its `default-worker` machine deployment class leaves the rollout strategy unset, so the default `maxSurge: 1` applies). A `TerraformCluster` itself is mostly mutable — only `spec.controlPlaneEndpoint` is immutable once it has a host — so a `clusterImage` or `identityName` change re-applies the cluster module in place rather than creating a new object. > [!NOTE] > > **See also** > > - [Identities and credentials]() > - [Module variables]() > - [Plan approval]() > - [The kinds]() > - [clusterctl variables]() # Machine Pools This page shows you how to create a `MachinePool` backed by a `TerraformMachinePool`: a group of nodes CAPTF provisions and scales as one native cloud scaling group (an autoscaling group, a scale set, an instance group, or similar), rather than as individual `TerraformMachine`s. It covers fixed and autoscaled replica counts, the group’s reported members, and deletion. See [The Kinds]() for how a `TerraformMachinePool` compares to a `TerraformMachine`, and the [machinepool contract]() for everything the module role must implement. > [!NOTE] > > **Before you begin** > > - The provider installed, and a `TerraformCluster` provisioned or being provisioned (see [Installation]()). > - A `TerraformClusterIdentity` allowed in your namespace (see [Identities and Credentials]()). > - A machinepool-role module image: one that implements the [machinepool contract](), managing one scaling group and reporting its provider IDs, desired capacity and members. ## Create a MachinePool A `MachinePool` has a single infrastructure object for its whole group, not one per member: `MachinePool.spec.template.spec.infrastructureRef` names one `TerraformMachinePool` directly, by kind and name, the same way `Cluster.spec.infrastructureRef` names one `TerraformCluster`. There is no per-replica cloning, so you create the `TerraformMachinePool` yourself rather than pointing at a `TerraformMachinePoolTemplate`; a `TerraformMachinePoolTemplate` exists only for a `MachinePool` a ClusterClass topology manages, covered in [Templates and ClusterClass](). The `TerraformMachinePool` must carry the cluster’s `cluster.x-k8s.io/cluster-name` label: CAPTF looks up the owning `Cluster` by that label, not by the `MachinePool`’s `spec.clusterName`. Cluster API’s `MachinePool` controller patches this label on from `spec.clusterName` once the `TerraformMachinePool` exists, but only on its own next reconcile; setting the label yourself avoids that wait. For example: ```yaml apiVersion: cluster.x-k8s.io/v1beta2 kind: MachinePool metadata: name: my-cluster-workers namespace: team-a spec: clusterName: my-cluster replicas: 3 template: spec: clusterName: my-cluster version: v1.31.4 bootstrap: configRef: apiGroup: bootstrap.cluster.x-k8s.io kind: KubeadmConfig name: my-cluster-workers infrastructureRef: apiGroup: infrastructure.cluster.x-k8s.io kind: TerraformMachinePool name: my-cluster-workers --- apiVersion: bootstrap.cluster.x-k8s.io/v1beta2 kind: KubeadmConfig metadata: name: my-cluster-workers namespace: team-a spec: joinConfiguration: {} --- apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: my-cluster-workers namespace: team-a labels: cluster.x-k8s.io/cluster-name: my-cluster spec: source: image: ghcr.io/example/machinepool-module:v0.1.0 identityRef: name: aws-prod ``` The `KubeadmConfig` is the bootstrap provider’s own object (one per pool, for the same reason: a pool has no per-member Machine to hold one each); its content is the bootstrap provider’s concern, not CAPTF’s. With `spec.replicas` unset, Cluster API defaults it to 1; set it to the fixed size you want, as above. Every field of `TerraformMachinePoolSpec` is mutable: changing `spec.source`, `spec.identityRef`, `spec.variables` or any other field re-applies the module on the next reconcile, and there is no `spec.applyPolicy`. Only an apply that renders a change of the cluster’s exports is guarded (see [When the cluster’s exports change](<#when-the-clusters-exports-change>) and [The Kinds]()). Pass module-specific configuration through `spec.variables` or `spec.variablesFrom` as for any other kind (see [Module Variables]()); `spec.jobs` tunes the Job the same way it does for a `TerraformMachine` (see [Tuning Jobs]()). A pool with none of its own inherits `spec.identityRef`, `spec.jobs` and `spec.drift.intervalSeconds` from the owning `TerraformCluster`’s `spec.defaults`. `node_labels` (rendered from `spec.template.metadata.labels`, not the `TerraformMachinePool`’s own metadata) and the other inputs a module sees are listed in the [machinepool contract]() and the [environment reference](); this page does not restate them. ## Choose fixed replicas or autoscaling With no autoscaler annotations, `MachinePool.spec.replicas` is the sole source of desired capacity, exactly as in the example above: change it to resize the group. To let the module’s own native autoscaling policy own the desired count instead, set both annotations Cluster API’s autoscaler contract defines, on the `MachinePool`: ```yaml apiVersion: cluster.x-k8s.io/v1beta2 kind: MachinePool metadata: name: my-cluster-workers namespace: team-a annotations: cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "3" cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "10" spec: clusterName: my-cluster template: spec: clusterName: my-cluster version: v1.31.4 bootstrap: configRef: apiGroup: bootstrap.cluster.x-k8s.io kind: KubeadmConfig name: my-cluster-workers infrastructureRef: apiGroup: infrastructure.cluster.x-k8s.io kind: TerraformMachinePool name: my-cluster-workers ``` Both annotations must be present, parse as non-negative integers, and satisfy min ≤ max, or the pool reports [`AutoscalingActive=False/AutoscalingAnnotationsInvalid`]() and applies without autoscaling. Valid, the pool reports `AutoscalingActive=True/ReplicasManagedByModule`: the module owns the group’s desired count and its own scaling policy (target tracking, scheduled, or whatever it implements), and the controller claims the `cluster.x-k8s.io/replicas-managed-by` annotation on the `MachinePool` so Cluster API stops treating `spec.replicas` as authoritative. On every reconcile the controller then writes the group’s observed desired capacity back to `MachinePool.spec.replicas`, emitting a [`ReplicasWrittenBack`]() event when it changes; see [Annotations, Labels and Finalizers]() for both keys. Leave `spec.replicas` unset in this mode: Cluster API defaults and clamps it from the annotations for the first apply, and the write-back takes over from there. Removing both annotations returns `spec.replicas` to being authoritative and releases `replicas-managed-by`. If another controller already owns `cluster.x-k8s.io/replicas-managed-by` on the `MachinePool` (its value is something other than the one CAPTF claims), CAPTF leaves the annotation alone and stops writing observed replicas back. The pool reports `AutoscalingActive=False`/`ReplicasManagedExternally`, with a `Warning` event, and `spec.replicas` stays under that other controller’s control. See [`AutoscalingActive`](). > [!WARNING] > > **The Kubernetes Cluster Autoscaler does not drive these pools** > > Its `clusterapi` cloud provider requires MachinePool Machines, which CAPTF does not implement. Running it against a CAPTF pool is unsupported, since its `spec.replicas` patches would be overwritten by the write-back above. ## Add an autoscaled pool to a generated cluster This walks through adding an autoscaled `MachinePool` to a cluster generated from the default flavor ([Templates and ClusterClass]()), since none of the shipped flavors creates one on their own. Generate the default flavor with no workers; a `MachineDeployment` is still created, with `spec.replicas: 0`, rather than omitted: ```sh clusterctl generate cluster my-cluster --infrastructure terraform \ --target-namespace team-a \ --kubernetes-version v1.31.4 \ --control-plane-machine-count 1 --worker-machine-count 0 \ | kubectl apply -f - ``` Add the pool alongside it: a `MachinePool` with the autoscaler annotations, its `KubeadmConfig` (the flavor’s bootstrap provider is kubeadm), and the `TerraformMachinePool`: ```yaml apiVersion: cluster.x-k8s.io/v1beta2 kind: MachinePool metadata: name: my-cluster-workers namespace: team-a annotations: cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "2" cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "5" spec: clusterName: my-cluster template: spec: clusterName: my-cluster version: v1.31.4 bootstrap: configRef: apiGroup: bootstrap.cluster.x-k8s.io kind: KubeadmConfig name: my-cluster-workers infrastructureRef: apiGroup: infrastructure.cluster.x-k8s.io kind: TerraformMachinePool name: my-cluster-workers --- apiVersion: bootstrap.cluster.x-k8s.io/v1beta2 kind: KubeadmConfig metadata: name: my-cluster-workers namespace: team-a spec: joinConfiguration: {} --- apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: my-cluster-workers namespace: team-a labels: cluster.x-k8s.io/cluster-name: my-cluster spec: source: image: ghcr.io/example/machinepool-module:v0.1.0 identityRef: name: aws-prod ``` `spec.replicas` is left unset on the `MachinePool`: with both autoscaler annotations present and valid, Cluster API defaults and clamps it from `min-size`/`max-size` for the first apply, and the pool’s own write-back takes over from there (see [Choose fixed replicas or autoscaling](<#choose-fixed-replicas-or-autoscaling>) above). Apply the three objects, then confirm as in [Confirm it worked](<#confirm-it-worked>) below. ## Set the membership refresh interval Between applies, the controller runs a refresh to pick up members joining or leaving the group, on `spec.membershipRefreshIntervalSeconds` (15–86400 seconds; unset or 0 means 60): ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool spec: membershipRefreshIntervalSeconds: 30 ``` A new member is unschedulable until its provider ID reaches `spec.providerIDList`, so a shorter interval gets new nodes ready for workloads sooner, at the cost of more frequent Jobs. The controller also refreshes right after every apply and, while the group has not converged (`spec.providerIDList`’s length differs from `status.replicas`), every 30 seconds, or the configured interval instead when it is shorter. ## Check the group’s members ```sh kubectl get terraformmachinepool -n ``` shows the owning cluster, `MachinePool`, desired replicas and the `Ready` condition. `spec.providerIDList` is every non-terminated member, sorted and deduplicated; `status.replicas` is the desired capacity as of the last refresh; `status.instances` is the module’s own per-member detail (provider ID, an optional instance ID, addresses, failure domain and health state), capped at 1000 entries: ```sh kubectl get terraformmachinepool -n \ -o jsonpath='{.status.instances}' ``` See [TerraformMachinePool]() for every status field. A pool reports provisioned, and the `Ready` condition (mirrored onto the `MachinePool`’s `InfrastructureReady`) true, once its state carries a successful apply and its module reports a health state other than pending; unlike a fixed-replica machine, that latch does not depend on `spec.providerIDList` being non-empty, since a pool may legitimately scale to zero. ## Drift on a pool > [!NOTE] > > **A pool’s drift check cannot be disabled** > > `spec.drift.intervalSeconds: 0` falls back to the manager’s default interval rather than turning checks off. The reason is that the drift Job’s own refresh is what feeds a plan; without it, a cloud-side scaling change would never register as drift. With `spec.drift.action: Remediate`, a detected difference re-applies the pool’s current inputs; with autoscaling enabled the module is responsible for excluding its own desired-count attribute from that plan, or every cloud-side scale reports as drift. See [Drift]() for setting the interval and action, and [Drift and Health]() for how a check runs and feeds health. ## When the cluster’s exports change A pool’s inputs include the cluster’s exports (`captf_cluster_outputs`). When they change, the pool applies the change, and that apply is guarded: if its plan deletes or replaces anything, it stops before the apply step and the change is **held**. Nothing else is guarded: the first apply, bootstrap rotations, version rolls, replica changes and spec edits apply as usual while the exports are unchanged. While a change is held, the pool keeps applying everything else with the exports of its last successful apply, and `ApplyJobSucceeded` shows `False`/`DestructivePlanBlocked` (so the pool’s `Ready` is `False`) with what the plan would delete or replace and the pool’s approval hash. After you have read the plan, approve it: ```sh kubectl annotate terraformmachinepool -n \ captf.io/approve-destructive-plan= --overwrite ``` The approval hash is the inputs hash without `bootstrap_data`, so it survives bootstrap rotations and changes on any other input change. Anyone who may patch the pool may approve it. The annotation is removed after the apply it approved succeeds. If the exports return to the applied ones, the change is withdrawn and its approval removed; if they move to another change, the old approval is removed. After a guarded apply fails part-way, every apply of the pool is guarded until one succeeds. See [The destructive-plan guard]() for the full behavior and its limits. ## Delete a MachinePool Deleting the `MachinePool` deletes its `KubeadmConfig` and `TerraformMachinePool` with it. Nothing blocks a `TerraformMachinePool`’s deletion: it destroys the group from its durable inputs and removes its finalizer once the destroy Job succeeds, whether or not the owning `MachinePool` or `Cluster` still exist. See [the reconcile lifecycle]() for how deletion and finalizers work across every kind. ## Confirm it worked > [!TIP] > > ```sh > kubectl get terraformmachinepool -n > ``` > > `Ready` reads `True` once the group is provisioned, and `Replicas` shows the desired capacity. `kubectl get machinepool -n ` shows the same replica count and `InfrastructureReady=True` once Cluster API has copied `spec.providerIDList` and `status.replicas` across, which happens only once the workload cluster is reachable. > [!NOTE] > > **See also** > > - [The Kinds]() for how a `TerraformMachinePool`’s mutability differs from a `TerraformMachine`’s. > - [The machinepool contract]() for every input and output a module role must implement. > - [Drift]() and [Drift and Health](). > - [Templates and ClusterClass]() for a `TerraformMachinePoolTemplate` used through a ClusterClass topology. > - [Module Variables]() and [Tuning Jobs](). > - [TerraformMachinePool](), [Conditions](), [Events]() and [Annotations, Labels and Finalizers](). # Tuning Jobs Every operation CAPTF runs — apply, destroy, drift, refresh, restore and plan — is one Kubernetes Job. `spec.jobs` on a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` tunes that Job: its resources, deadlines, history, environment, pull secrets, ServiceAccount and security contexts. Every field is optional and has a built-in default, so you only set the fields you need to change. For the full field list, types and validation see [`spec.jobs` in Common Fields](); for the Job’s exact shape (containers, mounts, args) see [Job Environment](). > [!NOTE] > > **Before you begin** > > - Know which object’s `spec.jobs` you are changing. A `TerraformCluster` also has `spec.defaults.jobs`, which its `TerraformMachine`s and `TerraformMachinePool`s inherit; see [Inheriting from a cluster’s defaults](<#inheriting-from-a-clusters-defaults>) below. > - A `TerraformCluster`’s own `spec.jobs` does not inherit `spec.defaults.jobs`: defaults are for its machines and pools, never for the cluster itself. ## Resources `spec.jobs.resources` sets the main container’s `resources` as a whole; when unset the controller applies its own default (250m CPU / 512Mi memory requested, 2Gi memory limit, deliberately no CPU limit — throttling a slow apply is worse than a slow apply). Raise it when a module pulls large provider plugins or holds a large plan in memory; the init container that copies the runner binary is not configurable, since it never varies with the module. See [Default resources]() for the exact values. ## Deadlines and lock waits - `activeDeadlineSeconds` bounds the whole Job, at most one day. Defaults to 3600 (one hour). Raise it for a module with a slow apply or many resources; a Job that hits its deadline is interrupted with SIGTERM, the same as a pod eviction or deletion. For an apply, destroy or plan Job, `ApplyJobSucceeded` is set False with reason `JobDeadlineExceeded`; for a drift or refresh Job, `DriftJobSucceeded` is set False with reason `DriftJobDeadlineExceeded`; for a restore Job, `RestoreJobSucceeded` is set False with reason `RestoreFailed`. A Job killed by its deadline is a failure: it counts toward retry backoff and the failed limit like any other failed Job. See [Conditions](). - `lockTimeoutSeconds` is passed to the runtime as `-lock-timeout`. Defaults to 300 (five minutes). Raise it when Jobs commonly queue behind each other on the same state lock. - `lockTimeoutSeconds` must be less than `activeDeadlineSeconds`, including inherited and built-in defaults: the webhook rejects a policy where a lock wait alone could fill the whole deadline. A policy that sets only one of the two is checked against the built-in default of the other (300 seconds for the lock wait, 3600 for the deadline), so `lockTimeoutSeconds: 4000` alone is rejected. The webhook sees one policy at a time, so it cannot see a machine’s or pool’s field-wise merge with a cluster’s `spec.defaults.jobs`. Reconcile checks the merged policy instead: an inconsistent result reports `ApplyJobSucceeded=False`/`JobPolicyInvalid` and no Job runs until you fix one of the two values. ## History limits `successfulJobsHistoryLimit` and `failedJobsHistoryLimit` cap how many finished Jobs of each kind are kept, both defaulting to 3. Each is scoped per object *and* per operation: apply, destroy, drift, refresh, restore and plan are pruned independently, so lowering one does not shrink another’s history. The newest Job of an operation is always kept, even at a limit of 0 — for a failed operation, until a newer Job of the same operation succeeds. `failedJobsHistoryLimit` also bounds retry backoff, since backoff looks at the same retained failures; see [The reconcile lifecycle]() for how backoff is computed. ## Environment variables `spec.jobs.env` adds environment variables to the main container. Entries named `TF_*` or `KUBE_*` are reserved for the runner and the Job’s own environment (`TF_IN_AUTOMATION`, `KUBE_NAMESPACE`, and so on): See [`spec.jobs.env` rejected names]() for the full list of names the Job already sets. > [!WARNING] > > **A reserved environment name is silently dropped** > > An entry using a `TF_*` or `KUBE_*` name is dropped instead of applied, with no error. ## Image pull secrets `spec.jobs.imagePullSecrets` covers both the role image (`spec.source.image`) and the runner’s own init image, so one list is enough even when they come from different registries. ## ServiceAccount `spec.jobs.serviceAccountName` overrides the runner ServiceAccount. Left unset, the controller creates and uses `captf-runner`, bound to the static `captf-runner` ClusterRole. An override must already exist and carry the label `captf.io/runner=true`; without it, no Job is created and `RunnerRBACReady` reports False with reason `ServiceAccountNotOptedIn` (see [Conditions]()). Setting up and opting in your own ServiceAccount, and the RBAC CAPTF manages around it, is covered in [RBAC](). ## Security contexts `spec.jobs.securityContext` sets the main container’s `securityContext`. The container holds cloud credentials, so the webhook rejects `privileged: true`, `allowPrivilegeEscalation: true`, any `capabilities.add`, `readOnlyRootFilesystem: false`, a `seccompProfile` of `Unconfined`, `procMount: Unmasked`, `windowsOptions.hostProcess` and an explicit `runAsUser: 0` or `runAsNonRoot: false` on it, whatever else the policy sets. Every other field you set is applied on top of the controller’s defaults (no privilege escalation, every capability dropped, a read-only root filesystem). `capabilities` always drops `ALL`, even when you supply a `capabilities` object. > [!WARNING] > > **An image that runs as root by default still runs as root** > > The webhook rejects only an explicit root setting; the Job sets neither `runAsUser` nor `runAsNonRoot`, so an image that runs as root by default is still admitted and runs as root. `spec.jobs.podSecurityContext` sets the Job pod’s `securityContext`. The controller defaults its `seccompProfile` to `RuntimeDefault` and its `fsGroup` to the runner’s UID, so a non-root image user can read the identity credential files (mode 0440) through that group; set your own `fsGroup` to override it, or your own `runAsNonRoot`/`runAsUser` to run the main container as a specific non-root user. For the trust boundary these defaults protect — why the Job is treated like any other pod that holds cloud credentials — see [The security model](). > [!NOTE] > > **Rules apply on create and on policy change** > > The `lockTimeoutSeconds` and `securityContext` rules on this page apply on create and whenever the jobs policy changes. An existing object with an older, weaker policy still accepts unrelated updates and can always be deleted. ## Inheriting from a cluster’s defaults A `TerraformMachine`’s or `TerraformMachinePool`’s `spec.jobs` is merged field by field with its `TerraformCluster`’s `spec.defaults.jobs`: a field the machine or pool sets wins, an unset one falls back to the cluster’s default, and a field neither sets gets the built-in default above. Two fields merge instead of falling back as a whole: - `env` is merged by name — the machine’s or pool’s entries first, then any of the cluster default’s entries whose name they do not already use. - `imagePullSecrets` is the union of both lists, without duplicates, the machine’s or pool’s first. `resources`, `securityContext` and `podSecurityContext` are each replaced as a whole: setting any one of them on the machine or pool drops the cluster default’s value for that field entirely, rather than merging individual keys inside it. See [What a cluster passes to its machines and pools]() for how this fits the rest of `spec.defaults`. Defaults are resolved at reconcile time and never persisted, so raising a cluster’s `spec.defaults.jobs` reaches every existing machine and pool on their next reconcile. ## Confirm it worked > [!TIP] > > ```sh > kubectl get job -n -l captf.infrastructure.cluster.x-k8s.io/owner-name= > kubectl get job -n -o jsonpath='{.spec.activeDeadlineSeconds}{"\n"}{.spec.template.spec.serviceAccountName}{"\n"}' > ``` > > `` and `` are the namespace and name of the `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` you changed; `` is one Job name from the first command’s output. Compare the Job’s `spec.template.spec.containers[0].resources`, `securityContext` and `env` against what you set; a field you expected to change but that still shows the built-in default usually means it was set on the wrong object, or dropped by the merge rules above. > [!NOTE] > > **See also** > > - [Other manual actions]() for `JobPolicyInvalid` and the other waits that need a person. > - [Common Fields]() for every `spec.jobs` field, its default and its validation. > - [Job Environment]() for the Job’s fixed fields, mounts, security contexts and runner args. > - [RBAC]() for the runner ServiceAccount, its RoleBinding and the opt-in sweep. > - [The security model]() for the trust boundary a Job’s security context protects. > - [The reconcile lifecycle]() for retry backoff and how a Job’s operation is chosen. # Drift and Health This page explains how CAPTF checks a provisioned `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` against reality: the drift schedule, what the module’s health output means, and how both feed the `InfrastructureHealthy` and `Ready` conditions and, for machines, Cluster API remediation. It is for anyone who wants to understand the behavior before changing it. To configure drift checking, see [Drift](); to configure machine remediation, see [Machine Remediation](); for the full reason tables, see [Conditions](). ## What runs, and when Two kinds of Job read an object’s infrastructure after it is provisioned: - A **refresh Job** runs `apply -refresh-only`: it updates the state from reality and produces a fresh health reading, but plans nothing and finds no drift. - A **drift Job** does the same refresh, then plans with `-refresh=false`: a plan with no changes means no drift, and any add, change or destroy in the plan is drift. A cluster or pool plans against its freshly rendered current inputs (falling back to the durable inputs Secret only when the current ones can’t be built); a machine, being immutable, always plans against its durable inputs. See [Job inputs](). Both count as one health sample once the object is provisioned, and both set the `DriftJobSucceeded` condition to report whether the Job succeeded; only a drift Job sets `DriftDetected`. A failed refresh or drift Job is retried under the reconciler’s normal backoff; see [Reconcile flow]() for how retries and requeues work in general. A drift check runs on a schedule: `spec.drift.intervalSeconds`, inherited from the `TerraformCluster`’s `spec.defaults.drift` for a machine or pool that sets none, else the manager’s `--drift-default-interval` (30 minutes by default; see [Manager flags]()). The first check after provisioning is due one interval after the last successful apply, not immediately; each object’s schedule is jittered deterministically by its UID so that objects created together don’t all check at once. `TerraformCluster` and `TerraformMachine` both let `spec.drift.intervalSeconds: 0` turn drift checks off entirely, which normally also stops routine health sampling after provisioning, since no more refresh or drift Jobs run on a schedule (a reading of `InstancePending` still gets occasional extra refreshes on its own, below). A `TerraformMachinePool`’s own interval must be at least 1 second — the CRD rejects `0` outright — and even a cluster’s `spec.defaults.drift.intervalSeconds: 0`, which does disable a machine’s drift, leaves a pool’s on the manager’s default instead: a pool’s periodic drift Job is what feeds its refreshed instance count into a plan, so disabling it would stop that count from ever reaching a plan; see [Machine pools](). Outside the drift schedule, a `TerraformMachine` or `TerraformMachinePool` also takes a refresh right after each successful apply, to get an immediate health reading (unless the apply’s own outputs already gave a definite one); a `TerraformMachinePool` refreshes again on its own `membershipRefreshIntervalSeconds` cadence to keep membership current regardless of drift (see [Machine pools]()), and, with `remediation.annotateMachine` set, a machine refreshes on `remediation.healthCheckIntervalSeconds` independent of drift too (see [Machine Remediation]()); none of these apply to a `TerraformCluster`. For a `TerraformCluster` or `TerraformMachine`, while a reading is `InstancePending`, the next refresh backs off on its own doubling schedule (30 seconds up to 5 minutes) instead of waiting for the next drift interval, so a newly launched instance is checked again quickly. A `TerraformMachinePool`’s pending reading instead refreshes every 30 seconds flat, the same fixed cadence its membership convergence uses, since the two converge together; see [Machine pools](). ## Report or remediate Drift found by a drift Job is either just recorded or automatically corrected, per the object’s drift action, `Report` or `Remediate`: - **`Report`** (the default for every kind) only sets `DriftDetected`; it records the finding but applies nothing. - **`Remediate`** additionally re-applies the object’s current inputs to remove the drift. Only a `TerraformCluster`’s own `spec.drift.action` and a `TerraformMachinePool`’s own `spec.drift.action` (never inherited from the cluster’s `spec.defaults.drift`, which has no action field at all) can select `Remediate`. A `TerraformMachine`’s drift policy has no action field either, and is always `Report`: a machine’s instance is immutable infrastructure, replaced by a Cluster API rollout, not reconciled in place by a re-apply. When a pool or cluster remediates, a finding sets `DriftDetected` True/`DriftPending` rather than True/`DriftReported`, and the reconciler starts an apply of the current inputs. While that apply runs, `DriftDetected` reads True/`DriftRemediating`; if the apply succeeds after the drift check that found the drift, `DriftDetected` clears to False/`NoDrift` at once, without waiting for the next drift check. If the apply instead fails, is blocked, or stops because its approved plan changed, `DriftDetected` falls back to True/`DriftPending` with a note of what happened, and the reconciler keeps retrying the remediation apply (under the normal backoff) until as many attempts have failed since the last successful drift check as the object’s failed-Job history keeps (see [Job tuning]()); past that, the drift stays pending until the next successful check re-evaluates it. A remediation apply is still an apply of that kind, with the same guard as any other: - A **`TerraformCluster`**’s remediation apply goes through the same destructive-plan guard as any other cluster apply — a plan that would delete or replace a resource blocks (`ApplyJobSucceeded=False`/ `DestructivePlanBlocked`) until it is approved — or, under `spec.applyPolicy: Manual`, the plan-approval flow instead of the guard. See [Plan preview and approval](). - A **`TerraformMachinePool`**’s remediation apply is not guarded: a pool has no `applyPolicy` and no destructive-plan guard, so it just runs. ## DriftDetected > [!NOTE] > > **Drift alone never makes an object un-Ready** > > `DriftDetected` is negative polarity: True means drift was found, and it is never an input to `Ready`. Before the first drift check it reads Unknown/`DriftNotChecked`. Its full reason table, including the messages each reason carries, is in [Conditions](). ## From module health to `InfrastructureHealthy` Every module role (`cluster`, `machine`, `machinepool`) declares a `health` output with a `state` (`pending`, `running`, `degraded`, `stopped`, `terminated` or `unknown`) and a `healthy` boolean, plus optional `message` and `reasons`. Each refresh or drift Job’s outcome — and, for a machine or pool, the reading an apply itself provides — maps to `InfrastructureHealthy`: | Module `health.state` | `healthy` | `InfrastructureHealthy` | | --- | --- | --- | | (no apply has started yet) | — | Unknown/`WaitingForProvisioning` | | (before provisioned, since the first apply started) | — | False/`Provisioning` | | `pending` | any | False/`InstancePending` | | `running` | `true` | True/`Healthy` | | `running` | `false` | False/`InstanceUnhealthy` | | `degraded` | any | False/`InstanceDegraded` | | `stopped` | any | False/`InstanceStopped` | | `terminated` | any | False/`InstanceTerminated` | | `unknown`, or no health output at all | — | Unknown/`HealthUnknown` | `message` and `reasons`, when the module sets them, are joined into the condition’s message. The full reason table is in [Conditions](). For a `TerraformMachine` specifically, a `provider_id` output that turns null (or empty) after provisioning is read in two steps. The first missing sample sets `Unknown`/`ProviderIDMissing` and changes nothing else. A second consecutive missing sample is read as `terminated` even though the module reported no such health state: the instance is gone, whatever the health output says, and `spec.providerID` is kept rather than cleared. ## `InfrastructureHealthy` and `Ready` `InfrastructureHealthy` is not itself mirrored to Cluster API, but once an object is provisioned it becomes, with `Deleting` (and for a pool also `ApplyJobSucceeded`), the entire set of conditions that `Ready` summarizes — down from the full set of dependency, credential, RBAC, apply, state and output conditions `Ready` watches beforehand. See [Reconcile flow]() for how `status.initialization.provisioned` latches, and [Conditions]() for exactly which conditions feed `Ready`, per kind and per phase. ## Unhealthy samples A `TerraformMachine` counts consecutive unhealthy samples in `status.unhealthySamples`: each successful refresh or drift Job after provisioning is one sample (an apply’s own reading counts too, when it already gives a definite reading and stands in for the post-apply refresh). How an `InfrastructureHealthy` reading changes the count: | Reading | Effect on the count | | --- | --- | | `Healthy` | Resets it to zero | | `InstanceUnhealthy`, `InstanceDegraded` or `InstanceStopped` | Adds one | | `InstancePending`, `HealthUnknown` or `InstanceTerminated` | Leaves it as it is | This count is what drives Cluster API machine remediation: see [Machine Remediation]() for the threshold, the terminated-instance shortcut, the `cluster.x-k8s.io/remediate-machine` annotation and its withdrawal. > [!NOTE] > > **See also** > > - [Drift]() > - [Machine Remediation]() > - [Plan preview and approval]() > - [Reconcile flow]() > - [Conditions]() # Drift CAPTF periodically refreshes and plans a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` to check whether the infrastructure still matches its module’s inputs. This page covers how to set the check interval and action for each kind, disable checks where that is possible, read the results, and remediate what a check finds. See [Drift and Health]() for how the checks work and how they feed the object’s health. > [!NOTE] > > **Before you begin** > > - A provisioned `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`. > - `kubectl` access to edit the object and read its status. ## Set the drift interval Every kind checks for drift on a timer: `spec.drift.intervalSeconds`. Left unset, it uses the manager’s `--drift-default-interval` flag (30 minutes by default; see [Manager Flags]()). A `TerraformMachine` and a `TerraformMachinePool` inherit an interval from the owning `TerraformCluster`’s `spec.defaults.drift.intervalSeconds` when they set none of their own: ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster spec: drift: intervalSeconds: 900 # (1)! ``` 1. The cluster’s own interval. ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster spec: defaults: drift: intervalSeconds: 1800 # (1)! ``` 1. Inherited by every `TerraformMachine` and `TerraformMachinePool` that sets no interval of its own. ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachine spec: drift: intervalSeconds: 900 ``` ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool spec: drift: intervalSeconds: 900 ``` A machine’s or pool’s own `intervalSeconds` always wins over the inherited default. `spec.drift` is operational policy and stays mutable for the object’s whole life, unlike the fields that define the machine or pool itself (see [The Kinds]() for what is mutable per kind). ## Set the drift action `spec.drift.action` decides what happens when a check finds changes: `Report` (the default) only records the finding, `Remediate` applies the object’s current inputs to remove it. Both a `TerraformCluster` and a `TerraformMachinePool` accept `action`: ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster spec: drift: action: Remediate ``` > [!NOTE] > > **A `TerraformMachine` only reports** > > A `TerraformMachine`’s drift is always reported, never remediated: the underlying instance is immutable infrastructure, replaced by a rollout rather than patched in place, so `spec.drift` has no `action` field. `spec.defaults.drift` on the `TerraformCluster` has no `action` field either: only a machine’s or pool’s own `spec.drift.action` sets it. ## Disable drift checks Setting `intervalSeconds: 0` disables drift checks on a `TerraformCluster` or a `TerraformMachine`. > [!WARNING] > > **Disabling drift checks also stops health sampling** > > Health is read only from a completed refresh or drift Job, so disabling drift stops every health sample taken after provisioning: the `InfrastructureHealthy` condition stops updating. A `TerraformMachine` still refreshes independently while its `spec.remediation.annotateMachine` is true; see [Machine Remediation](). A `TerraformMachinePool`’s drift cannot be disabled; see [Drift and health]() for why. ## Read the results `DriftDetected` reports the last check’s outcome: | Status | Reason | Meaning | | --- | --- | --- | | `Unknown` | `DriftNotChecked` | The first check has not completed. | | `False` | `NoDrift` | The last check found no changes. | | `True` | `DriftReported` | The last check found changes, and the action is `Report`. | | `True` | `DriftPending` | It found changes, the action is `Remediate`, and the remediation has not started or failed. | | `True` | `DriftRemediating` | A remediation apply is running. | ```sh kubectl get terraformcluster -n \ -o jsonpath='{.status.conditions[?(@.type=="DriftDetected")]}' ``` Its message names the drift Job and, when changes were found, the counts of resources to add, change and destroy. `status.lastRun.drift` carries the same counts as separate fields plus up to 20 affected resource addresses, and `status.lastDriftCheck` is when the last successful check completed. `DriftJobSucceeded` separately reports whether the check (or refresh) Job itself succeeded, independent of what it found. See [Conditions]() and [Last run]() for every reason and field. ## Remediate drift With `action: Remediate` set as above, a `TerraformCluster` or `TerraformMachinePool` re-applies its current inputs whenever a check finds changes, retried with the same backoff as any other apply (see [The reconcile lifecycle]()). > [!WARNING] > > **Remediation is capped** > > After `failedJobsHistoryLimit` ([job tuning]()) failed remediation applies since the last successful check, `DriftDetected` stays `DriftPending` until the next check succeeds. A `TerraformCluster` remediation apply also goes through the destructive-plan guard (or plan approval under `applyPolicy: Manual`); the guard does not apply to a `TerraformMachinePool` (see [Plan Approval]()). See [Drift and Health]() for exactly how a remediation runs and what each `DriftDetected` reason during it means. A `TerraformMachine`’s drift is never remediated this way: an unhealthy instance is instead signaled to Cluster API for replacement. See [Machine Remediation](). ## Confirm it worked > [!TIP] > > After a check completes, `status.lastDriftCheck` advances and `DriftDetected` reads `False`/`NoDrift` or `True` with a reason above. After a `Remediate` apply succeeds, `DriftDetected` returns to `False`/`NoDrift` and `status.lastRun.drift` is empty again. > [!NOTE] > > **See also** > > - [Drift and Health]() for how checks and health sampling work. > - [Machine Remediation]() for handling an unhealthy `TerraformMachine`. > - [Plan Approval]() for the destructive-plan guard a `TerraformCluster` remediation apply goes through. > - [Machine Pools]() for the pool’s separate membership refresh schedule. > - [Conditions]() and [Manager Flags](). # Machine Remediation A `TerraformMachine` reports its instance’s health on the `InfrastructureHealthy` condition. By itself that only ever turns the owning Machine’s `InfrastructureReady` condition `False`; nothing acts on it unless something is watching. `spec.remediation` optionally asks Cluster API to replace an unhealthy instance, through a `MachineHealthCheck` and the Machine’s owner. This page covers configuring it, how it reaches a replacement, and what it does to the Machine along the way. See [Drift and Health]() for how health itself is computed. > [!NOTE] > > **Before you begin** > > - A provisioned `TerraformMachine` whose owning Machine is selected by a `MachineHealthCheck` with an `unhealthyMachineConditions` check on `InfrastructureReady`. > - `kubectl` access to edit the `TerraformMachine` and read its owning Machine. ## Enable remediation `spec.remediation.annotateMachine` defaults to `false`, meaning CAPTF never touches the Machine. Set it to `true` to have CAPTF request remediation once an instance has been unhealthy for long enough: ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachine spec: remediation: annotateMachine: true ``` `spec.remediation` is operational policy: it stays mutable for the machine’s whole life, unlike `spec.source` and `spec.identityRef`, which are fixed at creation. ## Tune the threshold and sampling interval ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachine spec: remediation: annotateMachine: true unhealthyThreshold: 5 healthCheckIntervalSeconds: 120 ``` - `unhealthyThreshold` (1-100, default 3) is how many consecutive unhealthy health samples are needed before CAPTF signals remediation. A sample is one completed refresh or drift Job; a single transient reading is never enough on its own, since once a `MachineHealthCheck` acts on the signal, the Machine is replaced. A terminated instance skips the threshold; see [Terminated instances](<#terminated-instances>) below. - `healthCheckIntervalSeconds` (60-86400, default 300) is how often a provisioned instance is refreshed to resample its health while `annotateMachine` is `true`, independent of `spec.drift.intervalSeconds`. With `annotateMachine` set to `false` it has no effect, and health is then re-read only at the drift (or refresh) cadence set by `spec.drift.intervalSeconds`; see [Drift]() for that setting, including how `intervalSeconds: 0` stops health sampling entirely once `annotateMachine` is `false`. ## How it works with a MachineHealthCheck 1. Health outside `running`/`healthy`, read after provisioning, turns `InfrastructureHealthy` `False`, which turns `Ready` `False`, which Cluster API mirrors into the owning Machine’s `InfrastructureReady` condition. 2. A `MachineHealthCheck` with `spec.checks.unhealthyMachineConditions: [{type: InfrastructureReady, status: "False", timeoutSeconds: }]` selecting the Machine starts its own timeout once `InfrastructureReady` turns `False`. 3. When the timeout (or, with `annotateMachine`, the annotation below) marks the Machine unhealthy, `MachineHealthCheck` sets the Machine’s `OwnerRemediated` condition. Only the Machine’s **owner** acts on that condition: a `MachineDeployment`’s `MachineSet`, or a control-plane provider that implements remediation, such as `KubeadmControlPlane` or `RKE2ControlPlane`. A single-replica control plane refuses to remediate itself. 4. The owner that acts replaces the Machine (and, through its deletion, the `TerraformMachine`) with a new one rather than repairing it in place: `spec.source` is immutable, so there is nothing for CAPTF to patch. > [!NOTE] > > **`annotateMachine` feeds the flow, it does not replace it** > > It gives the flow a faster, more specific signal than the timeout alone (see the next section). ## The remediate-machine annotation With `annotateMachine: true`, once `unhealthyThreshold` consecutive samples are unhealthy, degraded or stopped (or the instance is terminated), CAPTF sets `cluster.x-k8s.io/remediate-machine` on the owning Machine. This asks the `MachineHealthCheck` reconciler to treat the Machine as unhealthy at once, bypassing `unhealthyMachineConditions`’ timeout; remediation still goes through the owner as described above. CAPTF marks the annotation as its own with a second annotation, `captf.io/remediation-requested`, whose value is why it was set, for example `InstanceUnhealthy for 5 consecutive samples` or `the instance is terminated`. CAPTF removes both annotations once the instance reads `Healthy` again and the Machine is not being deleted, withdrawing a request that has not yet been acted on. It never removes `cluster.x-k8s.io/remediate-machine` when `captf.io/remediation-requested` is absent: an annotation set by someone else is left alone. See [Annotations, Labels and Finalizers]() for both annotations’ full definitions. ## Terminated instances > [!WARNING] > > **A terminated instance skips the threshold** > > CAPTF sets the annotations on the very first sample that reads the instance terminated, since there is nothing to wait for and no risk of a transient reading. The exception is a `provider_id` output that goes missing: the first missing sample only sets `InfrastructureHealthy=Unknown`/`ProviderIDMissing`, and the second consecutive one is read as terminated and requests remediation. ## Confirm it worked > [!TIP] > > - `kubectl get machine -n -o jsonpath='{.metadata.annotations}'` shows `cluster.x-k8s.io/remediate-machine` and `captf.io/remediation-requested` once a request is made, and neither once the instance recovers or is replaced. > - A `RemediationRequested` event on the `TerraformMachine` marks a new request, and `RemediationWithdrawn` marks a withdrawal; see [Events](). > - `captf_remediation_requests_total{action="requested"|"withdrawn"}` counts both; see [Metrics](). > [!NOTE] > > **See also** > > - [Drift and Health]() for how health samples are produced. > - [Drift]() for `spec.drift.intervalSeconds`, which paces health sampling when `annotateMachine` is `false`. > - [TerraformMachine]() for every `spec.remediation` field. > - [Conditions]() for `InfrastructureHealthy`’s reasons. # Approvals and Gates CAPTF can stop a `TerraformCluster` before it changes infrastructure and wait for a person to say yes. It has two such gates, and they apply to the cluster kind only. This chapter explains what each gate does, what it binds, what is deliberately not gated, and what an operator needs to run them safely. The task-oriented commands are in [Plan Approval](). The pages: 1. This page: the gates, a decision table and the scaling questions. 2. [Manual plan approval](): `applyPolicy: Manual`, the plan Job, `status.plan` and the approved apply. 3. [What the plan hash binds](): the `p2:` fingerprint. 4. [The destructive-plan guard](): the check on every cluster apply under `Automatic`. 5. [Operating the gates](): commands, conditions, events, retries, upgrades and who can approve. 6. [Other manual actions](): restores, abandoning an object, and the fixes that are not approvals. 7. [Limits](): what an approval does and does not promise. ## The two gates | Gate | Set by | What it stops | Annotation that releases it | Binds | | --- | --- | --- | --- | --- | | **Manual plan approval** | `spec.applyPolicy: Manual` on a `TerraformCluster` or its template | Every apply except the first one | `captf.io/approve-plan=` | The plan: what each change does, and the output changes | | **Destructive-plan guard** | Always on, for every `TerraformCluster` apply | An apply whose plan deletes or replaces a resource | `captf.io/approve-destructive-plan=` | The inputs, not the plan | Under `Manual` the first gate subsumes the second: approving a plan also approves the deletes and replacements that plan lists, so the second annotation is never needed there. Both gates are specific to `TerraformCluster`, with one narrow extension. `applyPolicy` exists on no other kind, and the Job builder adds the guard to a `TerraformCluster`’s apply, and to a `TerraformMachinePool`’s apply **only when it renders cluster exports that differ from its last successful apply** (see [Machine pools]()). A `TerraformMachine`’s apply is never gated: its instance is immutable, and Cluster API replaces it rather than its apply changing it. Destroy, restore, refresh and drift Jobs are never gated on any kind. ## What is gated An apply decision has one reason; the reason decides whether a gate sees it. | Situation | Kind | `Automatic` | `Manual` | | --- | --- | --- | --- | | First apply of a new object (no state) | `TerraformCluster` | Guard runs, nothing to delete | Not gated, applies at once | | Changed inputs (image, spec, variables) | `TerraformCluster` | Guard: blocked if the plan deletes or replaces | Plan, then wait for `approve-plan` | | Retry after a failed apply | `TerraformCluster` | Guard | Plan, then wait (an earlier approval is reused if the re-plan matches) | | Drift remediation (`drift.action: Remediate`) | `TerraformCluster` | Guard | Plan, then wait | | State with no inputs hash | `TerraformCluster` | Guard | Plan, then wait | | Apply that renders a changed set of cluster exports | `TerraformMachinePool` | Guard: held if the plan deletes or replaces; the pool keeps applying the last exports | Not applicable: no `applyPolicy` | | Any other apply | `TerraformMachinePool` | Not gated | Not applicable: no `applyPolicy` | | Any apply | `TerraformMachine` | Not gated | Not applicable: no `applyPolicy` | | Destroy, restore, refresh, drift check | Any | Not gated | Not gated | ## Scaling is not gated Scaling does not go through an approval, so a large change to machine counts does not wait for one: - **Scaling a `MachineDeployment` up** creates `TerraformMachine` objects. Each one’s first apply creates its machine; there is nothing for a gate to check. - **Scaling it down** deletes `Machine` objects. After Cluster API drains them, each `TerraformMachine` runs a destroy Job, which is never gated. - **A fixed-replica `MachinePool`** re-applies its pool module when `spec.replicas` changes. Such an apply is not gated, unless it also renders changed cluster exports. - **An autoscaled `MachinePool`** scales in the cloud, with no Terraform run at all (see [Machine Pools]()). A 100-machine `MachineDeployment` therefore scales without an approval. If a change to shared infrastructure must wait for a person, that infrastructure belongs in the cluster module. See [Limits]() for how cluster outputs reach machines and pools. ## In this section - **Approve a Plan** --- The commands to review and approve a destructive plan or a manual plan. - **Manual Plan Approval** --- applyPolicy: Manual, the plan Job, status.plan and the approved apply. - **What the Plan Hash Binds** --- The p2: fingerprint an approval names, and what it covers. - **The Destructive-Plan Guard** --- The check on every cluster apply under Automatic. - **What Approval Does Not Guarantee** --- What an approval does and does not promise. ## Where the state lives An approval rests on three things that other chapters describe: - the plan key, a Secret that makes the plan hash a keyed value (see [Run inputs and the plan key]()); - `status.plan`, written by the controller from what the runner reports; - annotations on the object, which are the approval itself. > [!NOTE] > > **See also** > > - [Plan Approval]() for the commands. > - [Conditions]() and [Events]() for the exact reasons. > - [Annotations, Labels and Finalizers](). > - [Security Model](). # Plan Approval `TerraformCluster` guards two things before it changes infrastructure: a plan that deletes or replaces a resource waits for an approval, and with `spec.applyPolicy: Manual` every plan waits for one. This page is the how-to: the commands to review and approve. How the gates work, what the plan hash binds and what they do not cover are in the chapter [Approvals and Gates](). > [!NOTE] > > **Before you begin** > > - A `TerraformCluster` whose apply you want to guard or preview. > - `kubectl` access to annotate it and to read the logs of its Jobs. Anyone who can patch the object can approve; see [who can approve](). ## The destructive-plan guard Under the default `applyPolicy: Automatic`, every `TerraformCluster` apply, including a drift remediation, plans first and stops before applying if the plan deletes or replaces a resource. Nothing changes. `ApplyJobSucceeded` turns `False`/`DestructivePlanBlocked` and its message names the affected resources and the inputs hash. See [The destructive-plan guard]() for the full behavior. To approve: 1. Read the plan: `kubectl logs job/ -n -c source`. The condition message lists what it deletes or replaces. 2. Approve the **inputs hash** named in the condition: ```sh kubectl annotate terraformcluster -n \ captf.io/approve-destructive-plan= --overwrite ``` `` and `` are the `TerraformCluster`’s; `` is the hash from the condition message. > [!NOTE] > > **The approval covers exactly those inputs** > > The next change produces a new hash and is guarded again. The controller removes the annotation after the approved apply succeeds. `lifecycle { prevent_destroy = true }` in the module remains the stronger control for a resource that must never be replaced. ## Plan preview: applyPolicy Manual Set `spec.applyPolicy: Manual` on the `TerraformCluster` or its `TerraformClusterTemplate` to review every change except the first apply. Switching back to `Automatic` applies whatever was waiting. 1. A change plans first. `ApplyJobSucceeded` becomes `Unknown`/`PlanAwaitingApproval` and `status.plan` fills in. Nothing applies. 2. Review `status.plan` (counts, and up to 50 resources with their actions) and the plan Job’s log, which has the human-readable plan: ```sh kubectl get terraformcluster -n -o jsonpath='{.status.plan}' kubectl logs job/ -n -c source ``` 3. Approve by naming the plan hash, `status.plan.planHash`: ```sh kubectl annotate terraformcluster -n \ captf.io/approve-plan= --overwrite ``` 4. The apply plans again and applies only if the new plan has the same hash. If anything changed, it stops with `PlanChanged` and a new plan to approve. See [Manual plan approval](). 5. After the approved apply succeeds, the controller removes the annotation, clears `status.plan` and emits `PlanApplied`. > [!WARNING] > > **Approving a plan also approves its deletes and replacements** > > `captf.io/approve-destructive-plan` is not needed under `Manual`. A plan with no changes needs no approval. A plan that only changes outputs, or only imports or moves, does. ### Caveats - The plan hash binds what each change does, including old and new values, so a plan with different values needs its own approval. It reveals no value; read values in the plan Job’s log. See [What the plan hash binds](). - An approval is consumed only when the approved apply succeeds. A stale one can approve a later plan with the same hash. Remove it with `kubectl annotate terraformcluster -n captf.io/approve-plan-`. - Hashes start with `p2:`. After an upgrade from a release with `p1:` hashes, a waiting plan is planned again and needs a new approval. - `status.plan` is status, not durable state: after `clusterctl move` the controller plans again, and an approval still on the annotation applies if the new plan hashes the same. - In a GitOps setup, the annotation is set by a person, or by a pipeline after its own review of `status.plan` and the plan Job’s log. Do not keep it in Git: a controller that syncs annotations from Git would re-add a consumed approval. ## What is not guarded Neither gate applies to a `TerraformMachine`, and a `TerraformMachinePool` is guarded only when its apply renders a changed set of cluster exports (see [Machine pools]()). Destroy, restore, refresh and drift Jobs are never gated. See [Approvals and Gates]() and [Limits](). ## Confirm it worked > [!TIP] > > - After approving a destructive plan, `kubectl describe terraformcluster -n ` shows `ApplyJobSucceeded` back to `True` and the `captf.io/approve-destructive-plan` annotation gone. > - After approving a plan under `Manual`, `status.plan` is empty and `ApplyJobSucceeded` is `True`/`ApplySucceeded`. > [!NOTE] > > **See also** > > - [Approvals and Gates](), the chapter behind this page. > - [Operating the gates]() for conditions, events and RBAC. > - [Drift]() for `drift.action: Remediate`, which the destructive-plan guard also covers. > - [Reconcile Lifecycle](). > - [Annotations, Labels and Finalizers](), [Conditions]() and [Events](). # Manual Plan Approval With `spec.applyPolicy: Manual` on a `TerraformCluster` (or on its `TerraformClusterTemplate`), the controller plans every change first and applies it only after a person approves that plan. `Automatic` is the default. `applyPolicy` is mutable: switching back to `Automatic` applies whatever was waiting and clears `status.plan`. The first apply of a new cluster is not gated, since there is no state yet to damage. This page describes the flow; the commands are in [Plan Approval](). ## The flow ``` sequenceDiagram participant U as Operator participant C as Controller participant J as Job and runner C->>J: plan Job with the plan key J->>J: validate, plan, fingerprint J-->>C: counts, resources, plan hash C->>C: record status.plan, set PlanAwaitingApproval U->>C: annotate approve-plan with the plan hash C->>J: apply Job with expect-plan J->>J: plan again with refresh, fingerprint alt hash matches J->>J: apply the saved plan J-->>C: success C->>C: clear status.plan, remove approve-plan else hash differs J-->>C: stop before applying, new plan C->>C: update status.plan, set PlanChanged U->>C: approve the new hash end ``` 1. **Plan.** When an apply is due, the controller starts a plan Job (`status.activeJob.operation: plan`) instead. Before the Job, it ensures the plan key Secret exists, and the Job mounts it at `/captf/plan-key`. The runner runs `init`, `validate` and `plan -out`, then `show -json` with the output held in memory only. It never writes the plan JSON to disk or to a log; only the binary plan file exists in the working directory. 2. **Record.** The runner reports counts, a list of changed resources and the plan hash. The controller records them in `status.plan` (below), sets `ApplyJobSucceeded` to `Unknown`/`PlanAwaitingApproval` with a message that contains the exact `kubectl annotate` command, and emits one `PlanReady` event. 3. **Wait.** Nothing applies. The condition is `Unknown`, not `False`, so waiting never makes `Ready` false. The controller re-checks at least every ten minutes, and a changed annotation or new inputs trigger it at once. Drift and health checks continue while a plan waits. 4. **Approve.** You set `captf.io/approve-plan` to `status.plan.planHash`. 5. **Apply.** The controller starts the apply Job with `--expect-plan=`, the same plan key mount, and the Job annotation `captf.io/approved-plan`, and emits `PlanApproved`. The runner plans again, this time including the refresh, and computes the hash of that fresh plan. - If it equals the approved hash, the runner applies exactly the saved plan file. - If it differs, the runner stops before changing anything and reports the new plan. See [When the plan changes](<#when-the-plan-changes>). 6. **Done.** After the approved apply succeeds, the controller removes `captf.io/approve-plan`, clears `status.plan` and emits `PlanApplied`. The removal is its own patch with an optimistic lock, so a newer value that someone wrote in the meantime survives; on a conflict the reconcile requeues and tries again. ## What `status.plan` holds | Field | Meaning | | --- | --- | | `inputsHash` | The inputs hash the plan was made for. A plan is bound to it: new inputs make a new plan | | `job` | The plan Job | | `planHash` | The `p2:` hash to approve (see [What the plan hash binds]()) | | `add`, `change`, `destroy` | Counts of planned resource changes; a replacement counts as both an add and a destroy | | `outputChanges` | How many outputs change | | `resources` | Up to 50 entries of `
()`, sorted by address, never a value | | `truncated` | Set when more than 50 resources changed | | `createdAt` | When the controller recorded the plan | The labels in a `resources` entry are the action (`create`, `update`, `delete`, `replace`, `read` or `forget`), followed by `import` and then `move` where they apply: `aws_lb.x (import)`, `aws_instance.b (update, move)`. Imports and moves are not counted in `add`, `change` or `destroy`, so an entry is the way to see them. > [!NOTE] > > **Plan values never reach status, events or logs** > > They carry counts, addresses and the keyed hash only. The values are in the plan Job’s own log, which the `source` container prints in human-readable form. ## When approval is needed - **Everything except the first apply**, when `Manual` is set: changed inputs, a retry after a failed apply, a drift remediation and a state with no inputs hash. - **Output changes.** A plan that changes only outputs, or only imports or moves, is not an empty plan and waits for approval. - **Not an empty plan.** A plan with no resource, output, import or move change has a fixed hash, and the apply proceeds without approval (no `PlanApproved` event). It still plans again first, and stops if the plan is no longer empty. > [!WARNING] > > **Approving a plan also approves the deletes and replacements it lists** > > You saw them in `status.plan.resources`, and the apply runs only that plan. `captf.io/approve-destructive-plan` is not needed under `Manual`. `lifecycle { prevent_destroy = true }` in the module still fails an approved plan. ## When the plan changes If the plan the apply Job computes does not hash to the approved value, because the world moved since you reviewed it or because a partial earlier apply changed things, the Job stops before the apply step: - `status.lastRun.error.kind` is `plan-changed` and the Job is annotated `captf.io/plan-changed`; - `status.plan` is replaced by the new plan; - `ApplyJobSucceeded` becomes `Unknown`/`PlanChanged`, with the new command, and a `PlanChanged` warning event is emitted; - the change counts toward neither retry backoff nor the remediation failure cap, and the apply waits for approval of the new hash. A plan that comes back empty after you approved a non-empty one is a changed plan too. ## Retries and stale approvals An approval is consumed only when the approved apply succeeds. After a failed apply the annotation stays, so a retry whose new plan hashes the same runs without another approval, with the usual retry backoff. If the failed apply changed something, the plan differs and needs a new approval. > [!WARNING] > > **A leftover approval still approves a later plan with the same hash** > > This covers, for example, a plan that changed again or inputs that changed before the apply ran. Remove a stale one with `kubectl annotate terraformcluster -n captf.io/approve-plan-`. Because the hash binds values, a leftover approval matches only a plan that changes the same attributes to the same values. > [!NOTE] > > **See also** > > - [What the plan hash binds](). > - [Operating the gates](). > - [Plan Approval](). > - [Run inputs and the plan key](). # What the Plan Hash Binds The value you approve under `applyPolicy: Manual` is `status.plan.planHash`, a string that starts with `p2:`. It is a fingerprint of what the plan will do. This page defines what goes into it and why, so you can tell what an approval covers. ## Construction The runner builds the hash from one line per change: ```text || ``` - `` is the resource address, or `output.` for an output. - `` are the plan’s actions, comma-joined, in plan order (so the two orders of a replacement differ). - The last field is a keyed HMAC of the change’s *effect*, described below. The lines are sorted and joined with newlines, and the hash is `p2:` and the hex SHA-256 of that text. A resource appears when its action is not a no-op, or when it carries an import, or a moved-from address, since an import and a move are changes even where nothing else is. An output appears when its change is not a no-op. The HMAC key is the per-object plan key (see [Run inputs and the plan key]()). Because the hash is keyed, reading `status.plan.planHash` reveals nothing about any value, and the same plan hashes differently for two objects. A plan or approved-apply Job whose key is missing or shorter than 32 bytes fails before any step runs, so it cannot produce an unkeyed hash. ## The effect, by kind of change What each line’s HMAC covers depends on the action. The effect is serialized as canonical JSON: keys sorted, empty fields omitted, and numbers kept as the literals Terraform wrote. | Change | What the effect contains | | --- | --- | | **Update or replace** | Only the attributes the change touches, as leaf paths with their old and new values; the paths whose value is unknown until apply; changes in sensitivity; whether a value is absent or null; and whether a path step is a list index or a map key | | **Create or read** | The planned object, with its unknown and sensitive markers | | **Delete or forget** | Nothing about the resource: the address and the action are all that identify it | | **Import** | What the import imports (the identifier or identity) | | **Move** | The address the resource moves from | | **Output** | The same rules as a resource, under `output.` | A delete, forget or no-op line still carries any import or move on it. ## What this means for approval **An approval binds values, not only addresses.** A plan that changes the same resources to different values hashes differently and needs its own approval. A new image tag, a different CIDR or a changed count each changes the hash. > [!WARNING] > > **Untouched attributes are not part of the hash** > > For an update, only leaf paths that differ between the old and new object count. Attributes the change leaves alone, such as an autoscaler’s desired capacity under `ignore_changes` or a Kubernetes `resourceVersion`, can move between your review and the apply without invalidating the approval. Without this rule, any busy object would force a re-approval at every apply. **Unknown values are bound as unknown.** A value the provider will compute at apply time appears in the effect as “unknown after apply”, not as a value. The approval covers that the attribute will be computed, not what it will become. **Sensitive values are bound but hidden.** A change from sensitive to not sensitive changes the hash, and the value itself stays inside the HMAC. **Imports and moves always need approval.** They are changes even with no other difference, so a plan whose only change is an `import` or a `moved` block waits for approval. **Output changes need approval.** An output that changes appears in the hash, and `status.plan.outputChanges` counts them. Cluster outputs feed machines and pools; see [Limits](). ## What the hash does not bind It does not bind how the provider carries the change out. A reviewed update to a security group is bound to its old and new attributes, not to the API calls the provider makes, their order, or what the cloud does with them. It does not bind anything the module does outside the plan, such as provisioners that reach elsewhere. See [Limits](). ## The empty plan and old hashes A plan with no change at all has a fixed hash (`p2:` followed by the SHA-256 of the empty string), and an apply for it needs no approval. A hash that does not start with `p2:` is a plan recorded by an older release (`p1:`). The controller re-plans it automatically and the approval is then given again, against the new hash. An approval annotation written for a `p1:` hash does not match. ## Drift is classified separately Drift detection does not use this hash. It counts resources by action only, so an import or a move alone is not drift, and neither is an output change. That is why a drift check can report no drift while a `Manual` plan still has something to approve. > [!NOTE] > > **See also** > > - [Manual plan approval](). > - [Drift]() and [Drift and Health](). > - [Limits](). # The Destructive-Plan Guard A `TerraformCluster` is mutable: a new image tag, an edited field or a drift remediation re-applies its module. A module release that renames a resource, or a different module pointed at the same state, can plan to destroy infrastructure nobody meant to destroy. The guard stops that. It is on for every `TerraformCluster` apply under `applyPolicy: Automatic`, and it cannot be turned off. A `TerraformMachinePool` has a narrower form of the same guard, for applies that render a change of the cluster’s exports (see [Machine pools](<#machine-pools>)). The command to approve is in [Plan Approval](). ## The flow ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A["Apply Job starts
(guard on)"] --> B["init, validate,
plan, show"] B -->|no changes| S["Skip the apply step,
Job succeeds"] B -->|changes| D{"Does any change
delete or replace?"} D -->|no| AP["Apply the saved plan"] D -->|yes| H{"Is the inputs hash
approved?"} H -->|yes| AP H -->|no| BL["Stop: nothing applied
DestructivePlanBlocked"] BL -->|"approve-destructive-plan
names the inputs hash"| A ``` The guarded apply runs `init`, `validate`, `plan -detailed-exitcode -out`, `show -json` and then applies the saved plan, instead of applying directly. - A change counts as destructive when its planned actions include a delete: a removal, or a replacement in either order. A `forget`, an `import` or a `move` is not destructive. - If the plan has no changes, the apply step is skipped and the Job succeeds, adopting the inputs hash and pinning the image digest as a full apply would. - If the plan only creates or updates in place, it applies without any approval. - If it deletes or replaces anything and the current inputs hash is not approved, the Job stops before the apply step. Nothing changes and the state is untouched. The guard applies to every reason for an apply on a `TerraformCluster`: the first apply, changed inputs, a retry, a drift remediation and a state with no inputs hash. Under `Manual` it is replaced by the plan approval: a plan you approved is run as it is. ## What a block looks like - `status.lastRun.error.kind` is `blocked`. - `ApplyJobSucceeded` is `False`/`DestructivePlanBlocked`. The message names the affected addresses and actions (`
(delete)` or `(replace)`), says nothing was applied, and gives the inputs hash and the command to approve it. - A `DestructivePlanBlocked` warning event is emitted once per blocked Job, in place of the generic failure event. - The Job is annotated `captf.io/destructive-plan-blocked`. - A blocked Job counts toward neither the retry backoff nor the remediation failure cap, and does not mark the last apply as failed. While the newest apply of the current inputs hash is blocked, no apply of that hash starts, and the controller re-checks at least every ten minutes. What else pauses depends on why the apply was due: - A blocked **input change** pauses drift and health checks too, since they would render the unapplied inputs. - A blocked **drift remediation** leaves drift and health checks running. A new inputs hash starts a new guarded apply at once, and so does approving the blocked one. Editing the annotation re-triggers the reconcile at once. ## Approving Read the plan first: `kubectl logs job/ -c source` has its human-readable output, and the condition message lists what it deletes or replaces. Then name the **inputs hash** from the message: ```sh kubectl annotate terraformcluster -n \ captf.io/approve-destructive-plan= --overwrite ``` The value is an inputs hash, not a plan hash. It is the hash of everything the apply renders (the image reference and every module input; see [Job Inputs]()), so it approves exactly those inputs, never the object. Any later change produces a new hash that the annotation does not name, and the next destructive plan is blocked again without anyone removing the approval. Two things follow from approving inputs rather than a plan: - The approval covers the plan computed when the approved apply runs, not the plan you read. Something that changed in the meantime is covered too. If that matters, use `applyPolicy: Manual`, which binds the plan. - The controller passes the approval to the runner only when the annotation equals the hash it rendered, and the runner compares it again. Any other value approves nothing and is not an error. ### Consumption After any successful apply of the approved inputs hash, the controller removes the annotation, in its own patch with an optimistic lock, emits a `DestructivePlanApprovalConsumed` event and counts it in `captf_destructive_plan_approvals_consumed_total`. > [!WARNING] > > **An approval is consumed even when the plan was not destructive** > > A drift remediation re-applies the inputs the state already records, so an approval left behind would also cover a later destructive remediation of the same inputs. After consumption that remediation is blocked again and needs its own approval. ## Machine pools A `TerraformMachinePool` has no `applyPolicy`, but its apply is guarded in one case: when it renders cluster exports (`captf_cluster_outputs`) that **differ from those of the pool’s last successful apply**. The guard is the same: the apply stops before a plan that deletes or replaces anything. Nothing else about a pool is guarded on its own: the first apply, bootstrap rotations, version rolls, replica changes and spec edits all apply as before while the exports are unchanged. Machines are never guarded. See [Limits](). The controller keeps a record of the exports of each successful apply in the pool’s durable inputs Secret, and compares the current exports with it. A pool with no record is covered in [Pools that applied before the record existed](<#pools-that-applied-before-the-record-existed>). ### The approval hash A pool’s approval names its **approval hash**: the inputs hash without `bootstrap_data`. It survives bootstrap rotations, which change `bootstrap_data` roughly every few minutes, and changes on any other input change. `ApplyJobSucceeded` shows the current hash and the command: ```sh kubectl annotate terraformmachinepool -n \ captf.io/approve-destructive-plan= --overwrite ``` Approval is RBAC only: whoever may `patch` the pool may set it. It is removed after the successful apply it approved. ### When it is blocked: the change is held A blocked apply of a change of the exports is **held**, not retried: - The pool **keeps applying everything else** (rotations, upgrades, edits) with the exports of its last successful apply, so it keeps working. - **Refresh and drift render the held exports**, so the waiting change does not read as drift. - `ApplyJobSucceeded` stays `False`/`DestructivePlanBlocked`, with what the plan would delete or replace and the approval hash, so the pool’s `Ready` is `False` while a change is held. - One `Warning` event, `DestructivePlanBlocked`, is emitted per blocked Job. - The change is recorded in the durable Secret (`captf.io/pending-cluster-outputs`). ### Withdrawn and superseded - **Withdrawn.** If the exports return to the applied ones, nothing waits. The condition reports the last apply’s real outcome and names the withdrawn Job, and an approval of that change is removed. If the change comes back, it is held again and needs a fresh approval. - **Superseded.** If the exports move to a different change, the old change’s approval is removed. An approval is for one change. ### Partly applied If a guarded apply fails after it started, or its Job vanishes mid-run, the state may hold part of the change. The controller records it (`captf.io/partial-cluster-outputs`). > [!WARNING] > > **Until an apply succeeds, a partly applied pool holds nothing** > > **Until an apply succeeds, the pool holds nothing**: it cannot fall back to the last successful apply’s exports, and every apply is guarded by its own approval hash, rotations included. A destructive one waits for approval. This matters most for modules whose rotations replace resources, such as an instance configuration. ### Unrecorded The record of the applied exports shares the durable Secret’s budget (about 1,000,000 bytes) with the rendered inputs. If it did not fit while a change is pending, the pool cannot hold the change and waits, like a cluster, until it is approved or an input other than `bootstrap_data` changes. A hash annotation (`captf.io/applied-cluster-outputs-hash`) survives the record being dropped, so a later change is still guarded. ### Pools that applied before the record existed A pool that applied before CAPTF recorded applied exports has no baseline. On its first reconcile after the upgrade, CAPTF records the exports it can **prove** the pool applied: - its newest successful apply’s exports, when the durable inputs still hash to the state’s inputs hash; or - its current exports, when no apply Job is retained and the current inputs hash to the state’s. Without that proof (an apply Job was recorded as gone, or the Secret does not match the state), **every apply of the pool is guarded until one succeeds**. A plan that deletes or replaces nothing applies and records the baseline. A destructive plan waits for approval like a cluster’s, and `ApplyJobSucceeded` says the exports of the last successful apply are unknown and gives the approve command. See [Upgrades](). ### An apply Job deleted while it ran This applies to `TerraformCluster` and `TerraformMachinePool`, not to `TerraformMachine`. If an apply Job is deleted while it runs (for example `kubectl delete job`), the Job may have applied part of its change. CAPTF confirms with live reads that the Job is gone and that the object’s live status still names it, then records it in the durable inputs Secret (`captf.io/interrupted-apply=`). Then: - **An apply of the current inputs stays due**, even if the inputs equal the state’s. - **That apply is guarded** wherever the destructive guard applies: every cluster apply, and pool applies that change exports. A plan that would undo the lost Job’s work waits for approval. Under `applyPolicy: Manual` the due apply plans first and waits for plan approval as usual. - `ApplyJobSucceeded` is `False`/`ApplyFailed` with the message `Job : disappeared while it ran and may have applied part of its change; an apply of the current inputs is due`, followed by whether that apply is guarded. This condition takes precedence over an older blocked or plan-changed apply: it names the vanished Job until an apply started after it (one carrying `captf.io/after-interrupted-apply`) reports its own outcome, so an earlier “approve hash …” message no longer hides it. - It clears when an apply started afterwards succeeds (those Jobs carry `captf.io/after-interrupted-apply`). Nothing is recorded for a stuck Job that CAPTF deleted itself, and a `TerraformMachine` records nothing. For a pool whose lost Job rendered a change of the cluster’s exports, the change is also recorded as partly applied (see above). ### Limits - A deleting pool applies nothing: its condition says no apply runs. - `clusterctl move` does not carry Jobs or status, so the move re-runs the blocked plan once. ## Blocked after a failed apply For clusters and pools, an apply that was blocked after an earlier apply **failed** now waits for its approval, and the approval applies it. Before, such a blocked apply sat idle and the earlier failure was forgotten. ## Prevent destroy > [!TIP] > > **Use prevent\_destroy for a resource that must never be replaced** > > `lifecycle { prevent_destroy = true }` in the module is the stronger control: it fails even an approved plan, under either policy. > [!NOTE] > > **See also** > > - [Manual plan approval](), which binds a plan instead of the inputs. > - [Operating the gates](). > - [Drift](), whose `Remediate` action the guard also covers. > - [Conditions](). # Limits The gates are a review step, not a sandbox. This page says what an approval promises and what it does not, so you set your expectations and your module design to match. ## An approval binds a plan, not the apply The approval binds what the plan says will change, as far as the [plan hash binds it](). It cannot bind how the provider then carries that change out at apply time: the API calls it makes, their order, retries, or what the cloud does in response. A reviewed change can still fail halfway, and a partial apply leaves the world different from the plan you read, which is why the next attempt re-plans and may ask for a new approval. ## The plan key is readable by the module The plan key is mounted into the Job that plans, and the runner’s ServiceAccount can read every Secret in its namespace (see [Security considerations]()). So a module image can read the key and compute any hash it likes. The fingerprint protects against the world changing between your review and the apply, and against a person approving something other than the plan they read. > [!CAUTION] > > **A hostile module image can forge any hash** > > The fingerprint does not protect against it. Review the module you run, and keep modules you do not trust out of the namespace. ## What is guarded The cluster has both gates. `TerraformMachine` applies are never gated. A `TerraformMachinePool` apply is guarded in one case only: when it renders cluster exports that differ from its last successful apply (see [Machine pools]()). Destroy, restore, refresh and drift Jobs are never gated on any kind. The reasons are structural: a machine is immutable and Cluster API replaces it rather than changing it, and scaling would otherwise wait on a person for every machine (see [scaling]()). The consequence is that a machine or pool module can do anything its credentials allow with no approval step, apart from that one pool case. Keep shared and destructive infrastructure, such as networks, load balancers and databases, in the cluster module, where the gates apply, and keep machine and pool modules to what is meant to come and go. ## Cluster outputs reach pools and machines A cluster’s exports feed the inputs of every machine and pool. An output change in a cluster plan needs approval under `Manual`. After that apply succeeds, the new exports change the inputs of each `TerraformMachinePool`, which is mutable, and the pool applies them. That apply is **guarded**: if its plan deletes or replaces anything, it stops and the change is **held**. The pool keeps applying with the exports of its last successful apply until someone approves the change with the pool’s approval hash. See [Machine pools](). What the guard does not cover: - **A pool that applied before the guard existed** has no record of its last exports until its next successful apply, and is unguarded until then. - **Plans that only update in place** are not blocked, as for a cluster. - **A `TerraformMachine` is immutable** and is not re-applied; its inputs are fixed once it is provisioned. Nothing guards a machine. - **The pool apply waits for the cluster’s Job** (the cluster operation gate), but that is a scheduling lock, not an approval. > [!TIP] > > **Make a pool module defensive with prevent\_destroy** > > Still make a pool module defensive with `lifecycle { prevent_destroy = true }` on what must never go: it fails even an approved plan. ## Approving is a patch on the object Approval authority is Kubernetes RBAC, by design (see [Who can approve]()). The consequence is that whoever may `patch` the cluster can both change its spec and approve the change. Keep that right with the people who should hold both, and route everyone else’s spec changes through a reviewed path. ## Approvals do not cover state surgery Restoring a state, abandoning an object and unlocking a state lock are not approvals and are not gated: they act on whoever can annotate or patch the object. > [!CAUTION] > > **A restore can make the controller forget resources it created** > > A restore rewrites the state, so a person with that right can do this. Treat these annotations as equal in weight to the right to edit the object. A rebuild of a lost state, including imports, is a runbook, not an approval: see [Total State Loss and Import](). > [!NOTE] > > **See also** > > - [Security Model](). > - [Security considerations]() for the Secrets behind these limits. > - [Manual plan approval]() and [The destructive-plan guard](). # Deletion and Teardown Deleting a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` means running Terraform’s `destroy` in a Job, then removing what CAPTF stored about the object. Most of the time nothing needs doing: you delete the Cluster Cluster API owns and everything follows. This chapter explains the order things happen in, what holds a deletion, how a deletion ends, and what is left behind when you short-circuit it. The step-by-step recoveries stay in the [runbooks](); this chapter is the model behind them. The pages: 1. This page: the rules in one table, and the pages that follow. 2. [Order and finalizers](): who deletes what first, and when the finalizer comes off. 3. [The destroy Job](): what runs, what does not gate it, and how it fails. 4. [Held deletions](): missing or unreadable state, restore, and the abandon annotation. 5. [Terminating namespaces](): deletions that need no Job, and the Secrets that go with the namespace. 6. [Cleanup and garbage collection](): what the controller deletes, what Kubernetes deletes, and the RBAC sweep. 7. [`clusterctl move`](): the delete that skips the delete path. 8. [Stripping a finalizer by hand](): the consequences. 9. [My object will not delete](): a flowchart from the symptom to the action. ## In this section - **Order and Finalizers** --- Who deletes what first, and when the finalizer comes off. - **The Destroy Job** --- What runs, what does not gate it, and how it fails. - **Terminating Namespaces** --- Deletions that need no Job, and the Secrets that go with the namespace. - **Cleanup and Garbage Collection** --- What the controller deletes, what Kubernetes deletes, and the RBAC sweep. - **My Object Will Not Delete** --- A flowchart from the symptom to the action. ## The rules A deleting object follows these rules, in this order. The first that applies decides the pass. | Situation | What happens | Page | | --- | --- | --- | | The object or its Cluster is paused | Nothing: a paused object runs only Job bookkeeping, so it neither destroys nor releases | [Order]() | | A Job of the object is running | The deletion waits for it; the destroy starts after | [Order]() | | A `TerraformCluster` with machines or pools left | `DeletionBlocked=True`/`DependentsExist`, no destroy yet | [Order]() | | No state, and the object never applied | The finalizer comes off at once | [Held]() | | No state, and the object applied before | Held: `StateReadable=False`/`StateLost` | [Held]() | | State exists but cannot be read | Held: `StateCorrupt`, `StateEncrypted` or `StateInconsistent` | [Held]() | | Readable state | A destroy Job runs; no gate or approval applies | [Destroy]() | | The destroy succeeded | Cleanup runs and the finalizer comes off | [Cleanup]() | | The destroy failed or cannot start | Retried with backoff, forever; the abandon annotation releases it | [Held]() | ## What the controller never does - It never drops the finalizer while a state Secret may still describe live resources, except on an explicit abandon. - It never destroys against a state it cannot read. - It never invents inputs to destroy with: the destroy renders from the durable inputs Secret (or, for a cluster or pool without one, the current inputs when they build). - It never skips the destroy because it failed. There is no skip-destroy annotation, only the [abandon]() one. > [!NOTE] > > **See also** > > - [The Reconcile Lifecycle]() for the pass that runs all of this. > - [Terraform State]() and [Lifecycle walkthroughs]() for the Secrets. > - [Other manual actions]() for the abandon annotation next to the other manual fixes. > - [Stuck Destroy](), [Unreadable State](), [State Restore]() and [clusterctl move]() for the commands. # Order and Finalizers Three parties delete things in a fixed order: Cluster API, the admission webhook that guards the order, and the controller’s own delete pass. This page walks through each. ## The order ``` sequenceDiagram participant U as You participant C as Cluster API participant M as TerraformMachine participant K as TerraformCluster U->>C: delete the Cluster C->>C: delete Machines and MachinePools C->>C: drain, run lifecycle hooks C->>M: delete (after the Machine is gone) M->>M: destroy Job, cleanup, finalizer off C->>K: delete (after the control plane is gone) K->>K: wait for machines and pools, then destroy ``` - **Machines first.** Cluster API deletes the `Machine`, drains the node and runs the lifecycle hooks it owns (pre-drain, pre-terminate), and only then deletes the `TerraformMachine` the Machine’s `spec.infrastructureRef` names. CAPTF registers no hook of its own; the order comes from Cluster API and the webhook below. - **Machine pools** follow their `MachinePool` through the owner reference. Nothing refuses a direct delete of a `TerraformMachinePool`. - **The cluster last.** The `TerraformCluster` is deleted once the Cluster’s machines are gone, and its destroy waits for any that are left (see [below](<#the-cluster-waits-for-its-machines>)). ## The webhook refuses a direct machine delete Deleting a `TerraformMachine` directly would skip the drain and the hooks, so the validating webhook refuses it while a `Machine` that is not itself deleting references the object. The check lists the Machines in the namespace and compares each `spec.infrastructureRef` (API group, kind and name) with the object. It does not use the machine’s owner references or labels, because anyone who may update the object can change those. - The error names the Machine: `delete the Machine instead`. - A `TerraformMachine` that nothing references, such as one whose Machine is already gone, can always be deleted. - A `Machine` that is already deleting does not count, so the delete Cluster API issues after the drain goes through. - **The one exception is `clusterctl move`.** The `clusterctl.cluster.x-k8s.io/delete-for-move` annotation is honored only while the Cluster named by the machine’s `cluster.x-k8s.io/cluster-name` label has `spec.paused=true`. The annotation alone can be set by anyone who can update the object. See [clusterctl move](). > [!NOTE] > > **This is the only delete guard on the three kinds** > > A `TerraformCluster`, `TerraformMachinePool` and the template kinds allow every delete. A `TerraformClusterIdentity` is refused while any object uses it or while its credentials are still mirrored into a namespace. ## The cluster waits for its machines A `TerraformCluster` that is deleting lists, through the uncached reader, the `TerraformMachine` and `TerraformMachinePool` objects in its namespace that carry its cluster name in the `cluster.x-k8s.io/cluster-name` label. While any exist it sets `DeletionBlocked=True`/`DependentsExist` with a count and requeues every 30 seconds, and starts no destroy. When none are left it sets `DeletionBlocked=False`/`NotBlocked` and carries on. - The lists are unfiltered by `--watch-filter`: machines of another manager instance count too. - The cluster name is the object’s own label, else its owning Cluster’s name. With neither, nothing can reference the cluster and nothing blocks. - Machines and pools have no such wait: their `DeletionBlocked` never applies. The ownerless path runs the same check. A deleting object with no owner reference of the expected kind is released at once only when all of these hold: it has no state Secret, no Job and no live run lease, nothing blocks its deletion, and it never applied. If any fails, the full delete pass takes over: it waits for the dependents, holds on a lost state, or destroys. An owner reference whose target is gone, or a forged or stale one, does not get an object stuck either: while deleting, an owner gate falls through like a missing owner, so the destroy still runs from the durable inputs. ## The finalizer Each kind carries one finalizer (`terraformcluster.infrastructure.cluster.x-k8s.io` and its machine and pool counterparts; see [Annotations, Labels and Finalizers]()). The controller removes it only in three cases: 1. **After a successful destroy.** 2. **On a deletion with nothing to destroy:** no state and the object never applied (see [ever applied]()). 3. **On abandon:** the annotation named the object’s UID (see [abandon]()). And only once no Job of the object runs. Two checks guard that, because the Job list comes from the manager’s cache, which can lag a Job the controller just created: - `status.activeJob` names a Job the cached list lacks but the API server still has. The pass waits five seconds and looks again. - A **live run lease** is held by a Job the cache has not shown. The cleanup defers by five seconds too. The lease is read live only when it matters, which includes every deleting object. See [Run leases and the cluster operation gate](). The first reconcile with a deletion timestamp sets the `Deleting` condition to `True`/`Deleting` and emits a `DeletionStarted` event. `Deleting` has no other `True` reason. ## Pause stops a deletion > [!WARNING] > > **A deletion under a paused Cluster waits until it is unpaused** > > A paused object (its own `cluster.x-k8s.io/paused` annotation or a paused Cluster) runs only the paused branch: Job bookkeeping and the block-move clear. That branch never destroys and never removes the finalizer. The object keeps its finalizer, starts no Job, and `Deleting` reads `True`/`Deleting` with the message `Deletion waits until the object is unpaused (Cluster spec.paused or the cluster.x-k8s.io/paused annotation); a paused object never runs a destroy`. The reason is `clusterctl move`. It pauses the Cluster, then deletes the source objects with their finalizers stripped. Those deletes must not run a destroy, or the move would tear down infrastructure the target is taking over. Pause is the signal that makes the controller leave them alone. The exception is the ownerless release above, which runs before the pause check. > [!NOTE] > > **See also** > > - [The destroy Job](). > - [The Reconcile Lifecycle](). > - [My object will not delete](). # The Destroy Job Once a deleting object has readable state, no running Job and (for a cluster) no dependents, the controller decides a `destroy` Job. Every other decision a live object makes is skipped: a deleting object never applies, refreshes or checks drift. ## What is not in the way - **No approval.** `applyPolicy: Manual` and the destructive-plan guard belong to applies. A destroy is never gated, on any kind. See [Approvals and Gates](). - **No policy check.** The merged Job policy check that refuses a `lockTimeoutSeconds` not below `activeDeadlineSeconds` (`JobPolicyInvalid`) is skipped for a destroy, so a teardown never wedges on it. See [Tuning Jobs](). - **No input gates.** A deleting object builds no inputs and ignores `DependenciesReady`. The destroy renders from the **durable inputs Secret** `captf-inputs--`, the files of the last successful apply. A `TerraformCluster` or `TerraformMachinePool` whose Secret is gone falls back to building its current inputs when they build; a `TerraformMachine` never does, since its Machine and bootstrap Secret are usually gone by then. See [The durable inputs Secret is missing](). - **The pinned image.** For a machine, the destroy runs the image and the identity recorded in the durable Secret, not the current spec, so changing a template cannot change how an existing machine is torn down. ## What is in the way A destroy still takes the same leases as an apply: - The object’s **run lease**: one Job at a time per object. - The **cluster operation gate**, when it is on (the default): a `TerraformCluster` destroy takes the Cluster’s write lease and waits for any machine or pool apply or destroy in flight (`WaitingForMachineOperations`); a machine’s destroy waits while a `TerraformCluster` apply is running (`WaitingForClusterOperation`). The wait shows as `ApplyJobSucceeded=Unknown` with the wait reason, counts toward no backoff, and re-checks every 30 seconds. See [Leases and the operation gate](). ## Credentials, only when needed A live object prepares its runner credentials (the identity check, the credential mirror and the runner ServiceAccount and RoleBinding) on every pass. A deleting object does it only when a Job is about to start (`deletionCredentials`). A deletion that needs no Job, such as an object that never applied or one released by abandon, therefore never waits on credentials. Why this matters in a terminating namespace is on [that page](). If the credentials are not ready the destroy waits, and the pass reports the first condition that is not `True`: - The identity no longer allows the namespace, or is gone: `ApplyJobSucceeded=False`/`IdentityNotAllowed`, message `Destroy waits until the identity allows this namespace again`. - Otherwise the `Deleting` condition’s message reads `The destroy Job waits for its credentials: is ()`, naming `IdentityAllowed`, `CredentialsMirrored` or `RunnerRBACReady`. Both retry every 30 seconds. Both are among the cases the [abandon annotation]() releases. ## Results and retries | Outcome | `ApplyJobSucceeded` | What follows | | --- | --- | --- | | Destroy succeeded | `True`/`DestroySucceeded` | [Cleanup](), finalizer off | | Destroy failed | `False`/`DestroyFailed` | Retry with backoff, forever | | Killed at the deadline | `False`/`JobDeadlineExceeded` | Retry with backoff, forever | | Image could not be pulled | `False`/`ImagePullFailed` | Retry with backoff | | Image breaks the contract | `False`/`ImageInvalid` | Retry with backoff | | Stopped from outside (drain, eviction, Job delete) | `JobInterrupted` event | Retry at once, no backoff | The backoff is the one every operation uses: one minute after the first failure, doubling to a ten-minute cap, and the cap from the point the object’s retained failed Jobs reach `failedJobsHistoryLimit` (default 3, counted as at least 1). A destroy that the **deadline killed counts as a failure** even though the runner reports it interrupted: a step that always hangs must reach the cap. See [Failing Jobs](). > [!WARNING] > > **There is no retry limit** > > A destroy that can never succeed stays at `ApplyJobSucceeded=False`/`DestroyFailed` and `Ready=False`, and [`CAPTFDestroyStuck`]() fires. The way out is to fix the cause, to [abandon](), or to clean up and [strip the finalizer](). A failed destroy may have destroyed part of the infrastructure. The next attempt runs against the state the failed one left, so it continues rather than starting over. > [!NOTE] > > **See also** > > - [Held deletions](). > - [Stuck Destroy](). > - [Reference: ApplyJobSucceeded](). # Terminating Namespaces Deleting a namespace deletes everything in it, including the `Terraform*` objects and every Secret CAPTF keeps for them. The cloud resources are not touched. This page explains what a deletion can still do while its namespace is terminating, and what the namespace takes with it. ## What can still finish Kubernetes refuses to create new objects in a terminating namespace. A destroy needs new objects: a Job, its per-run Secret, and often the runner ServiceAccount, its RoleBinding and the credential mirror. So the controller sorts a deleting object by whether it needs a Job at all. | Situation | Needs a Job | Result | | --- | --- | --- | | The object never applied and has no state | No | The finalizer comes off at once | | Held on lost or unreadable state | No | Held; [abandon]() releases it | | Abandoned | No | The finalizer comes off at once | | A destroy of readable state | Yes | Waits, as below | | A restore of a held state | Yes | Waits, as below | The credentials a deleting object needs are prepared only for a Job that is about to start, so the first three never wait on them. The reconcile does not treat a credential failure as an error for a deleting object: the namespace lifecycle would reject the creation again on every retry. The object records the failure in its credential conditions instead. ## A destroy that waits When the credentials cannot be prepared, the pass does not start the Job. It sets the `Deleting` condition’s message to `The destroy Job waits for its credentials: is ()` and requeues every 30 seconds. The named condition is the first of `IdentityAllowed`, `CredentialsMirrored` and `RunnerRBACReady` that is not `True`. A restore says `restore` where the message says `destroy`. If the credentials already exist, the Job create itself is a create in a terminating namespace, which the API server rejects. The reconcile reports the error and controller-runtime retries it with backoff. A rejected create gives the run and cluster leases back, so nothing else waits on a Job that does not exist. Either wait is a case the abandon annotation releases, with the cause `it waits for its credentials`. Where the namespace deletion has already removed the state Secrets, the object is held, not released: `status.initialization.provisioned` still marks it as applied, so the missing state reads as `StateLost`, not as nothing to destroy. ## What goes with the namespace Every Secret CAPTF keeps is namespaced, so a deleted namespace removes: - the state Secrets and their backups, - the durable inputs Secret (and with it the `captf.io/applied` marker), - the plan key and the credential mirror, - the Jobs, their pods and their per-run Secrets, - the run and cluster Leases and the state lock Lease, - the runner ServiceAccount and RoleBinding. The credential **source** Secret named by a `TerraformClusterIdentity` lives in whichever namespace the operator put it, and the identity itself is cluster-scoped; neither goes unless that namespace is the one deleted. ### The known limit The applied marker is the only thing that tells a moved object that applied from one that never did, because status is not moved (see [`clusterctl move`]()). > [!CAUTION] > > **Deleting a moved object’s namespace can orphan its cloud resources** > > When the namespace of a **moved** object is deleted, the marker goes with the Secret, the state and the backups. The object, now without any of the four [ever applied]() signals, looks like one that never applied, and its finalizer comes off with nothing destroyed. An object that was not moved keeps `provisioned` in its status and is held instead. Move first, and delete the source namespace only when the target has taken over. > [!NOTE] > > **See also** > > - [Lifecycle walkthroughs: namespace deletion](). > - [Stuck Destroy: terminating namespaces](). > - [Cleanup and garbage collection](). # Cleanup and Garbage Collection Three things remove what CAPTF stored for an object: the controller’s own cleanup, Kubernetes garbage collection through owner references, and the namespace RBAC sweep. Which one removes what is what decides what survives a finalizer that comes off early. ## The controller’s cleanup Cleanup runs in exactly three situations: after a successful destroy, on a deletion with nothing to destroy, and on [abandon](). It never runs while a state Secret may still describe live resources, and it waits (five-second retries) while a Job of the object still holds a live run lease. Its steps, in order: 1. **The state Secrets**, every chunk, deleted one by one (the manager role has no `deletecollection`), then the **state lock Lease** `lock-tfstate-default-`. Neither Terraform nor OpenTofu deletes its own state: a destroy only empties it. 2. **The durable inputs Secret** `captf-inputs--`. 3. **The plan key Secret.** It is owned by the object too, but is removed here so it does not outlive a finished destroy while it waits for garbage collection. 4. **The run lease** `captf-run-` and, for a `TerraformCluster`, the **cluster write lease** `captf-cluster-`: every run or cluster lease carrying the object’s owner labels, whoever holds it. 5. **The mirror’s ownership.** The object is removed from the owners of the credential mirror `captf-creds-`. The mirror is deleted when this was its last owner, with a `MirrorRemoved` event. 6. **The finalizer**, in the same patch that writes the status. A destroy then reports a `Destroyed` event (“Infrastructure destroyed; state and inputs removed”) and `FinalizerRemoved`; an abandon reports the `InfrastructureAbandoned` warning; a no-state deletion reports only `FinalizerRemoved`. The object’s per-object metric series go with the finalizer. ## What Kubernetes collects These are owned by the object and carry no deletion of their own. They go when the object does, after the finalizer: | Object | Owner | Notes | | --- | --- | --- | | State backups `captf-state-backup--` | The object, not as a controller | Deliberately not deleted by cleanup; the finalizer keeps them for a held deletion | | Jobs and their pods | The object, as a controller | Also pruned earlier by the history limits | | Per-run Secrets | Their Job | Normally deleted much earlier, when the Job finishes and is bookkept | | Anything left of the state chunks, durable inputs or plan key | The object | Only when cleanup did not run: a stripped finalizer | > [!NOTE] > > **Owner references never delay the owner’s deletion** > > An owner reference without `blockOwnerDeletion` never delays the owner’s deletion. ## What nothing deletes - **The cloud resources**, after an abandon or a stripped finalizer. - **The credential source Secret** and the `TerraformClusterIdentity`. A `variablesFrom` ConfigMap or Secret. A Cluster Cluster API still owns. - **The runner ServiceAccount and RoleBinding**, until the namespace is empty of `Terraform*` objects (below). - **The credential mirror**, while another object of the namespace still uses it. ## The namespace RBAC sweep When a reconcile removed a finalizer, it then sweeps the object’s namespace. If the namespace holds no `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`, the sweep deletes the `captf.io/managed=true` ServiceAccounts, RoleBindings and Leases in it. If objects remain, it only prunes the RoleBinding’s subjects down to the ServiceAccounts some object still runs as. - The sweep runs after the finalizer’s removal is persisted, but the deleted object still counts until it is gone. When it is the last one, this sweep therefore may find the namespace non-empty, and the periodic sweep removes the runner objects instead: at manager start, then every `--sync-period` (default 10 minutes) on the leader, with 10% jitter. - It reads through the uncached API reader without `--watch-filter`, so an object another manager instance owns keeps the namespace. - A failed sweep is logged, never an error: the next tick retries. - Templates run no Jobs, so they do not keep a namespace. The same sweep is what clears the runner objects a `clusterctl move` leaves on the source. See [RBAC](). > [!NOTE] > > **See also** > > - [Secret Management]() for what each Secret holds. > - [Terraform State: state on deletion](). > - [Stripping a finalizer by hand](). # Cloud Modules # Cloud Modules The CAPTF project maintains reference modules for five clouds: AWS, Google Cloud, Azure, Oracle Cloud Infrastructure (OCI) and OpenStack. Each set lives in its own repository, implements the [v1alpha1 module contract](), and ships as module images you reference from a `TerraformCluster`, `TerraformMachineTemplate` or `TerraformMachinePool`. They are meant to be used as they are, and to be forked as the starting point for your own modules. A sixth set, [No-op](), provisions nothing: it runs real plans and keeps real state with no cloud and no credentials, for trying CAPTF and testing a management cluster. > [!WARNING] > > **Status: pre-release** > > The five cloud sets pass static analysis, mocked unit tests on Terraform and OpenTofu, and image smoke tests, but none has yet been applied to a real cloud. Each repository’s `DESIGN.md` lists the facts the first real apply must confirm. Read it before you rely on a module, and pin a release tag. ## The modules - **AWS** --- Network Load Balancer, security groups, IAM roles and an S3 bootstrap bucket; EC2 instances and Auto Scaling groups. - **Google Cloud** --- Proxy Network Load Balancer, firewall rules and service accounts; Shielded VMs and managed instance groups. - **Azure** --- Standard load balancer, network security groups and a resource group; VMs and scale sets. - **OCI** --- Network security groups, a network load balancer and a dynamic group; compute instances and instance pools. - **OpenStack** --- Octavia load balancer, security groups and a server group; Nova servers. No machinepool role. - **No-op** --- Real plans, state and outputs with nothing provisioned: try CAPTF or test a management cluster without a cloud. - **Shared Behavior** --- What all five sets have in common: endpoint, traffic, identities, bootstrap data, health, tags and destroy. | Cloud | Repository | Roles | Provider | | --- | --- | --- | --- | | [AWS]() | [`captf-io/aws-modules`]() | cluster, machine, machinepool | `hashicorp/aws` 6.67.0 | | [Google Cloud]() | [`captf-io/gcp-modules`]() | cluster, machine, machinepool | `hashicorp/google` 8.5.0 | | [Azure]() | [`captf-io/azure-modules`]() | cluster, machine, machinepool | `hashicorp/azurerm` 5.7.0 | | [OCI]() | [`captf-io/oci-modules`]() | cluster, machine, machinepool | `oracle/oci` 9.8.0 | | [OpenStack]() | [`captf-io/openstack-modules`]() | cluster, machine | `terraform-provider-openstack/openstack` 3.4.0 | | [No-op]() | [`captf-io/noop-modules`]() | cluster, machine, machinepool | None: `terraform_data` is built in | OpenStack has no machinepool role: it has no native scaling group, and a `MachineDeployment` of individual machines covers the same need. Each role is a separate image: - **cluster** creates what one workload cluster needs around its nodes: the security rules, the API server load balancer and, except on OpenStack, the node identities. It publishes the API endpoint, the failure domains and the ids the other roles need (its `exports`). - **machine** creates one instance for one `Machine` and registers control-plane instances with the API load balancer. - **machinepool** creates one native scaling group (an Auto Scaling group, a managed instance group, a scale set or an instance pool) for one `MachinePool`. The no-op roles create none of this: each records its inputs in its state and returns placeholder outputs. None of them creates a network. You bring the network, the subnets and the egress path, and the modules make nodes in them. ## Images and tags Every role is published as `ghcr.io/captf-io/-`, for example `ghcr.io/captf-io/aws-machine`, for `linux/amd64` and `linux/arm64`, in two flavours: one on the Terraform base image and one on the OpenTofu base image. Each image carries a mirror of the providers its role needs, so a Job never downloads a provider at run time. | Tag | Meaning | | --- | --- | | `vX.Y.Z-opentofu`, `vX.Y.Z-terraform` | Release `vX.Y.Z` on that runtime; never moves | | `opentofu`, `terraform` | The newest release on that runtime | | `edge-opentofu`, `edge-terraform` | The newest build of `main` | Pin a release tag, or a digest, in anything you keep. The examples in each cloud repository pin `v0.1.0-opentofu`. ## What you bring The no-op set needs only the identity. Every cloud set needs all of these: - **A network.** The VPC, VNet, VCN or network, its subnets, and an egress path (NAT gateway, Cloud NAT, router) exist before the cluster. Each cloud page lists what the subnets need. - **Node images.** Images built for Cluster API, such as the [image-builder]() images, with cloud-init and a kubelet that runs with `--cloud-provider=external`. - **A cloud controller manager and a CNI** in the workload cluster: the nodes stay `NotReady` and tainted until both run. The modules set up what the cloud controller manager needs (node identities, tags, provider IDs) but do not install it; on Azure they also write its configuration file on each node. On OpenStack you also supply its credentials. - **An identity.** A `TerraformClusterIdentity` whose Secret holds the cloud credentials the provider reads; each cloud page shows the keys. See [Identities and Credentials](). ## Using the modules 1. Read the cloud page for the prerequisites and the identity Secret. 2. Create the identity Secret and the `TerraformClusterIdentity`. 3. Generate a cluster from the repository’s `examples/cluster-kubeadm.yaml` with `clusterctl generate yaml --from`, filling in the network ids and the node image. 4. Install the CNI and the cloud controller manager in the workload cluster once its API server answers. The no-op set skips all of this: see its [quick start](). Module settings beyond the contract inputs are ordinary module variables, set with `spec.variables` or `spec.variablesFrom`; see [Module Variables](). Each role page lists them. ## Forking a module Each repository is built to be forked. Its `CONVENTIONS.md` is the rule book the five repositories share; `make verify` checks most of it: the file layout, tags on every resource, `terraform validate` and `tofu validate` (including the oldest supported runtimes), the unit tests, tflint, [`tfcapi-lint`](), shellcheck and a trivy scan. To make a fork your own, change the registry and image names in the `Makefile` and the workflow, run `make lock`, then `make verify`. The no-op modules are the smallest modules that meet the contract, and the starting point of [Your First Module](). [Shared Behavior]() describes what all five sets have in common. # Shared Behavior The five sets of [cloud modules]() follow one set of conventions, so they behave the same way wherever the cloud allows it. This page describes that common behavior; each cloud’s pages describe only what differs. ## The API endpoint - **Internal by default.** The cluster role puts the API server load balancer on your subnets, reachable from inside the network. Setting `api_load_balancer_public = true` makes it internet-facing and requires a non-empty `api_allowed_cidrs`. Nodes then reach the endpoint through your egress path, so the list must include its public addresses; each cloud page lists what else a public endpoint needs. - **Your own endpoint.** When the `Cluster` already has a control-plane endpoint (kube-vip, an external load balancer), the cluster role creates no load balancer and publishes that endpoint unchanged. - **The port.** The load balancer listens on `Cluster.spec.clusterNetwork.apiServerPort` (default 6443). With `distribution = "rke2"` it also listens on the RKE2 supervisor port, 9345, and the kube-apiserver backends stay on 6443, where RKE2 always listens. - **Hairpin.** A control-plane node must reach the endpoint itself while it joins (see the [cluster contract]()). Each cloud’s load balancer handles this differently; the cloud pages say how. - **The endpoint guard.** Cluster API never updates a cluster’s endpoint after its first value, so a moved endpoint breaks every kubeconfig and certificate. Each cluster role records, when the load balancer is created, every input that decides the endpoint (visibility, subnets, port, address) and fails any later plan that would change one, with a message naming the recorded and the requested value. CAPTF’s [destructive-plan guard]() catches replacements only; the endpoint guard also catches changes a cloud makes in place. AWS lets you add load balancer subnets, and removing one is caught only by the destructive-plan guard; see [AWS](). ## Network traffic Nodes of one cluster accept all traffic from each other, scoped to the cluster’s own security group, network security group, network tag, service account or application security group, so any CNI works without port lists. Nothing else is open by default except the API port from the load balancer, and on OpenStack the NodePorts from the node subnet (see [OpenStack Cluster]()). Some clouds have variables that open more, such as `nodeport_allowed_cidrs` on OCI. SSH is closed unless you list sources in `ssh_allowed_cidrs` (Google Cloud uses OS Login instead). ## Node identities The cluster role creates the cloud identities the cloud controller manager and the CSI driver need on the nodes: instance profiles, service accounts, managed identities or a dynamic group. Variables take existing identities instead. On OCI the dynamic group covers the control-plane nodes only, and needs a defined tag to match them by. OpenStack has no equivalent: its cloud controller manager needs credentials you supply in the workload cluster. ## Bootstrap data - Both `cloud-config` and `ignition` bootstrap formats are accepted. Gzipped Ignition is refused. A precondition checks the cloud’s user-data size limit. - Control-plane bootstrap data holds the cluster’s CA and service-account keys. Where instance user data is readable outside the instance, the machine role can stage it in a secret store the node’s identity reads: S3 on AWS, Secret Manager on Google Cloud. The instance gets a small `#cloud-boothook` script that fetches the payload and installs it as `/etc/cloud/cloud.cfg.d/99-captf-bootstrap.cfg`, which cloud-init reads before it runs its modules. `bootstrap_delivery = "inline"` turns staging off for images without the fetch tools. - Elsewhere the payload goes in user data, and the cloud page says who can read it. Azure keeps custom data out of its API and instance metadata. On OCI and OpenStack the instance metadata service serves it, which each page lists as an exception. ## Machine pools - **Bootstrap rotation in place.** Bootstrap data rotates about every 7.5 minutes. A rotation updates the group’s launch configuration (or the staged object) in place, and never replaces a running instance; new instances pick up the current data. - **Version rolls.** A change of the `MachinePool`’s Kubernetes version, compared verbatim so that an RKE2 `+rke2rN` bump counts, rolls the instances. Any other change that replaces instances is listed in the role’s Exceptions. - **Zones.** A pool that names no failure domains takes the cluster’s at its first apply and keeps them, so a later change to the cluster’s zones never moves running instances. - **Autoscaling.** With the autoscaler annotations set, the group’s minimum and maximum come from them and an apply never resets the desired count. - **Node labels.** Cluster API does not put a `MachinePool`’s labels on its nodes, so the module does: a boot script adds `--node-labels` to the kubelet’s arguments (both `/etc/default/kubelet` and `/etc/sysconfig/kubelet`) and writes an RKE2 `config.yaml.d` file. Labels the NodeRestriction admission plugin forbids are dropped and listed in the `dropped_node_labels` output. Ignition with labels is refused. ## Health Every role reports the contract’s `health` output from the cloud’s own state, with a machine-readable reason: - A resource that is gone, or in a state that means deletion, is `terminated` with a reason like `InstanceNotFound` or `LoadBalancerNotFound`. - Cluster health comes from the API load balancer, never from its backends, which are unhealthy during every normal control-plane bring-up. - Pool health is, in this order: the group gone gives `terminated`; a desired capacity of 0 gives `running` and healthy; no members yet gives `pending` (`NoMembers`); any degraded, stopped or unknown member gives the worst of those; otherwise `running`, healthy only when every member runs and the count matches, with `ScalingInProgress` while it does not. A starting member never makes the pool `pending`. On Google Cloud an autoscaled group still at size 0 while its minimum is above 0 also reports `pending`, so a group created empty is not marked provisioned early. See [Drift and Health]() for how CAPTF uses these readings. ## Tags and labels Every resource that can carry tags or labels carries the contract’s `captf_tags`, plus your `additional_tags`. The `captf.io/` keys are not valid on every cloud, so each maps them the same way everywhere: | Cloud | Mapping | Example | | --- | --- | --- | | AWS | unchanged | `captf.io/cluster` | | Google Cloud labels | lowercase; `.` to `-`; other invalid characters to `_` | `captf-io_cluster` | | Azure | `/` to `_` | `captf.io_cluster` | | OCI free-form tags | `.` and space to `_`; at most 4 `additional_tags` | `captf_io/cluster` | | OpenStack | Nova metadata `/` to `:`; Neutron and Octavia tags as `key=value` | `captf.io:cluster` | ## Exports and externally managed clusters The cluster role’s `exports` output carries a `schema` string such as `captf.io/aws-cluster/v1`, the region, the failure domains, the API endpoint and the ids the other roles need. The machine and machinepool roles check the schema. When the `TerraformCluster` is externally managed its exports are empty; set the user variable `external_cluster_exports` on the machine or pool, in the same shape, to run them anyway. ## Destroy A cluster’s resources go only after all its machines and pools are gone. Each role reads what you brought (subnets, networks, identities) through listings that come back empty rather than failing, so deleting the network first does not leave a `destroy` stuck, except where a cloud page’s Limitations say otherwise. Resources are counted so that one deleted outside Terraform reads as gone, not unknown. ## Common variable names The same concept has the same variable name in every repository: | Concept | Variable | | --- | --- | | Public API load balancer | `api_load_balancer_public` | | API client allowlist | `api_allowed_cidrs` | | Static private API address | `api_load_balancer_private_ip` | | Kubernetes distribution | `distribution` (`kubeadm` or `rke2`) | | SSH allowlist | `ssh_allowed_cidrs` | | Spot capacity | `spot`, or the cloud’s term (`preemptible`) | | Image name placeholders | `{version}` (`v1.31.4`) and `{semver}` (`1.31.4`); on Google Cloud, whose image names hold no dots, also `{slug}` and `{fullslug}` (which keeps an RKE2 `+rke2rN` suffix) | | Staged bootstrap delivery | `bootstrap_delivery` | | Extra tags or labels | `additional_tags` | | Externally managed cluster | `external_cluster_exports` | # AWS The AWS modules build a workload cluster’s infrastructure inside a VPC you bring. The cluster role creates the API server’s Network Load Balancer, the security groups, IAM roles and instance profiles for the nodes, and an S3 bucket the nodes fetch their bootstrap data from. The machine role creates one EC2 instance per `Machine`, and the machinepool role one Auto Scaling group per `MachinePool`. The code lives in [`captf-io/aws-modules`](). Like every set, these modules are pre-release; see the status note in [Cloud Modules](). - **Cluster** --- The API endpoint, security rules, node identities and exports of one workload cluster. - **Machine** --- One instance per `Machine`, registered with the API load balancer. - **MachinePool** --- One native scaling group per `MachinePool`. ## Images | Role | Image | Page | | --- | --- | --- | | cluster | `ghcr.io/captf-io/aws-cluster` | [Cluster]() | | machine | `ghcr.io/captf-io/aws-machine` | [Machine]() | | machinepool | `ghcr.io/captf-io/aws-machinepool` | [MachinePool]() | The images pin the `hashicorp/aws` provider at 6.67.0. ## Prerequisites - **A VPC and one subnet per availability zone** for the nodes. Each zone becomes a failure domain. The subnets need a route to everything the nodes pull from: the internet through a NAT gateway, or VPC endpoints for S3, ECR, EC2 and STS plus a registry mirror. An internet-facing API endpoint needs public subnets of its own in the same zones. For `Service` load balancers, the cloud controller manager finds subnets by the `kubernetes.io/role/elb` and `kubernetes.io/role/internal-elb` tags, which you set. - **Permissions for the identity.** The credentials create and delete everything the three roles manage. The repository’s [`examples/identity-policy.json`]() covers all three roles; it scopes IAM changes to roles under `/captf/` and S3 to `captf-bootstrap-*` buckets. Its [README]() explains what it lets the holder do and how to narrow it. - **Node images.** cloud-init, the AWS CLI v2 on the `PATH` (for the staged bootstrap data), and the Kubernetes binaries for the version. [image-builder]()’s AWS images have all of it. Without an image ID the modules look up the newest [Cluster API Provider AWS image]() for the version, which that project builds for testing, not production. - **In the workload cluster**: [cloud-provider-aws](), the cloud controller manager that initializes the nodes and sets their provider IDs, and a CNI. For persistent volumes, the [Amazon EBS CSI driver](); the node roles carry no EBS permissions by default, so attach `AmazonEBSCSIDriverPolicy` with the cluster’s `node_role_policy_arns`. ## Identity Secret The provider block sets only the region; the AWS SDK’s default chain reads every credential from the identity’s Secret, as environment variables or as files under `/var/run/captf/credentials/`. | Key | Purpose | | --- | --- | | `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY` | Static keys | | `AWS_REGION` | The region, when the cluster’s `region` variable is unset | | `AWS_CONFIG_FILE`, `AWS_SHARED_CREDENTIALS_FILE`, `AWS_PROFILE` | A role to assume through shared config files | | `config`, `credentials` | The shared config files themselves | identity.yaml ```yaml apiVersion: v1 kind: Secret metadata: name: aws namespace: captf-system type: Opaque stringData: AWS_ACCESS_KEY_ID: REPLACE_WITH_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY: REPLACE_WITH_SECRET_ACCESS_KEY AWS_REGION: us-east-1 ``` A role to assume, which you should prefer over long-lived keys: identity.yaml ```yaml apiVersion: v1 kind: Secret metadata: name: aws-assume-role namespace: captf-system type: Opaque stringData: AWS_CONFIG_FILE: /var/run/captf/credentials/config AWS_SHARED_CREDENTIALS_FILE: /var/run/captf/credentials/credentials AWS_PROFILE: captf AWS_REGION: us-east-1 config: | [profile captf] role_arn = arn:aws:iam::123456789012:role/captf-provisioner source_profile = base role_session_name = captf credentials: | [base] aws_access_key_id = REPLACE_WITH_ACCESS_KEY_ID aws_secret_access_key = REPLACE_WITH_SECRET_ACCESS_KEY ``` The Job’s `HOME` is `/captf/work`, so `~/.aws` never exists and the paths are explicit. IRSA and EKS Pod Identity do not apply: the Job mounts no projected service account token. See [Identities and Credentials]() for the `TerraformClusterIdentity` that names the Secret. ## Quick start 1. Create the identity. Edit the placeholder values in [`examples/identity.yaml`](), then apply it: ```sh export TERRAFORM_IDENTITY_NAME=aws NAMESPACE=team-a AWS_REGION=us-east-1 clusterctl generate yaml --from examples/identity.yaml | kubectl apply -f - ``` 2. Create the cluster from [`examples/cluster-kubeadm.yaml`](): a `TerraformCluster`, a `KubeadmControlPlane`, a `MachineDeployment` and an autoscaled `MachinePool`. The only variables a cluster needs are the VPC and the zone-to-subnet map: ```sh export CLUSTER_NAME=demo KUBERNETES_VERSION=v1.34.1 export CONTROL_PLANE_MACHINE_COUNT=3 WORKER_MACHINE_COUNT=2 export AWS_VPC_ID=vpc-0123456789abcdef0 export AWS_SUBNETS='{"us-east-1a": "subnet-0aaa0000000000001", "us-east-1b": "subnet-0bbb0000000000002", "us-east-1c": "subnet-0ccc0000000000003"}' clusterctl generate yaml --from examples/cluster-kubeadm.yaml | kubectl apply -n team-a -f - ``` 3. Once the workload cluster’s API server answers, install cloud-provider-aws and a CNI in it. The example’s kubeadm configuration names each node after its private DNS name and sets its provider ID, as cloud-provider-aws expects. ## API endpoint - **Internal by default.** The Network Load Balancer sits in the node subnets, or in `api_load_balancer_subnets`, which must cover every node zone: a Network Load Balancer sends no traffic to targets in a zone it is not in. An internal endpoint admits the VPC’s primary CIDR unless you set `api_allowed_cidrs`; add the management cluster’s network there when it is outside the VPC. - **Public.** With `api_load_balancer_public = true` you must name public subnets in `api_load_balancer_subnets`, and `api_allowed_cidrs` must list the nodes’ public egress addresses (the NAT gateways’ Elastic IPs, as `/32`s) besides the clients: the endpoint’s DNS name resolves to public addresses, so the nodes’ own traffic arrives from those. - **Hairpin.** The target groups turn client IP preservation off, because a Network Load Balancer does not hairpin a target’s traffic back to itself while it preserves client IPs. The load balancer then connects to the control-plane nodes from its own addresses, which its security group stands for in the control-plane rules. - **The endpoint guard** records the scheme (`api_load_balancer_public`), the port (`cluster_network.api_server_port`) and the load balancer’s subnets. Subnets can be added later, in place; a subnet present at creation can never be removed. Removing a subnet added later is not caught by the guard and replaces the load balancer, which only CAPTF’s [destructive-plan guard]() stops. [Shared Behavior]() covers the rest: your own endpoint, the port, and RKE2’s supervisor port. ## Exports The cluster role’s `exports` carry the schema `captf.io/aws-cluster/v1`. | Key | Value | | --- | --- | | `schema` | `captf.io/aws-cluster/v1` | | `region` | The region the provider resolved | | `vpc_id` | The VPC | | `kubernetes_cluster_id` | The cluster ID the cloud controller manager reads from the `kubernetes.io/cluster/` tag | | `failure_domains` | Zone name to `{ subnet_id }` | | `security_group_ids` | `{ control_plane = [, ], worker = [] }` | | `instance_profiles` | `{ control_plane = , worker = }` | | `api` | `{ host, port, target_groups = { kube_apiserver = { arn, port }, rke2_supervisor = ... } }`: the endpoint and, per target group, the backend port to register; null with a user-supplied endpoint | | `bootstrap_bucket` | The S3 bucket machines and pools stage bootstrap data in | ## Tags AWS accepts the `captf.io/` tag keys unchanged. Every taggable resource carries `captf_tags` merged over `additional_tags`, so yours cannot override them; `additional_tags` takes at most 40 tags, from the AWS tag character set, and no `aws:`, `captf.io/` or `kubernetes.io/cluster/` keys. Instances, their volumes, the node security group and the Auto Scaling group also carry `kubernetes.io/cluster/ = owned`, which cloud-provider-aws needs on exactly one security group per instance. Cannot carry tags: `aws_iam_role_policy`, `aws_iam_role_policy_attachment`, the bucket’s public access block, encryption configuration and policy, target group attachments, and scaling policies. S3 objects take at most 10 tags, so the bootstrap objects carry only `captf_tags`. > [!NOTE] > > **Design notes** > > - **Bootstrap data is staged in S3.** A pool’s bootstrap data rotates about every 7.5 minutes, and each user-data change would add a launch template version, of which AWS allows 10,000: about 52 days. Staging also keeps the control-plane CA keys out of instance metadata. > - **Workers cannot read control-plane payloads.** The bucket policy grants each node role its own key prefix and explicitly denies the worker role `control-plane/*`, whatever its own policies allow. > - **The bucket policy grants the node roles, not inline role policies,** so brought instance profiles need no S3 permission of their own; their role ARNs are inputs, never read, so deleting them first never blocks a destroy. > - **The AMI lookup returns nothing rather than failing** once an image is deregistered, so a refresh never fails on it; a precondition reports a missing image when one is needed. > - **One security group carries the cluster tag.** cloud-provider-aws fails when an instance has two security groups tagged for the cluster, so the control-plane and load balancer groups go without it. > - **Pools roll only on a version change**, by replacing the launch template, which starts a launch-before-terminate instance refresh; every other change adds a version under `$Latest` for new instances only. [`DESIGN.md`]() records the evidence and the alternatives rejected. > [!WARNING] > > **Not yet verified** > > - cloud-init merging the staged payload from `cloud.cfg.d` as system configuration on a real boot, and the AWS CLI v2 being on the AMIs used. > - The kubelet merging `--node-labels` with your `kubeletExtraArgs`, and RKE2 appending `node-label+` labels. > - Instance refresh rounding at 100 % and 200 % healthy for very small groups. > - Deregistering a target whose instance is already terminated, at destroy. > - `Service` load balancers left by the cloud controller manager blocking the node security group’s deletion. > - The Network Load Balancer’s 350-second idle timeout with long `kubectl logs -f` sessions. > - The NAT gateways’ Elastic IPs being the source of node traffic to a public endpoint in every routing setup. > - `examples/identity-policy.json` being complete for create, update and destroy of all three roles. > - An Ignition 3.0.0 stub replacing itself with a config of a newer spec. > - An Auto Scaling group launching Spot through the launch template’s market options. # Cluster The `ghcr.io/captf-io/aws-cluster` image implements the [cluster role](). In the VPC and subnets you bring, it creates the API endpoint (a Network Load Balancer), the security groups, the node identities and the S3 bucket the nodes fetch their bootstrap data from, and publishes them in its `exports` for the [machine]() and [machinepool]() roles. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `aws_lb.api_load_balancer` | Network Load Balancer for the Kubernetes API | Without a user-supplied endpoint | | `aws_lb_target_group.api_target_groups` | TCP target groups on the backend ports: the API port with kubeadm, 6443 for RKE2’s kube-apiserver, 9345 for the RKE2 supervisor | One per API port, with the load balancer | | `aws_lb_listener.api_listeners` | TCP listeners on the endpoint ports, forwarding to the target groups | One per API port, with the load balancer | | `terraform_data.api_endpoint_guard` | Records the scheme, port and subnets; fails any plan that would change them | With the load balancer | | `aws_security_group.api_load_balancer_security_group` | The load balancer’s security group, given at creation | With the load balancer | | `aws_security_group.control_plane_security_group` | Control-plane nodes: the API backend ports | Always | | `aws_security_group.node_security_group` | Every node; the only group tagged for the cloud controller manager | Always | | `aws_vpc_security_group_ingress_rule.node_ingress_rules` | All traffic between the cluster’s nodes; SSH from `ssh_allowed_cidrs` | Always; SSH rules per CIDR | | `aws_vpc_security_group_egress_rule.node_egress_rules` | All outbound traffic from the nodes | Always | | `aws_vpc_security_group_ingress_rule.control_plane_ingress_rules` | API backend ports from the load balancer and the IPv4 pod CIDRs, or from `api_allowed_cidrs` with a user-supplied endpoint | Per port and source | | `aws_vpc_security_group_ingress_rule.api_load_balancer_ingress_rules` | Endpoint ports from the nodes and `api_allowed_cidrs` | With the load balancer | | `aws_vpc_security_group_egress_rule.api_load_balancer_egress_rules` | Traffic and health checks to the control-plane nodes | With the load balancer | | `aws_iam_role.node_roles` | Control-plane and worker roles under `/captf/` | Each unless its instance profile is brought | | `aws_iam_role_policy.node_role_policies` | The cloud-provider-aws policies: the Node Policy for workers, the Control Plane and Node Policies for control-plane nodes | With each role | | `aws_iam_role_policy_attachment.node_role_policy_attachments` | The managed policies in `node_role_policy_arns` | Per policy, with each role | | `aws_iam_instance_profile.node_instance_profiles` | The instance profiles machines and pools launch with | With each role | | `aws_s3_bucket.bootstrap_bucket` | Holds the staged bootstrap data | Always | | `aws_s3_bucket_public_access_block.bootstrap_bucket_public_access` | Blocks every form of public access | Always | | `aws_s3_bucket_server_side_encryption_configuration.bootstrap_bucket_encryption` | SSE-S3 encryption | Always | | `aws_s3_bucket_policy.bootstrap_bucket_policy` | TLS only; each node role reads only its own key prefixes; workers denied the control-plane payloads | Always | It reads the VPC, the node subnets and the load balancer’s own subnets with listings that return nothing rather than fail, so a destroy still runs after your network is gone; it reads the VPC’s CIDR only when the VPC exists and an internal endpoint defaults to it. ## Inputs Contract inputs it uses: `captf_cluster` (the names), `captf_tags`, `control_plane_endpoint` (non-null skips the load balancer) and `cluster_network` (`api_server_port`, default 6443, and the `pods` CIDRs). It declares `captf_object`, `kubernetes_version` and `control_plane_initialized` without using them. User variables, from [variables.tf](): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_tags` | `map(string)` | `{}` | Extra tags for every taggable resource; at most 40, no `aws:`, `captf.io/` or `kubernetes.io/cluster/` keys. | | `api_allowed_cidrs` | `list(string)` | `[]` | IPv4 networks allowed to reach the endpoint besides the nodes. Empty: the VPC’s primary CIDR for an internal endpoint. Required for a public one, and must then include the nodes’ NAT gateway Elastic IPs. | | `api_load_balancer_public` | `bool` | `false` | Make the load balancer internet-facing. | | `api_load_balancer_subnets` | `map(string)` | `{}` | Zone to subnet ID for the load balancer; empty means `subnets`. Must cover every zone of `subnets`; required (public subnets) for a public endpoint. | | `control_plane_instance_profile` | `object({ name = string, role_arn = string })` | `null` | An existing instance profile for control-plane nodes, with its role’s ARN, instead of creating one. | | `distribution` | `string` | `"kubeadm"` | `kubeadm` or `rke2`; `rke2` adds the supervisor listener on 9345. | | `node_role_permissions_boundary` | `string` | `null` | Permissions boundary ARN for the roles the module creates; it must allow `s3:GetObject` on the bootstrap bucket. | | `node_role_policy_arns` | `object({ control_plane = optional(list(string), []), worker = optional(list(string), []) })` | `{}` | Managed policies to attach to the roles the module creates, for example `AmazonEBSCSIDriverPolicy`. | | `region` | `string` | `null` | The region; null uses `AWS_REGION` from the identity Secret. | | `ssh_allowed_cidrs` | `list(string)` | `[]` | IPv4 networks allowed to reach the nodes on SSH; none by default. | | `subnets` | `map(string)` | `null` | Required. Zone to node subnet ID, one subnet per zone; each zone is a failure domain. | | `vpc_id` | `string` | `null` | Required. The VPC of the subnets. | | `worker_instance_profile` | `object({ name = string, role_arn = string })` | `null` | An existing instance profile for workers, with its role’s ARN; its role must differ from the control-plane one. | ## Outputs | Output | Value | | --- | --- | | `control_plane_endpoint` | The load balancer’s DNS name and `api_server_port`, or the input passed through | | `failure_domains` | One per subnet zone, sorted, each eligible for the control plane, with the `subnet_id` attribute | | `exports` | `captf.io/aws-cluster/v1`; see [AWS]() | | `health` | See below | | `api_load_balancer_id` (extra) | The load balancer’s ARN | ## Health | AWS state | Contract state | Reason | | --- | --- | --- | | The load balancer was deleted out of band (when the module owns the endpoint) | `terminated` | `LoadBalancerNotFound` | | The bootstrap bucket was deleted out of band | `degraded` | `BootstrapBucketNotFound` | | Both exist | `running`, healthy | none | Target health is never consulted; see [Shared Behavior](). ## Limitations > [!CAUTION] > > **Service load balancers outlive the cluster** > > `Service` load balancers created by cloud-provider-aws are not deleted with the cluster; delete the `Service` objects first. - The endpoint is fixed once the load balancer exists: the scheme, the port and the subnets present at creation cannot change. Removing a subnet added later replaces the load balancer, which only CAPTF’s destructive-plan guard stops. - Nodes accept all traffic from each other, so a worker can reach control-plane ports such as etcd. - Any principal of the account with broad S3 read access, other than the worker role, can read the control-plane payloads. - The node roles have fixed names derived from the cluster’s namespace and name, so two management clusters cannot create the same cluster in one account. - Nodes egress everywhere (`0.0.0.0/0`); restrict egress in the network you bring. ## Exceptions None: the role follows the conventions as written. ## Example terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo spec: source: image: ghcr.io/captf-io/aws-cluster:v0.1.0-opentofu identityRef: name: aws defaults: identityRef: name: aws variables: region: us-east-1 vpc_id: vpc-0123456789abcdef0 subnets: us-east-1a: subnet-0aaa0000000000001 us-east-1b: subnet-0bbb0000000000002 us-east-1c: subnet-0ccc0000000000003 ``` # Machine The `ghcr.io/captf-io/aws-machine` image implements the [machine role](): one EC2 instance per `Machine`, control plane or worker. It takes the subnets, security groups, instance profiles, API target groups and bootstrap bucket from the [cluster]() role’s exports. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `aws_instance.node_instance` | The node, in its zone’s subnet, with IMDSv2 and an encrypted root volume | Always | | `aws_s3_object.bootstrap_object` | The bootstrap data, staged as `control-plane/` or `worker/` in the cluster’s bucket | `bootstrap_delivery = "s3"` (the default) | | `aws_lb_target_group_attachment.api_target_attachments` | Registers the instance behind the API endpoint, so destroying the machine deregisters it | Control-plane machines, one per API port | Without `machine_image.id` it looks up the image with a listing that returns nothing rather than fails. ## Inputs Contract inputs it uses: `captf_cluster_outputs`, `captf_tags`, `machine_name` (the `Name` tag and the object key), `bootstrap_data`, `bootstrap_format`, `failure_domain`, `kubernetes_version` (the image lookup) and `control_plane`. With no `failure_domain` it picks a zone from the hash of `machine_name`. User variables, from [variables.tf](): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_security_group_ids` | `list(string)` | `[]` | Extra security groups for the instance. | | `additional_tags` | `map(string)` | `{}` | Extra tags for every taggable resource; at most 40, no `aws:`, `captf.io/` or `kubernetes.io/cluster/` keys. | | `bootstrap_delivery` | `string` | `"s3"` | `s3` stages the bootstrap data behind a small stub; `inline` sends it as user data, 16 KiB at most. | | `external_cluster_exports` | `any` | `null` | The exports of an externally managed `TerraformCluster`, used when `captf_cluster_outputs` is empty. | | `instance_metadata_hop_limit` | `number` | `1` | IMDSv2 hop limit; 2 lets pods without host networking reach the metadata service. | | `instance_type` | `string` | `"m6i.large"` | EC2 instance type; matches the image’s capacity labels. | | `machine_image` | `object({ id = optional(string), owner = optional(string, "819546954734"), name_format = optional(string, "capa-ami-ubuntu-24.04-?{semver}-*"), architecture = optional(string, "x86_64") })` | `{}` | `id` pins the AMI; without it, the newest AMI from `owner` named `name_format`, with `{semver}` or `{version}` filled in. | | `public_ip` | `bool` | `false` | Give the instance a public IPv4 address. | | `root_volume_kms_key_id` | `string` | `null` | KMS key ARN for the root volume; null uses the account’s default EBS key. | | `root_volume_size_gib` | `number` | `40` | Root volume size. | | `root_volume_type` | `string` | `"gp3"` | `gp3` or `gp2`. | | `spot` | `bool` | `false` | Run a one-time Spot Instance; refused for control-plane machines. | | `ssh_key_name` | `string` | `null` | EC2 key pair; the cluster’s `ssh_allowed_cidrs` decides who reaches SSH. | ## Outputs | Output | Value | | --- | --- | | `provider_id` | `aws:////`, the format cloud-provider-aws writes to the `Node` | | `addresses` | `InternalIP`, `InternalDNS` and `Hostname` from the private IP and DNS name; `ExternalIP` and `ExternalDNS` with a public address | | `failure_domain` | The zone the instance runs in | | `interruptible` | `true` for a Spot Instance | | `health` | See below | ## Health | EC2 instance state | Contract state | Reason | | --- | --- | --- | | `pending` | `pending` | `InstancePending` | | `running` | `running`, healthy | none | | `stopping` | `stopped` | `InstanceStopping` | | `stopped` | `stopped` | `InstanceStopped` | | `shutting-down` | `terminated` | `InstanceShuttingDown` | | `terminated` | `terminated` | `InstanceTerminated` | | Gone from state | `terminated` | `InstanceNotFound` | | Anything else | `unknown` | `UnknownState` | The provider drops a terminated instance from state on refresh, so a terminated machine usually reads `InstanceNotFound`. ## Lifecycle A machine is immutable. Changes to the looked-up image and to the user-data stub (a newer module release) are ignored on an existing instance instead of stopping and starting it; they reach the cluster through new machines. ## Bootstrap - **Staged (the default).** The bootstrap data goes to the cluster’s S3 bucket, readable only by the matching node role: control-plane nodes read `control-plane/*`, workers `worker/*`. The user data is a small `#cloud-boothook` that fetches it once and installs it in `/etc/cloud/cloud.cfg.d`, as [Shared Behavior]() describes. With Ignition, the user data is a config that replaces itself with the object; gzipped Ignition is refused. - **Inline.** `bootstrap_delivery = "inline"` sends the data as user data, a plain cloud-config gzipped, anything else as it is, up to 16 KiB. Anyone who can call `ec2:DescribeInstanceAttribute` or reach the instance metadata service can then read it, control-plane CA keys included. ## Limitations > [!WARNING] > > **Nodes must register under their private DNS names** > > Nodes must register under their private DNS names, which cloud-provider-aws looks them up by. - Spot interruptions are not drained beyond what a termination handler in the cluster does. - The staged data stays on the node as `/etc/cloud/cloud.cfg.d/99-captf-bootstrap.cfg`, root only, and in the Terraform state. - With a customer managed `root_volume_kms_key_id`, the identity needs `kms:CreateGrant`, `kms:Decrypt`, `kms:GenerateDataKeyWithoutPlaintext`, `kms:ReEncrypt*` and `kms:DescribeKey` on the key. ## Exceptions None: the role follows the conventions as written. ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 spec: template: spec: source: image: ghcr.io/captf-io/aws-machine:v0.1.0-opentofu variables: instance_type: m6i.large ``` # MachinePool The `ghcr.io/captf-io/aws-machinepool` image implements the [machinepool role](): one Auto Scaling group of worker instances per `MachinePool`. It takes the subnets, the worker security group and instance profile, and the bootstrap bucket from the [cluster]() role’s exports. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `aws_autoscaling_group.pool_autoscaling_group` | The group across the pool’s zones, launching the newest launch template version | Always | | `aws_launch_template.pool_launch_template` | Image, instance type, worker identity and security groups, IMDSv2, encrypted root volume, tags, and the user-data stub | Always; replaced only to roll | | `aws_autoscaling_policy.pool_scaling_policy` | Target tracking on average CPU | With autoscaling enabled | | `aws_s3_object.bootstrap_object` | The bootstrap data as `pool/`, rewritten in place on every rotation | With complete exports | | `terraform_data.kubernetes_version_roll` | The Kubernetes version and an explicit `machine_image.id`; replacing it replaces the launch template | Always | | `terraform_data.inherited_zones` | The cluster’s zones at the pool’s first apply | Always; used without `failure_domains` | It reads the members with one `aws_instances` listing per cluster zone and two group-wide listings for their states, and looks up the image like the [machine role](). ## Inputs Contract inputs it uses: `captf_object` (the namespace in the group’s name), `captf_cluster_outputs`, `captf_tags`, `machinepool_name`, `replicas`, `bootstrap_data`, `bootstrap_format`, `failure_domains`, `cluster_failure_domains`, `kubernetes_version`, `node_labels` and `autoscaling`. User variables, from [variables.tf](): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_security_group_ids` | `list(string)` | `[]` | Extra security groups for the instances. | | `additional_tags` | `map(string)` | `{}` | Extra tags for every taggable resource; at most 40, no `aws:`, `captf.io/` or `kubernetes.io/cluster/` keys. | | `autoscaling_target_cpu_percent` | `number` | `60` | Average CPU the scaling policy holds the group at while autoscaling is enabled. | | `external_cluster_exports` | `any` | `null` | The exports of an externally managed `TerraformCluster`, used when `captf_cluster_outputs` is empty. | | `instance_metadata_hop_limit` | `number` | `1` | IMDSv2 hop limit; 2 lets pods without host networking reach the metadata service. | | `instance_type` | `string` | `"m6i.large"` | EC2 instance type; applies to instances launched afterwards. | | `machine_image` | `object({ id = optional(string), owner = optional(string, "819546954734"), name_format = optional(string, "capa-ami-ubuntu-24.04-?{semver}-*"), architecture = optional(string, "x86_64"), root_device_name = optional(string, "/dev/sda1") })` | `{}` | `id` pins the AMI, and changing it rolls the instances; without it, the newest AMI from `owner` named `name_format`. `root_device_name` is the image’s root device: `/dev/sda1` for Ubuntu, `/dev/xvda` for Amazon Linux and Flatcar. | | `public_ip` | `bool` | `false` | Give the instances public IPv4 addresses. | | `rollout_instance_warmup_seconds` | `number` | `300` | Seconds a roll waits after each new instance is in service. | | `root_volume_kms_key_id` | `string` | `null` | KMS key ARN for the root volumes; a customer managed key must grant `AWSServiceRoleForAutoScaling`. | | `root_volume_size_gib` | `number` | `40` | Root volume size. | | `root_volume_type` | `string` | `"gp3"` | `gp3` or `gp2`. | | `spot` | `bool` | `false` | Launch Spot Instances, with capacity rebalancing. | | `ssh_key_name` | `string` | `null` | EC2 key pair; the cluster’s `ssh_allowed_cidrs` decides who reaches SSH. | ## Outputs | Output | Value | | --- | --- | | `provider_id` | The Auto Scaling group’s ARN | | `provider_id_list` | `aws:////` of every pending, running, stopping or stopped instance tagged with the group’s name, sorted | | `replicas` | The group’s desired capacity as last observed | | `instances` | Per member: provider ID, instance ID, `InternalIP` address, zone and state | | `health` | See below | | `autoscaling_group_name` (extra) | The group’s name | | `dropped_node_labels` (extra) | The labels left out of the kubelet’s registration | | `launch_template_id` (extra) | The current launch template; it changes only when the instances roll | ## Health | AWS state | Contract state | Reason | | --- | --- | --- | | The group was deleted out of band | `terminated` | `GroupNotFound` | | Desired capacity 0 | `running`, healthy | none | | No member yet | `pending` | `NoMembers` | | A member `stopping` or `stopped` | `stopped` | `InstanceStopped:` per member | | Every member running, as many as desired | `running`, healthy | none | | Members still `pending`, or fewer or more than desired | `running`, unhealthy | `InstancePending:` per member, `ScalingInProgress` | ## Lifecycle | Change | Launch template | Running instances | | --- | --- | --- | | `bootstrap_data` (about every 7.5 minutes) | Unchanged; only the S3 object is rewritten | Kept | | `node_labels`, `instance_type`, a looked-up image, `spot`, root volume, security groups | New version under `$Latest` | Kept; new instances get it | | `kubernetes_version`, its `+rke2rN` suffix included, or an explicit `machine_image.id` | Replaced, create before destroy | Replaced by a rolling instance refresh that launches each replacement first | | `replicas`, with autoscaling off | Unchanged | The group grows or shrinks | | A zone the cluster adds or removes, for a pool without `failure_domains` | Unchanged | Kept: the pool keeps the cluster’s zones of its first apply | | A zone added to `failure_domains` | Unchanged | Kept: new instances go there over time (zone rebalancing is suspended) | | A zone removed from `failure_domains` | Unchanged | The instances in it are replaced elsewhere, undrained | With autoscaling enabled, the group’s minimum and maximum come from the `MachinePool` annotations and an apply never resets the desired capacity; with it off, the minimum and maximum equal `replicas`. ## Bootstrap The bootstrap data is staged in S3 as `pool/`, readable by the worker role, and rewritten in place on every rotation, so the launch template never changes with it. The user data is a `#cloud-boothook` that writes the node labels on every boot and, once per instance, fetches the data and installs it in `/etc/cloud/cloud.cfg.d`, as [Shared Behavior]() describes. The stub must stay within EC2’s 16 KiB of user data, which only many node labels can exceed. Pools have no inline delivery. With Ignition the stub is a config that replaces itself with the object, and node labels are refused. ## Limitations > [!WARNING] > > **Scale-in, a roll and Spot rebalancing do not drain nodes** > > Scale-in, a roll and Spot rebalancing terminate instances without draining their nodes. Run a termination handler such as aws-node-termination-handler with an Auto Scaling lifecycle hook if your workloads need it. - CPU target tracking is a proxy: pods waiting for capacity do not raise CPU. The Kubernetes Cluster Autoscaler cannot drive these pools. - Running members keep the labels they registered with until they are replaced. - Membership is read in every cluster zone on every refresh: one `DescribeInstances` call per zone, plus two for member states. - The staged data stays in the Terraform state. ## Exceptions - A change of an explicit `machine_image.id` rolls the instances: the pinned image carries the kubelet, and with ClusterClass it can render in a later apply than the version. - Removing a zone from `failure_domains` replaces the instances in it: the group cannot keep instances in a zone it no longer spans. ## Example terraformmachinepool.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: demo-pool-0 labels: cluster.x-k8s.io/cluster-name: demo spec: source: image: ghcr.io/captf-io/aws-machinepool:v0.1.0-opentofu variables: instance_type: m6i.large ``` # Google Cloud The Google Cloud modules create, in a VPC network you bring, a Cluster API cluster’s API load balancer (an internal proxy Network Load Balancer by default), its firewall rules and its node service accounts; one Shielded VM per `Machine`; and one regional managed instance group per `MachinePool`. They live in [`captf-io/gcp-modules`]() and pin the `hashicorp/google` provider at 8.5.0. Like every reference module set, they are pre-release: read the status note in [Cloud Modules]() first. - **Cluster** --- The API endpoint, security rules, node identities and exports of one workload cluster. - **Machine** --- One instance per `Machine`, registered with the API load balancer. - **MachinePool** --- One native scaling group per `MachinePool`. ## Images | Role | Image | Page | | --- | --- | --- | | cluster | `ghcr.io/captf-io/gcp-cluster` | [Cluster]() | | machine | `ghcr.io/captf-io/gcp-machine` | [Machine]() | | machinepool | `ghcr.io/captf-io/gcp-machinepool` | [MachinePool]() | The machine image’s capacity labels describe its default machine type, `n2-standard-4`: 4 vCPU and 16 GiB on `amd64`. ## Prerequisites ### Network You bring the network; the modules create nothing in it but firewall rules. - A VPC network and a regional subnetwork for the nodes. The internal API address comes from the same subnetwork. - Cloud NAT on that subnetwork’s router: nodes have no external address and need egress for container images and packages. - For the default internal endpoint, a proxy-only subnet in the same network and region (`--purpose=REGIONAL_MANAGED_PROXY --role=ACTIVE`, an unused /23). One serves every Envoy-based load balancer there; the plan fails, with the `gcloud` command in its message, when there is none. - Pod and Service CIDRs that overlap neither the subnetworks nor the proxy-only subnet. An auto-mode network uses 10.128.0.0/9. - The management cluster must reach the internal API address: from the same VPC, a peered network, a VPN or Interconnect. Clients in other regions need `api_global_access`. - Shared VPC: set the cluster’s `network_project` to the host project. The firewall rules go there, and cloud-provider-gcp needs `network-project-id` in its `gce.conf`. ### Permissions of the identity The role set below is expected to suffice; no real project has confirmed it yet. | Role | Granted on | For | | --- | --- | --- | | `roles/compute.loadBalancerAdmin` | the project | Addresses, forwarding rules, backend services, health checks, target proxies | | `roles/compute.instanceAdmin.v1` | the project | Instance groups, instances and memberships, instance templates, managed instance groups, autoscalers | | `roles/compute.securityAdmin` | the project and the network’s project | Firewall rules, the Cloud Armor policy | | `roles/compute.networkViewer` | the network’s project | Listing the subnetworks | | `roles/compute.networkUser` | the node subnetwork (Shared VPC only) | Instances in the host project’s subnetwork | | `roles/iam.serviceAccountAdmin` | the project | Creating the node service accounts (not needed when you bring both) | | `roles/resourcemanager.projectIamAdmin` | the project | Granting the node service accounts their roles (not needed when you bring both) | | `roles/iam.serviceAccountUser` | the two node service accounts only | Running instances as them; never grant it project-wide | | `roles/secretmanager.admin` | the project | Staging control-plane bootstrap data | Enable the Compute Engine and Secret Manager APIs in the project. ### Node images Build an image per Kubernetes version, for example with [image-builder](), and name it so that a placeholder finds it: `capi-ubuntu-2404-v1-33-4` for `projects//global/images/capi-ubuntu-2404-{slug}`. The image needs cloud-init (or Ignition), and `curl`, `sed`, `base64` and `gzip` for staged control-plane bootstrap data. Secure Boot is on by default, so its kernel modules must be signed. ### In the workload cluster - [cloud-provider-gcp](), the cloud controller manager, with a `gce.conf` from the cluster’s exports: `project-id`, `network-name`, `subnetwork-name`, and `node-tags = `. The kubelets run with `cloud-provider: external`, so the CCM writes the `gce://` provider IDs the modules report. - The [Compute Engine persistent disk CSI driver](), whose controller runs as the control-plane service account. - A CNI. ## Identity Secret | Key | Value | | --- | --- | | `GOOGLE_CREDENTIALS` | A service account key JSON, or an `external_account` (workload identity federation) configuration | | `GOOGLE_PROJECT` | The default project, when the `TerraformCluster` does not set `project` | | `GOOGLE_REGION` | The default region, when the `TerraformCluster` does not set `region` | Instead of `GOOGLE_CREDENTIALS`, a `credentials.json` key plus `GOOGLE_APPLICATION_CREDENTIALS=/var/run/captf/credentials/credentials.json` works too. From [examples/identity.yaml](): identity.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: gcp spec: secretRef: name: gcp namespace: captf-system allowedNamespaces: list: - team-a --- apiVersion: v1 kind: Secret metadata: name: gcp namespace: captf-system type: Opaque stringData: GOOGLE_PROJECT: my-project GOOGLE_REGION: us-central1 GOOGLE_CREDENTIALS: | { "type": "service_account", "project_id": "replace-me", "private_key_id": "replace-me", "private_key": "replace-me", "client_email": "captf@replace-me.iam.gserviceaccount.com", "client_id": "replace-me", "token_uri": "https://oauth2.googleapis.com/token" } ``` See [Identities and Credentials]() for how the Secret reaches a Job. ## Quick start 1. Create the network: VPC, node subnetwork, Cloud NAT and the proxy-only subnet. 2. Build a node image per Kubernetes version, named for the `{slug}` placeholder. 3. Create the identity: ```sh export NAMESPACE=team-a GCP_PROJECT=my-project GCP_REGION=us-central1 clusterctl generate yaml --from examples/identity.yaml | kubectl apply -f - ``` 4. Create the cluster: a `TerraformCluster`, a `KubeadmControlPlane`, a `MachineDeployment` of workers and an autoscaled `MachinePool`: ```sh export CLUSTER_NAME=demo KUBERNETES_VERSION=v1.33.4 \ GCP_NETWORK=captf-vpc GCP_SUBNETWORK=captf-nodes \ GCP_IMAGE='projects/my-images/global/images/capi-ubuntu-2404-{slug}' clusterctl generate yaml --from examples/cluster-kubeadm.yaml | kubectl apply -n team-a -f - ``` 5. Once the API server answers, install cloud-provider-gcp, the persistent disk CSI driver and a CNI in the workload cluster. [examples/README.md]() lists every variable of the two example files. ## API endpoint - **Internal (default).** A regional internal proxy Network Load Balancer with a reserved internal address in the node subnetwork. Clients reach it from inside the VPC, or from any region with `api_global_access`. - **Public.** With `api_load_balancer_public = true`, a global external proxy Network Load Balancer on a global address. A VPC firewall cannot filter clients behind a proxy, so a Cloud Armor policy allows only `api_allowed_cidrs` and denies the rest. Nodes reach the endpoint through Cloud NAT, so its static egress addresses must be in `api_allowed_cidrs`. - **Hairpin.** Both are proxies: the proxy opens a new connection to a healthy backend, so a control-plane node reaches the endpoint it is itself behind. An internal passthrough load balancer would route the node’s packets back to the node itself. - **Backends.** One unmanaged instance group per zone; control-plane machines join their zone’s group in their own state. - **The endpoint guard** records `api_load_balancer_public`, `network`, `subnetwork`, `project`, `region`, the API port and the address when the load balancer is created, and fails any later plan that would change one of them. See [Shared Behavior](). ## Exports Schema `captf.io/gcp-cluster/v1`: | Key | Value | | --- | --- | | `schema` | `"captf.io/gcp-cluster/v1"` | | `project`, `region` | Where machines and pools are created | | `network`, `subnetwork` | Self links of the network and the node subnetwork | | `name_prefix` | `captf--`, truncated, plus 8 hex characters of the cluster key’s sha256 | | `failure_domains` | `{"" = {zone = ""}}` | | `node_network_tag` | `-node`, on every node: the `node-tags` value for cloud-provider-gcp | | `control_plane`, `worker` | `{service_account = , network_tags = [, ]}` | | `api` | `{host, port, backend_port, instance_groups = {"" = }}`; `null` with a user endpoint | `api.port` is the endpoint’s port; `api.backend_port` is the kube-apiserver port on the nodes, 6443 with RKE2. Both listeners share the per-zone instance groups, so there is no per-listener registration target. ## Tags GCP labels: keys match `[a-z][a-z0-9_-]{0,62}` and values `[a-z0-9_-]{0,63}`, so the `captf_tags` keys become `captf-io_cluster` and so on, and values are lowercased with invalid characters turned into `_`. A value over 63 characters keeps 54 characters, then `-` and 8 hex characters of the original’s sha256. `additional_tags` must already be valid labels, may not start with `captf-io_`, and holds at most 58 entries (GCP allows 64 per resource). These resource types cannot carry labels: firewall rules, instance groups, managed instance groups, autoscalers, health checks, backend services, target proxies, service accounts and IAM members. Their names carry the cluster’s or pool’s name prefix, and their descriptions name the owning object. > [!NOTE] > > **Design notes** > > - **A proxy load balancer, not passthrough,** because only a proxy lets a control-plane node reach the endpoint it is behind. > - **Firewall rules match service accounts, not network tags,** so an instance cannot join the cluster’s traffic by setting a tag. Nodes still carry a node tag, which cloud-provider-gcp needs for Service firewall rules. > - **Node service accounts use predefined roles only:** a deleted custom role’s ID cannot be reused for weeks, which would break recreating a cluster of the same name. > - **Descriptions never hold `captf_tags`:** `description` forces replacement on addresses, forwarding rules and instance groups, and `captf.io/template` changes on a ClusterClass rebase. > - **Control-plane bootstrap data is staged in Secret Manager:** instance metadata is readable through `compute.instances.get`, which `roles/viewer` includes. > - **A pool’s bootstrap data lives in the group’s all-instances config,** so a rotation is a metadata refresh that never restarts or replaces an instance. > - **A Kubernetes version change rolls a pool through the image name:** the image must carry a version placeholder, because a template change that only touches metadata is applied as a refresh. [DESIGN.md]() has the evidence for each, and the alternatives rejected. > [!WARNING] > > **Not yet verified** > > - The proxy-only subnet filter `purpose = "REGIONAL_MANAGED_PROXY"` on a real project. > - The minimal role set for cloud-provider-gcp and the PD CSI driver, and the identity’s role set. > - That the managed instance group applies an all-instances config change as a refresh. > - Who can read control-plane bootstrap data that is still in instance metadata (Ignition, or `bootstrap_delivery = "inline"`). > - That `DEPROVISIONING` appears only while an instance is deleted, not while it stops. > - cloud-init’s handling of a gzip part, and of a `## template: jinja` payload, inside a pool’s multipart user data. > - That the managed instance group repairs a preempted, stopped Spot VM. > - Firewall rules naming a service account created moments earlier. > - The staged bootstrap fetch on a real image, and IAM propagation within its five minutes of retries. > - Secret Manager replication under a resource-location policy, and the secret’s deletion with the machine. > - That a kubeadm payload behaves the same as cloud-init system configuration as it does as user data. # Cluster The `ghcr.io/captf-io/gcp-cluster` image implements the [cluster role]() for a `TerraformCluster` on Google Cloud. It creates the API load balancer, the firewall rules and the node service accounts in the network you bring, and publishes the endpoint, one failure domain per zone and the exports the machine and machinepool roles read. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `google_service_account.node_service_accounts` | Control-plane and worker node service accounts | Each one unless `control_plane_service_account` or `worker_service_account` brings it | | `google_project_iam_member.node_project_roles` | `control_plane_roles` and `worker_roles` on the project | For the accounts the module created | | `google_service_account_iam_member.node_service_account_users` | `roles/iam.serviceAccountUser` for the control-plane account, so the PD CSI controller can attach disks | On each account the module created | | `google_compute_firewall.node_internal_firewall` | Every protocol between the node service accounts | Always | | `google_compute_firewall.pod_ingress_firewalls` | Pod CIDRs to the nodes, one rule per address family | When `Cluster.spec.clusterNetwork.pods` is set | | `google_compute_firewall.api_ingress_firewall` | API backend ports on control-plane nodes from the load balancer and Google’s health checkers | Without a user endpoint | | `google_compute_instance_group.api_instance_groups` | One unmanaged instance group per zone, the API backends | Without a user endpoint | | `terraform_data.api_endpoint_guard` | Records what decides the endpoint and fails a plan that would change it | Without a user endpoint | | `google_compute_address.api_address` | The internal API address, shared by both listeners | Internal endpoint (default) | | `google_compute_region_health_check.api_region_health_checks` | One TCP health check per listener | Internal endpoint | | `google_compute_region_backend_service.api_region_backend_services` | One `INTERNAL_MANAGED` backend service per listener, over every zone’s instance group | Internal endpoint | | `google_compute_region_target_tcp_proxy.api_region_target_tcp_proxies` | One target proxy per listener | Internal endpoint | | `google_compute_forwarding_rule.api_forwarding_rules` | One forwarding rule per listener, on the API address | Internal endpoint | | `google_compute_global_address.api_global_address` | The public API address | `api_load_balancer_public` | | `google_compute_security_policy.api_security_policy` | Cloud Armor allowlist of `api_allowed_cidrs`, deny for everyone else | `api_load_balancer_public` | | `google_compute_health_check.api_health_checks` | One TCP health check per listener | `api_load_balancer_public` | | `google_compute_backend_service.api_backend_services` | One `EXTERNAL_MANAGED` backend service per listener | `api_load_balancer_public` | | `google_compute_target_tcp_proxy.api_target_tcp_proxies` | One target proxy per listener | `api_load_balancer_public` | | `google_compute_global_forwarding_rule.api_global_forwarding_rules` | One forwarding rule per listener, on the public address | `api_load_balancer_public` | The listeners are `kube_apiserver` and, with `distribution = "rke2"`, `rke2_supervisor` on port 9345. Backend services have a 3600-second idle timeout: a TCP proxy’s default of 30 seconds cuts idle watches and `kubectl logs -f`. The role reads the provider’s project and region, the node subnetwork and the proxy-only subnets through listings, and the region’s zones. It does not read the network: its self link is built from `network` and `network_project`. ## Inputs Contract inputs it uses: - `captf_cluster`: names and descriptions. - `captf_tags`: labels. - `control_plane_endpoint`: a non-null value means no load balancer. - `cluster_network`: `pods` for the pod firewall rules, `api_server_port` for the endpoint’s port. `captf_contract` is validated; `captf_object`, `kubernetes_version`, `control_plane_initialized` and `captf_cluster_outputs` are declared and unused. User variables, set with `spec.variables` ([variables.tf]()): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_tags` | `map(string)` | `{}` | Extra GCP labels for every labelable resource. Keys and values must already be valid GCP labels; the `captf-io_` keys are reserved for `captf_tags`, which win. | | `api_allowed_cidrs` | `list(string)` | `[]` | Client CIDRs allowed to reach the public API endpoint, enforced by a Cloud Armor policy. Required when `api_load_balancer_public` is true; include the Cloud NAT egress addresses so nodes can reach the endpoint. | | `api_global_access` | `bool` | `false` | Let clients in any region of the VPC reach the internal API endpoint. Off by default: only clients in the cluster’s region can. | | `api_load_balancer_public` | `bool` | `false` | Serve the API through a global external proxy load balancer instead of the internal one. Off by default; requires `api_allowed_cidrs`. | | `control_plane_roles` | `list(string)` | `["roles/compute.instanceAdmin.v1", "roles/compute.loadBalancerAdmin", "roles/compute.securityAdmin", "roles/compute.storageAdmin", "roles/compute.viewer", "roles/logging.logWriter", "roles/monitoring.metricWriter"]` | Project roles for a module-created control-plane service account: what cloud-provider-gcp and the PD CSI controller need. Ignored with `control_plane_service_account`. | | `control_plane_service_account` | `string` | `null` | Email of an existing service account for control-plane nodes. Null creates one with `control_plane_roles`. | | `distribution` | `string` | `"kubeadm"` | Kubernetes distribution of the control plane: `kubeadm`, or `rke2`, which adds the RKE2 supervisor port 9345 on the API address and always uses kube-apiserver port 6443 on the nodes. | | `network` | `string` | `null` | Name of the existing VPC network the cluster runs in. Required. | | `network_project` | `string` | `null` | Project that owns the network: the host project of a Shared VPC. Null means the cluster’s own project. Firewall rules are created here. | | `project` | `string` | `null` | Project to create the cluster in. Null uses the provider’s project: `GOOGLE_PROJECT` from the identity Secret, or the credentials’ project. | | `region` | `string` | `null` | Region of the cluster. Null uses the provider’s region: `GOOGLE_REGION` from the identity Secret. | | `subnetwork` | `string` | `null` | Name of the existing regional subnetwork, in `network`, that nodes and the internal API address use. Required. | | `worker_roles` | `list(string)` | `["roles/logging.logWriter", "roles/monitoring.metricWriter"]` | Project roles for a module-created worker service account: logs and metrics only. Ignored with `worker_service_account`. | | `worker_service_account` | `string` | `null` | Email of an existing service account for worker nodes. Null creates one with `worker_roles`. | | `zones` | `list(string)` | `[]` | Zones of the region to use as failure domains. Empty means every zone of the region; set it when the machine type is not offered everywhere. | ## Outputs | Output | Value | | --- | --- | | `control_plane_endpoint` | The user’s endpoint when given; else the API address (internal or global) and `cluster_network.api_server_port`, 6443 by default | | `failure_domains` | One per zone: `{name = , control_plane = true, attributes = {zone = }}`. The zones are `zones`, or every zone of the region | | `exports` | The [exports]() object, schema `captf.io/gcp-cluster/v1` | | `health` | Below | Extra outputs, for operators and the tests: | Output | Value | | --- | --- | | `api_address_id` | ID of the internal or global API address; null with a user endpoint | | `api_backend_service_ids` | IDs of the backend services, by listener | | `api_forwarding_rule_ids` | IDs of the forwarding rules, by listener | | `api_instance_group_ids` | IDs of the per-zone instance groups, by zone | | `firewall_rule_ids` | IDs of the firewall rules, by purpose | | `node_service_account_ids` | IDs of the module-created node service accounts, by node role | | `proxy_subnet_cidrs` | Ranges of the proxy-only subnets the API firewall rule allows | ## Health From the cluster’s own resource, the `kube_apiserver` forwarding rule: | Situation | Contract state | Reason | | --- | --- | --- | | User endpoint: no load balancer | `running`, healthy | none | | The forwarding rule exists | `running`, healthy | none | | A refresh no longer finds the forwarding rule | `terminated` | `LoadBalancerNotFound` | ## Limitations > [!WARNING] > > **The endpoint is fixed once the load balancer exists** > > The endpoint is fixed once the load balancer exists; the endpoint guard refuses a later change of `api_load_balancer_public`, `network`, `subnetwork`, `project`, `region`, the API port or the address. Revert the change, or create a new cluster. - With a user endpoint the module creates no instance groups, so control-plane machines join nothing. - Removing a zone from `zones` fails while control-plane machines are still members of its instance group. Machine pools without their own failure domains keep the zones they were created with. - With a brought worker service account and a module-created control-plane account, grant the control-plane account `roles/iam.serviceAccountUser` on the worker account yourself; the module grants it only on accounts it created. - In a Shared VPC the module binds `control_plane_roles` in `project` only; cloud-provider-gcp also needs `roles/compute.loadBalancerAdmin` and `roles/compute.securityAdmin` in the host project. - No SSH firewall rule and no external addresses: use OS Login with IAP TCP forwarding or a bastion of your own. - The API address has no `prevent_destroy`, which would block deleting the cluster too; the endpoint guard and the [destructive-plan guard]() catch a replacement. ## Exceptions - `exports.api` carries per-zone `instance_groups` instead of one registration target per listener keyed `kube_apiserver` and `rke2_supervisor`: a Google Cloud backend is an instance group, and one group serves both listeners. ## Example terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo namespace: team-a spec: source: image: ghcr.io/captf-io/gcp-cluster:v0.1.0-opentofu identityRef: name: gcp defaults: identityRef: name: gcp variables: network: captf-vpc subnetwork: captf-nodes ``` # Machine The `ghcr.io/captf-io/gcp-machine` image implements the [machine role]() for a `TerraformMachine`, cloned from a `TerraformMachineTemplate`. It creates one Compute Engine instance per `Machine`, control plane or worker, and joins control-plane instances to the API load balancer. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `google_compute_instance.node_instance` | A Shielded VM with no external address, running as the cluster’s control-plane or worker service account | Always | | `google_compute_instance_group_membership.api_instance_group_membership` | Joins the API instance group of the instance’s zone, in the machine’s own state, so destroying the machine leaves the group | Control-plane machines of a cluster with a module-owned endpoint | | `google_secret_manager_secret.bootstrap_secret` | Holds the bootstrap data, replicated in the cluster’s region and labeled | Control-plane machines with `cloud-config` data and `bootstrap_delivery = "secret-manager"` | | `google_secret_manager_secret_version.bootstrap_secret_version` | The bootstrap data itself | Same | | `google_secret_manager_secret_iam_member.bootstrap_secret_accessor` | `roles/secretmanager.secretAccessor` for the control-plane service account on that one secret | Same | The instance runs `machine_type` from `image`, with a `boot_disk_size_gib` GiB boot disk of `boot_disk_type`, optionally encrypted with `boot_disk_kms_key_id`. It has Shielded VM with vTPM and integrity monitoring, Secure Boot unless `secure_boot` is off, OS Login on, project-wide SSH keys blocked and the serial console off. It runs with the `cloud-platform` scope, so the service account’s IAM roles alone decide its access, and carries the cluster’s node tag, its role tag and `additional_network_tags`. The instance name is the Node name, because cloud-provider-gcp looks instances up by Node name: `machine_name` when it is a valid Compute Engine name, else `machine_name` with invalid characters replaced, `m-` in front when it starts with a digit, cut to 54 characters, then `-` and 8 hex characters of its sha256. ## Inputs Contract inputs it uses: - `machine_name`: the instance name. - `bootstrap_data`, `bootstrap_format`: the instance’s user data or the staged secret. - `failure_domain`: the zone. Without one, the module picks a zone from the sha256 of `machine_name`. - `kubernetes_version`: fills the image placeholders. - `control_plane`: the service account, the network tag, staging and the API registration. - `captf_cluster_outputs`: the cluster’s exports. - `captf_cluster`, `captf_object`, `captf_tags`: descriptions and labels. User variables, set with `spec.template.spec.variables` on the `TerraformMachineTemplate` ([variables.tf]()): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_network_tags` | `list(string)` | `[]` | Extra network tags for the instance, on top of the cluster’s node and role tags. | | `additional_tags` | `map(string)` | `{}` | Extra GCP labels for the instance and its boot disk. Keys and values must already be valid GCP labels; the `captf-io_` keys are reserved for `captf_tags`, which win. | | `boot_disk_kms_key_id` | `string` | `null` | Cloud KMS key (`projects/.../cryptoKeys/...`) that encrypts the boot disk. Null uses Google-managed encryption; the Compute Engine service agent needs encrypt and decrypt on the key. | | `boot_disk_size_gib` | `number` | `50` | Boot disk size in GiB: room for the image, container images and logs. | | `boot_disk_type` | `string` | `"pd-balanced"` | Boot disk type. `pd-balanced` suits the default N2 machine type; C3, N4 and newer series need `hyperdisk-balanced`. | | `bootstrap_delivery` | `string` | `"secret-manager"` | How a control-plane machine’s `cloud-config` data reaches it: `secret-manager` stages it in a per-machine secret behind a small user-data script, which needs `curl`, `sed`, `base64` and `gzip` in the image; `inline` puts it in instance metadata. Workers and Ignition are always inline. | | `can_ip_forward` | `bool` | `false` | Let the instance send and receive packets for other addresses, as CNIs that route pod CIDRs through GCP routes need. | | `external_cluster_exports` | `any` | `null` | Exports (schema `captf.io/gcp-cluster/v1`) to use when the `TerraformCluster` is externally managed and `captf_cluster_outputs` is `{}`. Ignored otherwise. | | `image` | `string` | `null` | Boot image: a name, `family/`, `projects/

/global/images/` or a self link, with optional placeholders (below). Required. | | `machine_type` | `string` | `"n2-standard-4"` | Machine type of the instance. The default matches the image’s capacity labels (4 vCPU, 16 GiB). | | `secure_boot` | `bool` | `true` | Shielded VM Secure Boot. Turn it off only for images whose kernel modules are unsigned (some GPU drivers). | | `spot` | `bool` | `false` | Run as a Spot VM: cheaper, preemptible at any time; a preempted instance stops and reports `stopped`. Refused for control-plane machines. | The `image` placeholders take `kubernetes_version`: `{version}` gives `v1.33.4`, `{semver}` `1.33.4` and `{slug}` `v1-33-4`, all without an RKE2 `+rke2rN` suffix; `{fullslug}` keeps it, `v1-33-4-rke2r1`. Image names allow no dots or `+`. ## Outputs | Output | Value | | --- | --- | | `provider_id` | `gce:////`, what cloud-provider-gcp writes and parses ([`gce_util.go`]() `providerIDRE`) | | `addresses` | The NIC’s `InternalIP`, and an `ExternalIP` when an access config exists (never, as created): what cloud-provider-gcp reports ([`gce_instances.go`]()) | | `failure_domain` | The instance’s zone: the requested `failure_domain`, or the module’s pick | | `interruptible` | `true` for a Spot VM | | `health` | Below | Extra outputs: `api_instance_group_membership_id`, `instance_id` (the numeric ID, which changes only when the instance is replaced) and `instance_self_link`. ## Health From the instance’s `current_status`, read on every refresh: | Instance status | Contract state | Reason | | --- | --- | --- | | `RUNNING` | `running`, healthy | none | | `PENDING`, `PROVISIONING`, `STAGING` | `pending` | `InstancePending`, `InstanceProvisioning`, `InstanceStaging` | | `REPAIRING` | `degraded` | `InstanceRepairing` | | `PENDING_STOP`, `STOPPING`, `STOPPED`, `SUSPENDING`, `SUSPENDED`, `TERMINATED` | `stopped` | `InstancePendingStop`, `InstanceStopping`, `InstanceStopped`, `InstanceSuspending`, `InstanceSuspended`, `InstanceTerminated` | | `DEPROVISIONING` | `terminated` | `InstanceNotFound` | | any other | `unknown` | `UnknownState` | | The instance no longer exists | `terminated` | `InstanceNotFound` | `TERMINATED` is Compute Engine’s word for stopped: the instance and its disk still exist. A preempted Spot VM is `TERMINATED`, so it reports `stopped` and a MachineHealthCheck can replace it. ## Lifecycle A machine is immutable: CAPTF applies it once. The image is a create-time property: changes to `image` are ignored on an existing instance, so a drift check never reports an image family’s newer image as a replacement. Roll a new image through a new `TerraformMachineTemplate`. The instance group membership is replaced with the instance, since a replaced instance keeps its self link but leaves the group. ## Bootstrap | Machine | Format | Delivery | | --- | --- | --- | | Control plane, `bootstrap_delivery = "secret-manager"` (default) | `cloud-config`, plain or gzipped | Staged in Secret Manager; the instance metadata `user-data` holds a small `#cloud-boothook` script | | Worker, or `bootstrap_delivery = "inline"` | `cloud-config`, plain or gzipped | The base64 data in `user-data` with `user-data-encoding=base64`, which cloud-init’s GCE data source decodes | | Any | Ignition | Decoded into `user-data`, which Ignition reads as it is; gzipped Ignition is refused | The staging script takes the instance’s access token from the metadata server, reads the secret version over the Secret Manager REST API with `curl`, and installs the data as `/etc/cloud/cloud.cfg.d/99-captf-bootstrap.cfg` (see [Shared Behavior]()). It retries for five minutes, which covers the access binding’s IAM propagation. Destroying the machine deletes the secret. Size limits: a secret version holds at most 64 KiB, and an instance metadata value 256 KB. A larger payload fails a precondition; gzip it. Who can read the data: - Staged: the control-plane service account, and anything on the node that can take its token from the metadata server, pods included. Block pod egress to `169.254.169.254` with a NetworkPolicy, except for the system components that need it. - Inline: anyone with `compute.instances.get` in the project, which `roles/viewer` and `roles/compute.viewer` include, and anything on the node that reaches the metadata server. Worker data holds only a short-lived join token. ## Limitations > [!WARNING] > > **Instance names collide across namespaces in a shared project** > > Instance names are unique per project and zone, so two `Machine`s of the same name in different namespaces collide in a shared project. - A failure domain must be one of the cluster’s zones, and a control-plane machine’s zone must have an API instance group. - Control-plane machines refuse `spot`: a preempted control-plane node takes an etcd member with it. - The staged script needs `curl`, `sed`, `base64` and `gzip` in the image; set `bootstrap_delivery = "inline"` for images without them. ## Exceptions - **Ignition control-plane data stays in instance metadata.** Ignition cannot run the staging script’s authenticated fetch, so Ignition data goes inline, readable through `compute.instances.get` and the metadata server. `bootstrap_delivery = "inline"` makes the same trade for `cloud-config`. ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 namespace: team-a spec: template: spec: source: image: ghcr.io/captf-io/gcp-machine:v0.1.0-opentofu variables: image: projects/my-images/global/images/capi-ubuntu-2404-{slug} machine_type: n2-standard-4 ``` # MachinePool The `ghcr.io/captf-io/gcp-machinepool` image implements the [machinepool role]() for a `TerraformMachinePool` on Google Cloud. It creates one regional managed instance group per `MachinePool`, spread over the pool’s zones, with an optional autoscaler. Pool instances are workers. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `google_compute_region_instance_template.pool_instance_template` | Everything an instance is made of except the bootstrap data | Always | | `google_compute_region_instance_group_manager.pool_instance_group_manager` | The group, over the pool’s zones; holds the bootstrap data in its all-instances config | Always | | `google_compute_region_autoscaler.pool_autoscaler` | The autoscaler, between the annotations’ minimum and maximum | While the `MachinePool`’s autoscaler annotations enable autoscaling with a maximum above 0 | | `terraform_data.pool_default_zones` | The cluster’s zones at the first apply, pinned | Always; used when the pool names no failure domains | A data source lists the group’s members, in every state, on each refresh. Instances run as the cluster’s worker service account, with its network tags, no external address, Shielded VM, OS Login, project-wide SSH keys blocked and the serial console off. ## Inputs Contract inputs it uses: - `machinepool_name`: names of the group and its instances. - `replicas`: the group’s target size without autoscaling. - `bootstrap_data`, `bootstrap_format`: the user data. - `failure_domains`, `cluster_failure_domains`: the zones. - `kubernetes_version`: fills the image placeholders. - `node_labels`: registered by every new instance’s kubelet. - `autoscaling`: the autoscaler. - `captf_cluster_outputs`: the cluster’s exports. - `captf_cluster`, `captf_object`, `captf_tags`: names, descriptions and labels. User variables, set with `spec.variables` on the `TerraformMachinePool` ([variables.tf]()): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_network_tags` | `list(string)` | `[]` | Extra network tags for the pool’s instances, on top of the cluster’s node and role tags. | | `additional_tags` | `map(string)` | `{}` | Extra GCP labels for the pool’s instances, their boot disks and the instance template. Keys and values must already be valid GCP labels; the `captf-io_` keys are reserved for `captf_tags`, which win. | | `autoscaling_initialization_seconds` | `number` | `300` | Seconds a new instance needs to boot and join before the autoscaler reads its CPU; 300 covers image boot, cloud-init and kubeadm join. | | `autoscaling_target_cpu_percent` | `number` | `60` | Average CPU utilization, in percent, the autoscaler keeps the pool at while the `MachinePool`’s autoscaler annotations enable autoscaling. 60 leaves headroom for a node to fail. | | `boot_disk_kms_key_id` | `string` | `null` | Cloud KMS key (`projects/.../cryptoKeys/...`) that encrypts the boot disks. Null uses Google-managed encryption; the Compute Engine service agent needs encrypt and decrypt on the key. | | `boot_disk_size_gib` | `number` | `50` | Boot disk size in GiB: room for the image, container images and logs. | | `boot_disk_type` | `string` | `"pd-balanced"` | Boot disk type. `pd-balanced` suits the default N2 machine type; C3, N4 and newer series need `hyperdisk-balanced`. | | `can_ip_forward` | `bool` | `false` | Let the instances send and receive packets for other addresses, as CNIs that route pod CIDRs through GCP routes need. | | `external_cluster_exports` | `any` | `null` | Exports (schema `captf.io/gcp-cluster/v1`) to use when the `TerraformCluster` is externally managed and `captf_cluster_outputs` is `{}`. Ignored otherwise. | | `image` | `string` | `null` | Boot image: a name, `projects/

/global/images/` or a self link, with a version placeholder (below) so a `kubernetes_version` change rolls the pool. Required. | | `machine_type` | `string` | `"n2-standard-4"` | Machine type of the pool’s instances: 4 vCPU and 16 GiB by default. | | `secure_boot` | `bool` | `true` | Shielded VM Secure Boot. Turn it off only for images whose kernel modules are unsigned (some GPU drivers). | | `spot` | `bool` | `false` | Run the pool on Spot VMs: cheaper, preemptible at any time; a preempted instance stops and the group repairs it. | With `kubernetes_version` set, `image` must contain `{version}` (`v1.33.4`), `{semver}` (`1.33.4`), `{slug}` (`v1-33-4`) or `{fullslug}` (`v1-33-4-rke2r1`), and `{fullslug}` when the version has a `+rke2rN` suffix: only `{fullslug}` keeps the suffix, so only it makes a suffix-only upgrade change the image. ## Outputs | Output | Value | | --- | --- | | `provider_id` | The group’s ID, `projects/

/regions//instanceGroupManagers/` | | `provider_id_list` | `gce:////` of every member, in any state, except members being deleted (`DEPROVISIONING`), sorted | | `replicas` | The group’s target size as observed: the autoscaler’s while it scales; 0 once the group is gone | | `instances` | Per member: `provider_id`, `instance_id` (the instance name), `addresses = []`, `failure_domain` (zone), `state` | | `health` | Below | Stopped members (`TERMINATED`) stay in `provider_id_list`: dropping one would make Cluster API delete its Node. The member listing carries no addresses. Extra outputs: `autoscaler_id`, `dropped_node_labels` (keys of `node_labels` the kubelet may not set on itself), `instance_group_manager_id` and `instance_template_id`. ## Health Each member’s status maps as on the [Machine]() page. The pool’s health, in this order: | Situation | Contract state | Reason | | --- | --- | --- | | The group no longer exists | `terminated` | `InstanceGroupNotFound` | | Target size 0, and no autoscaler minimum above 0 still to reach | `running`, healthy | none | | No members yet | `pending` | `NoMembers` | | A member degraded, stopped or unknown | The worst of those three | `:` per affected member, then per starting member, plus `ScalingInProgress` while the count differs | | Otherwise | `running`, healthy only when every member runs and the count equals the target size | The starting members, plus `ScalingInProgress` while the count differs | An autoscaled group is created empty and grows to its minimum, so until a member exists it reports `pending`, not “scaled to zero”. ## Lifecycle | Change | What happens | | --- | --- | | `bootstrap_data` (rotates about every 7.5 minutes with kubeadm) | The all-instances metadata updates in place: a refresh that never restarts or replaces an instance; new instances boot with the current data | | `kubernetes_version` | The placeholder changes the image, so a new instance template is created and the group’s proactive update replaces every instance, one surge instance per zone at a time, none removed first | | `node_labels` | The all-instances metadata updates in place: new instances register the new labels; existing ones keep theirs | | `replicas`, without autoscaling | The group’s target size | | `replicas`, with autoscaling | Ignored: the target size is left to the autoscaler | | Autoscaling on or off | Creates or deletes the autoscaler; a maximum of 0 means no autoscaler and a target size of 0 | | `failure_domains` | Replaces the group, and every instance at once | | The cluster’s zones | Nothing: a pool without its own failure domains keeps the zones of its first apply | | `captf_tags`, `additional_tags` | The all-instances config relabels every instance in place; the template keeps its creation labels | | `machine_type`, `boot_disk_*`, `spot`, `secure_boot`, `additional_network_tags`, `can_ip_forward` | A new instance template; the group updates instances with the least disruptive action Compute Engine allows, which may replace them | Nothing drains an instance first: `MachinePool` Machines are out of scope for the contract. ## Bootstrap The bootstrap data lives in the group’s all-instances config, as the instance metadata key `user-data`. | Format | Labels to register | Sent as | | --- | --- | --- | | `cloud-config` | None | The base64 data, with `user-data-encoding=base64` | | `cloud-config` | Some | A MIME multipart message: a `text/cloud-boothook` part with the shared node-labels script, then the data as a base64 part (`application/x-gzip` when gzipped) | | Ignition | None | The decoded data; gzipped Ignition is refused | | Ignition | Some | Refused: Ignition has no boothook to register them | The node labels script is the one every CAPTF pool shares; see [Shared Behavior](). The user data must fit the 256 KB instance metadata value limit. Pool data holds a join token, not key material: it is readable by anyone with `compute.instances.get` in the project and by anything on the node that reaches the metadata server. ## Limitations > [!WARNING] > > **Changing failure\_domains replaces the group at once** > > Changing `failure_domains` replaces the group at once. A pool without them keeps the cluster’s zones of its first apply, even if the cluster later drops one. - A pool holds at most 500 instances, the one page the member listing returns; a precondition checks `replicas` or the autoscaler’s maximum. - The group keeps zones balanced and may delete an instance to rebalance; nothing drains it. - Ignition pools cannot register `node_labels`. - The template and its boot disks keep the labels of the template’s creation: their labels force replacement, so they are not updated. ## Exceptions - **Any instance template change rolls the pool.** Every template argument forces a new template, so a change of `machine_type`, `boot_disk_*`, `spot`, `secure_boot`, `additional_network_tags` or `can_ip_forward` updates instances with the least disruptive action Compute Engine allows, which may replace them. - **A version change rolls through the image name**, not a version trigger: a template change that only touches metadata is applied as a refresh and rolls nothing. So `image` needs a version placeholder, and `{fullslug}` for versions with a `+rke2rN` suffix. - **The autoscaler’s count is kept without `ignore_changes`** (the allowed `tfcapi-lint` warning `pool/autoscaling-ignore-changes`): the group’s target size is null while an autoscaler runs, which leaves the observed value alone, and `replicas` otherwise. ## Example terraformmachinepool.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: demo-pool-0 namespace: team-a labels: cluster.x-k8s.io/cluster-name: demo spec: source: image: ghcr.io/captf-io/gcp-machinepool:v0.1.0-opentofu variables: image: projects/my-images/global/images/capi-ubuntu-2404-{slug} ``` # Azure The Azure modules run a Kubernetes cluster on Azure virtual machines. The cluster role creates one resource group per cluster, holding the API server’s Standard load balancer, the nodes’ network security groups, application security groups and managed identities with their role assignments. The machine role creates one Linux VM per Machine, and the machinepool role one virtual machine scale set per MachinePool. The source is [`captf-io/azure-modules`](). The modules are pre-release; read the status note in [Cloud Modules]() first. - **Cluster** --- The API endpoint, security rules, node identities and exports of one workload cluster. - **Machine** --- One instance per `Machine`, registered with the API load balancer. - **MachinePool** --- One native scaling group per `MachinePool`. ## Images | Role | Image | Page | | --- | --- | --- | | cluster | `ghcr.io/captf-io/azure-cluster` | [Cluster]() | | machine | `ghcr.io/captf-io/azure-machine` | [Machine]() | | machinepool | `ghcr.io/captf-io/azure-machinepool` | [MachinePool]() | The images carry `hashicorp/azurerm` 5.7.0, pinned exactly. ## Prerequisites - **Network.** A virtual network with a subnet for the nodes, and optionally a second subnet in the same network for workers, in the subscription the identity works in. Nodes get no public IP, so the subnets need an egress path, such as a NAT gateway or a firewall, to the image registries and Azure’s endpoints. With the default internal API endpoint, the management cluster must reach the control-plane subnet (through peering or a VPN, for example). - **Resource providers** registered in the subscription: `Microsoft.Compute`, `Microsoft.Network`, `Microsoft.ManagedIdentity`, `Microsoft.Authorization` and `Microsoft.Insights` (autoscale settings). The modules never register them. - **Permissions** of the identity’s service principal: - Contributor on the subscription: the cluster role creates a resource group, and joins the subnets, which must be in the same subscription. Narrow Contributor, and the network’s resource group needs Network Contributor. - Role Based Access Control Administrator on the subscription, best with a condition that limits it to assigning Contributor, Network Contributor and AcrPull: the cluster role grants the node identities. - Managed Identity Operator on any node identity you bring, so the machine and pool roles can attach it. - **Node images.** A Linux image built for Cluster API, with kubeadm, the kubelet, a container runtime, cloud-init, `curl` and `iptables`. The [CAPZ reference images]() work: `/communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver}`. - **In the workload cluster:** [cloud-provider-azure]() (the cloud controller manager and cloud-node-manager), with `--configure-cloud-routes=false` for an overlay CNI, and the [Azure Disk CSI driver]() for persistent volumes. Both read `/etc/kubernetes/azure.json`, which the modules write on every node from cloud-config bootstrap data. ## Identity Secret The Secret holds the azurerm provider’s environment variables: | Key | Value | | --- | --- | | `ARM_TENANT_ID` | The Microsoft Entra tenant of the service principal | | `ARM_SUBSCRIPTION_ID` | The subscription the cluster and its network live in | | `ARM_CLIENT_ID` | The service principal’s application (client) ID | | `ARM_CLIENT_SECRET` | Its client secret | | `ARM_USE_CLI` | `"false"`: the images have no Azure CLI | For a client certificate, set `ARM_CLIENT_CERTIFICATE_PATH` to a file under `/var/run/captf/credentials/` and `ARM_CLIENT_CERTIFICATE_PASSWORD` instead of `ARM_CLIENT_SECRET`. From [`examples/identity.yaml`](): identity.yaml ```yaml apiVersion: v1 kind: Secret metadata: name: azure namespace: captf-system type: Opaque stringData: ARM_TENANT_ID: 00000000-0000-0000-0000-000000000000 ARM_SUBSCRIPTION_ID: 00000000-0000-0000-0000-000000000000 ARM_CLIENT_ID: 00000000-0000-0000-0000-000000000000 ARM_CLIENT_SECRET: replace-me ARM_USE_CLI: "false" --- apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: azure spec: secretRef: name: azure namespace: captf-system allowedNamespaces: list: - team-a ``` See [Identities and Credentials]() for how the Secret reaches the Jobs. ## Quick start 1. Create the identity from the repository’s example: ```sh export NAMESPACE=team-a AZURE_TENANT_ID=... AZURE_SUBSCRIPTION_ID=... \ AZURE_CLIENT_ID=... AZURE_CLIENT_SECRET=... clusterctl generate yaml --from examples/identity.yaml | kubectl apply -f - ``` 2. Generate the cluster from [`examples/cluster-kubeadm.yaml`](): a KubeadmControlPlane of three, a MachineDeployment of two and a MachinePool of two. Quote the image ID so the shell keeps the braces: ```sh export CLUSTER_NAME=demo KUBERNETES_VERSION=v1.34.1 export AZURE_SUBNET_ID=/subscriptions//resourceGroups/network/providers/Microsoft.Network/virtualNetworks/hub/subnets/nodes export AZURE_SSH_PUBLIC_KEY="$(cat ~/.ssh/id_ed25519.pub)" export AZURE_IMAGE_ID='/communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver}' clusterctl generate yaml --from examples/cluster-kubeadm.yaml | kubectl apply -n "$NAMESPACE" -f - ``` 3. Once the first control-plane node answers, install a CNI (for example Calico with VXLAN) and cloud-provider-azure in the workload cluster. cloud-provider-azure sets each Node’s provider ID; only then do the Machines get their Nodes and the workers join. The KubeadmConfigs in the example name each Node after its VM (`nodeRegistration.name: '{{ ds.meta_data["local_hostname"] }}'`), which cloud-provider-azure needs to find the VM. ## API endpoint - **Internal** (the default): a Standard load balancer frontend in the control-plane subnet, zone-redundant where the region has zones, with a dynamic address unless `api_load_balancer_private_ip` sets one. - **Public** (`api_load_balancer_public = true`): a static Standard public IP. Include the management cluster’s egress and the nodes’ NAT gateway addresses in `api_allowed_cidrs`: the load balancer keeps the client’s source address, so the control plane’s security group sees them. - **Probes.** HTTPS `GET /readyz` every 5 seconds with kubeadm; TCP with `distribution = "rke2"`, which may disable anonymous access, and on the supervisor port 9345. - **Hairpin.** An Azure internal load balancer drops a flow from a backend to its own frontend when it maps the flow back to that backend. The machine role therefore installs a small service on control-plane nodes (cloud-config bootstrap data only): while the node’s own API server answers `/readyz`, an iptables rule sends the node’s traffic for the frontend to it directly; otherwise the rule is gone and the load balancer takes the traffic to another control-plane node, since this node’s probe is down too. With RKE2 it covers port 9345 as well. Turn it off with `api_server_hairpin_workaround = false`; with Ignition, add the equivalent to your bootstrap configuration. - **The endpoint guard** records `api_load_balancer_public`, the endpoint port, the control-plane subnet and the private address the frontend got, and fails any later plan that changes one (see [Shared Behavior]()). Making the current dynamic address static is allowed, since it moves nothing. ## Exports The cluster role’s `exports` carry the schema `captf.io/azure-cluster/v1`. | Key | Value | | --- | --- | | `schema` | `captf.io/azure-cluster/v1` | | `tenant_id` | The identity’s tenant | | `subscription_id` | The subscription, lowercase | | `region` | The virtual network’s Azure location, such as `westeurope` | | `resource_group_name`, `resource_group_id` | The cluster’s resource group (name lowercase) | | `failure_domains` | One entry per zone, `{ "1" = {}, "2" = {}, "3" = {} }`; `{}` without zones | | `virtual_network` | `{id, name, resource_group_name}` of the brought network | | `subnet_id`, `subnet_name` | The control-plane subnet | | `worker_subnet_id`, `worker_subnet_name` | The worker subnet (the control-plane subnet without `worker_subnet_id`) | | `admin_username`, `admin_ssh_public_key` | The nodes’ admin user (`captf`) and key | | `control_plane` | `{identity_id, identity_client_id, network_security_group_id, network_security_group_name, application_security_group_id, availability_set_id}`; the availability set only in a region without zones | | `worker` | The same keys without `availability_set_id` | | `api` | `{host, port, backend_port, frontend_ip, backend_pool_id, hairpin_workaround, supervisor_port}`; `port` is the endpoint port, `backend_port` the kube-apiserver’s; `null` with a user endpoint | | `cloud_provider_config` | The cloud-provider-azure configuration (`/etc/kubernetes/azure.json`) without `userAssignedIdentityID`, which each node adds for its identity | Nothing in the exports is secret: the cloud provider authenticates with the nodes’ managed identities. ## Tags Azure tag names cannot contain `/`, so `captf.io/` becomes `captf.io_` (see [Shared Behavior]()). Tag names are case-insensitive, so `additional_tags` rejects any key that starts with `captf.io_` or `captf.io/` in any case, and takes at most 44 tags: Azure allows 50 per resource. Values longer than 256 characters keep 247, then `-` and 8 hex characters of their sha256. Not taggable: role assignments, security rules, load balancer backend pools, probes and rules, NIC associations, the OS disk Azure creates with a VM, and the instances, NICs and OS disks Azure creates for a scale set. All of them live in the cluster’s resource group, which is tagged. > [!NOTE] > > **Design notes** > > - **One resource group per cluster, always created.** It scopes the node identities’ Contributor rights to what the cluster owns, collects what cloud-provider-azure creates for Services, and makes destroy fail loudly while those are still there instead of deleting disks with data. > - **Network security groups on the NICs, not the subnet.** The subnet is yours; application security groups name the nodes, so intra-cluster traffic is open whatever the addresses. Workers keep Azure’s default allow-from-network rule, because cloud-provider-azure adds none for internal Service load balancers. > - **The module writes `/etc/kubernetes/azure.json`.** Without it the cloud controller manager never starts and no Machine gets its Node; the derived names it needs are not something you can easily put in a KubeadmConfig. It is written only when absent, so a file of your own wins. > - **Custom data, never user data.** Azure exposes user data to every process on the node through the instance metadata service; custom data it does not. > - **Scale sets update their model in place.** The provider’s `roll_instances_when_required` and `reimage_on_manual_upgrade` are off: with azurerm’s defaults every bootstrap token rotation would reimage every instance. A version change therefore creates a new scale set. > - **Azure Autoscale holds a pool’s capacity in both modes,** because the scale set ignores changes to its instance count, which an autoscaled pool must. > - **Trusted launch is off by default:** the CAPZ reference images do not support it. The evidence for each is in [DESIGN.md](). > [!WARNING] > > **Not yet verified** > > - Whether Azure reports zones on a frontend or public IP differently from what was sent (harmless: the modules ignore later zone changes). > - The minimal role set for cloud-provider-azure and Azure Disk CSI. > - Azure Autoscale holding minimum, maximum and default equal with no rules, including 0, and how long it takes to apply a change. > - The custom data limit of 65,535 decoded bytes (87,380 base64 characters). > - Hairpin through a public load balancer; whether RKE2 needs the hairpin workaround; the DNAT rule alongside every CNI’s iptables rules. > - cloud-init running CABPK’s Jinja template from a `text/plain` part of a multipart message, on the CAPZ images’ cloud-init version. > - Azure returning a role assignment’s subnet or registry scope with the casing it was created with. > - cloud-provider-azure with `vmType: "vmss"` handling standalone VMs next to scale sets. > - The CAPZ gallery publishing an image for every Kubernetes version you roll to. > - A cluster destroy after its network was deleted. > - RKE2’s API server answering `/readyz` with 401 or 403 when anonymous access is off, and whether the 9345 DNAT is needed. > - Azure Resource Manager read throttling of pool refreshes beyond about 200 instances. # Cluster The `ghcr.io/captf-io/azure-cluster` image implements the [cluster role]() on Azure. It creates one resource group for the cluster and, in it, the API server load balancer, the nodes’ network security groups and application security groups, and the nodes’ managed identities with their role assignments. It publishes the endpoint, one failure domain per availability zone, and the [exports]() the machine and pool roles read. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `azurerm_resource_group.cluster_resource_group` | The cluster’s own group: scope of the node identities’ rights, home of what cloud-provider-azure creates for Services | Always | | `azurerm_user_assigned_identity.node_identities` | Managed identities of control-plane and worker nodes | Per role, unless you bring one (`control_plane_identity_id`, `worker_identity_id`) | | `azurerm_role_assignment.node_role_assignments` | Contributor on the group and Network Contributor on the node subnets for the control-plane identity; AcrPull on each of `container_registry_ids` | For the identities the module creates | | `azurerm_network_security_group.node_security_groups` | One per role, attached to each node’s NIC | Always | | `azurerm_network_security_rule.node_security_rules` | SSH denied unless from `ssh_allowed_cidrs`; all traffic between nodes; the API server ports and a final deny of the virtual network on the control plane | Always; the SSH allow only with `ssh_allowed_cidrs`, the CIDR allow only with `api_allowed_cidrs` | | `azurerm_application_security_group.node_application_security_groups` | Name control-plane and worker NICs in the rules | Always | | `azurerm_availability_set.control_plane_availability_set` | Spreads control-plane VMs over fault domains | In a region without availability zones | | `azurerm_public_ip.api_public_ip` | Static Standard public IP of the endpoint | With `api_load_balancer_public` | | `azurerm_lb.api_load_balancer` | Standard load balancer of the endpoint | Without a user endpoint | | `azurerm_lb_backend_address_pool.api_backend_pool` | Control-plane machines join it from their own state | Without a user endpoint | | `azurerm_lb_probe.api_probes` | HTTPS `/readyz` (kubeadm) or TCP probe, every 5 seconds | One per listener, without a user endpoint | | `azurerm_lb_rule.api_rules` | The endpoint port to the kube-apiserver, and 9345 to 9345 with RKE2 | One per listener, without a user endpoint | | `terraform_data.api_endpoint_guard` | Records what fixes the endpoint and fails a later plan that would move it | Without a user endpoint | It reads the identity’s tenant and subscription, and finds the brought virtual network and any brought identities through subscription-wide listings that come back empty instead of failing, so a network deleted first never blocks a destroy ([Shared Behavior]()). Every other plan stops at a precondition naming what is missing. ## Inputs Contract inputs used: `captf_cluster` (names), `captf_tags` (tags), `control_plane_endpoint` (a user endpoint skips the load balancer) and `cluster_network.api_server_port` (the endpoint port, default 6443; the kube-apiserver backend port too with kubeadm, while RKE2’s is always 6443). `captf_contract` is validated; `captf_object`, `captf_cluster_outputs`, `kubernetes_version` and `control_plane_initialized` are not used. User variables, set with `spec.variables` on the `TerraformCluster` ([variables.tf]()): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_tags` | `map(string)` | `{}` | Extra Azure tags on every taggable resource. Keys starting with `captf.io_` or `captf.io/` (any case) are rejected; at most 44 | | `admin_ssh_public_key` | `string` | `null` | Required. OpenSSH public key (`ssh-rsa` or `ssh-ed25519`) of every node’s admin user: Azure Linux VMs need a key or a password, and the module invents neither. SSH stays closed unless `ssh_allowed_cidrs` opens it | | `api_allowed_cidrs` | `list(string)` | `[]` | IPv4 CIDRs, besides the virtual network, allowed to reach the API server ports. Required with `api_load_balancer_public`: include the management cluster’s egress and the nodes’ NAT gateway addresses | | `api_load_balancer_private_ip` | `string` | `null` | Static private IPv4 address of the internal frontend, in the control-plane subnet. `null` takes a dynamic address, stable for the load balancer’s life | | `api_load_balancer_public` | `bool` | `false` | Put the API load balancer on a public IP | | `api_server_hairpin_workaround` | `bool` | `true` | Control-plane nodes send their own traffic for the frontend to their local API server while it is ready ([API endpoint]()) | | `container_registry_ids` | `list(string)` | `[]` | Azure Container Registry IDs the identities this module creates may pull from (AcrPull) | | `control_plane_identity_id` | `string` | `null` | Existing user-assigned identity for control-plane nodes. It needs Contributor on the cluster’s resource group and Network Contributor on the node subnets, which you grant | | `distribution` | `string` | `"kubeadm"` | `kubeadm` or `rke2`. RKE2 adds the supervisor port 9345 to the load balancer and the control-plane security group, probes the API server with TCP, and fixes the kube-apiserver backend port at 6443 | | `resource_group_name` | `string` | `null` | Name of the cluster’s resource group, lowercase, because cloud-provider-azure lowercases it in provider IDs. `null` derives `captf---` | | `ssh_allowed_cidrs` | `list(string)` | `[]` | IPv4 CIDRs allowed to reach SSH on every node, for example an Azure Bastion subnet. Empty: SSH is denied, between nodes too | | `subnet_id` | `string` | `null` | Required. Existing subnet for control-plane nodes and the internal frontend, and for workers without `worker_subnet_id` | | `worker_identity_id` | `string` | `null` | Existing user-assigned identity for worker nodes | | `worker_subnet_id` | `string` | `null` | Existing subnet for worker nodes, in the same virtual network as `subnet_id`. `null` uses `subnet_id` | | `zones` | `list(string)` | `[]` | Availability zones to publish as failure domains. Empty uses every zone of the region | Subnet, identity and registry IDs are matched case-insensitively and rebuilt with Azure’s canonical segment names. Give resource group and resource names in the casing Azure shows (`az network vnet subnet show --query id`): role assignment scopes are compared case-sensitively. ## Outputs | Output | Value | | --- | --- | | `control_plane_endpoint` | The user endpoint, or the frontend address (private, or the public IP) and the endpoint port | | `failure_domains` | One entry per zone of the region (or of `zones`), eligible for the control plane; `[]` in a region without zones | | `exports` | See [Exports]() | | `health` | See Health | | `api_load_balancer_id` | Extra: the load balancer’s ARM ID; `null` with a user endpoint | | `resource_group_id` | Extra: the resource group’s ARM ID | ## Health Cluster health comes from the cluster’s own resources, never from the control-plane nodes behind the load balancer. | Azure state | Contract state | Reason | | --- | --- | --- | | Load balancer and frontend address present, or the resource group with a user endpoint | `running`, healthy | none | | Load balancer or its frontend address gone | `terminated` | `LoadBalancerNotFound` | | Resource group gone | `terminated` | `ResourceGroupNotFound` | ## Limitations > [!WARNING] > > **Destroy fails while cloud-provider-azure resources remain** > > Destroy fails while cloud-provider-azure’s load balancers, public IPs or disks are still in the resource group: delete the cluster’s `LoadBalancer` Services and `PersistentVolumes` first, or remove what is left by hand. - The network, its egress, DNS and peering are yours. - Azure public cloud only. - The endpoint is an IP address; there is no DNS name. - With `distribution = "rke2"` the endpoint port cannot be 9345, the supervisor’s. ## Exceptions None: `tfcapi-lint module --strict` passes without allowed warnings. The tests cannot cover the `terminated` readings, because a mock provider never drops a resource on refresh. ## Example terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo namespace: team-a spec: source: image: ghcr.io/captf-io/azure-cluster:v0.1.0-opentofu identityRef: name: azure defaults: identityRef: name: azure variables: subnet_id: /subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/network/providers/Microsoft.Network/virtualNetworks/hub/subnets/nodes admin_ssh_public_key: ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIExample ops@example.com ``` # Machine The `ghcr.io/captf-io/azure-machine` image implements the [machine role]() on Azure: one Linux VM per `Machine`, control plane or worker, in the cluster’s resource group. `control_plane` picks the subnet, identity and security groups, and registers a control-plane VM with the API load balancer before it boots. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `azurerm_network_interface.node_network_interface` | The VM’s NIC in the control-plane or worker subnet, with accelerated networking by default | Always | | `azurerm_network_interface_security_group_association.node_security_group_association` | Puts the NIC under its role’s network security group | Always | | `azurerm_network_interface_application_security_group_association.node_application_security_group_association` | Makes the NIC a member of its role’s application security group | Always | | `azurerm_network_interface_backend_address_pool_association.api_backend_pool_association` | Registers the NIC in the API backend pool, in this machine’s own state, so destroy deregisters it | Control-plane machines, unless the endpoint is your own | | `azurerm_linux_virtual_machine.node_virtual_machine` | The node, with password login and VM extensions off and a user-assigned identity | Always | The VM is created after the associations, so it boots in its security groups and, on the control plane, behind the load balancer. Without exports (an externally managed cluster and no `external_cluster_exports`) the role creates nothing and fails a precondition. ## Inputs Contract inputs used: `captf_cluster_outputs` (the cluster’s exports), `captf_tags`, `machine_name`, `bootstrap_data`, `bootstrap_format`, `failure_domain`, `kubernetes_version` (fills `{version}` and `{semver}` in `image_id`) and `control_plane`. `captf_contract` is validated; `captf_cluster` and `captf_object` are not used. User variables, set with `spec.template.spec.variables` on the `TerraformMachineTemplate` ([variables.tf]()): | Name | Type | Default | Description | | --- | --- | --- | --- | | `accelerated_networking` | `bool` | `true` | Accelerated networking on the NIC; turn it off for a `vm_size` without it | | `additional_tags` | `map(string)` | `{}` | Extra Azure tags on the VM and NIC. Keys starting with `captf.io_` or `captf.io/` (any case) are rejected; at most 44 | | `boot_diagnostics` | `bool` | `true` | Keep the serial console log in Azure-managed storage. It shows boot output, which may include kubeadm’s join command | | `encryption_at_host` | `bool` | `false` | Encrypt temporary disks and caches on the host too; needs the `EncryptionAtHost` feature on the subscription. Managed disks are encrypted at rest either way | | `external_cluster_exports` | `any` | `null` | Exports (schema `captf.io/azure-cluster/v1`) for an externally managed `TerraformCluster` | | `image_id` | `string` | `null` | Required. A managed image, Compute Gallery image (version), or community or shared gallery image (version) ID. `{version}` and `{semver}` become the Machine’s version, `v1.31.4` and `1.31.4`, without any `+suffix` | | `ip_forwarding` | `bool` | `false` | IP forwarding on the NIC, for CNIs that route pod addresses natively | | `os_disk_size_gib` | `number` | `128` | OS disk size, 30 to 4095 GiB | | `os_disk_storage_account_type` | `string` | `"Premium_LRS"` | `Standard_LRS`, `StandardSSD_LRS`, `StandardSSD_ZRS`, `Premium_LRS` or `Premium_ZRS`; premium needs a `vm_size` with premium storage | | `spot` | `bool` | `false` | A Spot VM, deallocated on eviction, at most the on-demand price, reported interruptible. A control-plane machine with `spot` fails a precondition | | `trusted_launch` | `bool` | `false` | Secure boot and vTPM; needs a generation 2 image built for trusted launch, which the CAPZ reference images are not | | `vm_size` | `string` | `"Standard_D4s_v5"` | Azure VM size. The image’s capacity labels describe the default: 4 CPUs, 16 GiB, amd64 | Community and shared gallery IDs are case-sensitive in azurerm 5.7.0: `/communityGalleries//images//versions/`. ## Outputs | Output | Value | | --- | --- | | `provider_id` | `azure:///subscriptions//resourceGroups//providers/Microsoft.Compute/virtualMachines/`, the format [cloud-provider-azure]() writes to the Node | | `addresses` | `InternalIP`, the NIC’s private address, and `Hostname`, the VM name | | `failure_domain` | The VM’s zone: the requested failure domain, else one picked from the `sha256` of `machine_name`; `null` without zones | | `interruptible` | `true` for a Spot VM | | `health` | See Health | | `network_interface_id` | Extra: the NIC’s ARM ID | | `virtual_machine_id` | Extra: the VM’s ARM ID | cloud-provider-azure finds the VM by the Node’s name, so the Node must be named after the VM: set `nodeRegistration.name: '{{ ds.meta_data["local_hostname"] }}'` in the KubeadmConfig, as the examples do. The VM and its hostname are `machine_name` when Azure accepts it (at most 64 lowercase letters, digits and inner hyphens); any other name is made valid, cut to 55 characters and suffixed with `-` and 8 hex characters of its `sha256`. In a region without zones, control-plane VMs join the cluster’s availability set. ## Health The module lists the VM first and reads its power state only when the listing finds it, because Azure’s VM read fails on a missing VM, which would fail every refresh and destroy. | Azure state | Contract state | Reason | | --- | --- | --- | | Power state `running` | `running`, healthy | none | | `starting`, or no power state yet | `pending` | `PowerState/starting`, `PowerState/unknown` | | `stopping`, `stopped`, `deallocating`, `deallocated` | `stopped` | `PowerState/` | | Any other power state | `unknown` | `UnknownState` | | In the state but not listed yet (the first apply; listings lag) | `pending` | `VirtualMachineNotListed` | | Gone (a refresh dropped it) | `terminated` | `VirtualMachineNotFound` | The first apply therefore reports `pending`; the controller’s next refresh, 30 seconds later, reads the power state. An evicted Spot VM is deallocated, reports `stopped`, and a `MachineHealthCheck` replaces it. ## Bootstrap - **Delivery.** The bootstrap data goes in the VM’s custom data, never in user data, which Azure shows to every process on the node through the instance metadata service. `cloud-config` data is wrapped in a MIME multipart message: a boot hook first, then the payload as an opaque base64 part, never decoded by the module (`application/x-gzip` when it is gzipped). The boot hook writes `/etc/kubernetes/azure.json` for the node’s identity when the file is absent and, on control-plane nodes, starts the [hairpin workaround](). Ignition goes in unchanged; gzipped Ignition fails a precondition. - **Size limit.** Azure takes 65,535 bytes of custom data, 87,380 base64 characters; a precondition stops a larger message. Compress the payload (CAPRKE2 `gzipUserData`) if a control-plane payload grows too big. - **Who can read it.** Azure keeps custom data out of the instance metadata service and does not return it when the VM is read; on the node it sits in files only root reads. Like every input, it is also stored in the machine’s state Secret on the management cluster. With `boot_diagnostics` on, the serial console log, readable by anyone who may read the VM’s boot diagnostics, shows boot output. ## Limitations - One NIC, one OS disk and no data disk: etcd shares the OS disk. - No public IP; nodes reach the internet through the subnet’s egress. - Azure public cloud only. - An image for `{semver}` must exist in the gallery for every version you roll to. ## Exceptions None: `tfcapi-lint module --strict` passes without allowed warnings. One trivy finding is ignored with its reason: AZU-0068 on the NIC, whose security group is attached by a separate association resource. The tests cannot cover the `terminated` reading, because a mock provider never drops a resource on refresh. ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 namespace: team-a spec: template: spec: source: image: ghcr.io/captf-io/azure-machine:v0.1.0-opentofu variables: image_id: /communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver} vm_size: Standard_D8s_v5 spot: true ``` # MachinePool The `ghcr.io/captf-io/azure-machinepool` image implements the [machinepool role]() on Azure: one uniform Linux virtual machine scale set of worker nodes per `MachinePool`, in the cluster’s resource group, with an Azure Autoscale setting that holds its capacity. See [Machine Pools]() for the objects to create. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `azurerm_linux_virtual_machine_scale_set.pool_scale_set` | The workers: manual upgrades, no overprovisioning, no single placement group, in the worker subnet, security groups and identity, with termination notifications | Always | | `azurerm_monitor_autoscale_setting.pool_autoscale_setting` | Holds the scale set’s capacity: pinned to `replicas`, or between the autoscaling bounds with CPU rules | Always | | `terraform_data.pool_default_zones` | The cluster’s zones as of the first apply, kept for a pool without its own failure domains | Always (used without `failure_domains`) | Without exports (an externally managed cluster and no `external_cluster_exports`) the scale set and autoscale setting are not created and a precondition fails. ## Inputs Contract inputs used: `captf_cluster_outputs` (the cluster’s exports), `captf_tags`, `machinepool_name`, `replicas`, `bootstrap_data`, `bootstrap_format`, `failure_domains`, `cluster_failure_domains`, `kubernetes_version`, `node_labels` and `autoscaling`. `captf_contract` is validated; `captf_cluster` and `captf_object` are not used. User variables, set with `spec.variables` on the `TerraformMachinePool` ([variables.tf]()): | Name | Type | Default | Description | | --- | --- | --- | --- | | `accelerated_networking` | `bool` | `true` | Accelerated networking on the instances’ NICs | | `additional_tags` | `map(string)` | `{}` | Extra Azure tags on the scale set and autoscale setting. Keys starting with `captf.io_` or `captf.io/` (any case) are rejected; at most 44 | | `autoscaling_scale_in_cpu_percent` | `number` | `25` | With autoscaling, scale in by one instance below this average CPU over 10 minutes; must be below the scale-out threshold | | `autoscaling_scale_out_cpu_percent` | `number` | `75` | With autoscaling, scale out by one instance above this average CPU over 10 minutes | | `boot_diagnostics` | `bool` | `true` | Keep the serial console logs in Azure-managed storage | | `encryption_at_host` | `bool` | `false` | Encrypt temporary disks and caches on the host too; needs the `EncryptionAtHost` feature | | `external_cluster_exports` | `any` | `null` | Exports (schema `captf.io/azure-cluster/v1`) for an externally managed `TerraformCluster` | | `image_id` | `string` | `null` | Required. A managed image, Compute Gallery image (version), or community or shared gallery image (version) ID. `{version}` and `{semver}` become the pool’s version, `v1.31.4` and `1.31.4`, without any `+suffix` | | `ip_forwarding` | `bool` | `false` | IP forwarding on the NICs, for CNIs that route pod addresses natively | | `os_disk_size_gib` | `number` | `128` | OS disk size per instance, 30 to 4095 GiB | | `os_disk_storage_account_type` | `string` | `"Premium_LRS"` | `Standard_LRS`, `StandardSSD_LRS`, `StandardSSD_ZRS`, `Premium_LRS` or `Premium_ZRS` | | `spot` | `bool` | `false` | Spot instances, deleted on eviction, at most the on-demand price | | `trusted_launch` | `bool` | `false` | Secure boot and vTPM; needs a generation 2 image built for trusted launch | | `vm_size` | `string` | `"Standard_D4s_v5"` | Azure VM size of the instances | ## Outputs | Output | Value | | --- | --- | | `provider_id` | `azure:///subscriptions//resourceGroups//providers/Microsoft.Compute/virtualMachineScaleSets/`; changes when a new generation replaces the scale set | | `provider_id_list` | Every instance the scale set lists, whatever its power state, as `/virtualMachines/`, the format cloud-provider-azure writes to the Node | | `replicas` | The scale set’s capacity as last refreshed, which Azure Autoscale sets; `null` once the scale set is gone | | `instances` | Per instance: `provider_id`, `instance_id`, `addresses` (`InternalIP`, `Hostname`), `failure_domain` (its zone) and `state` | | `health` | See Health | | `autoscale_setting_id` | Extra: the autoscale setting’s ARM ID | | `dropped_node_labels` | Extra: `node_labels` keys left out because the kubelet may not set them on itself | | `scale_set_id`, `scale_set_name` | Extra: the current scale set’s ARM ID and name | The scale set is named `-<8 hex characters>`, and its instances’ hostnames start with that name; cloud-provider-azure finds an instance by its hostname, so name the Nodes after it (`nodeRegistration.name: '{{ ds.meta_data["local_hostname"] }}'`). ## Health The module lists the scale set first and reads its instances only when the listing finds it. Each instance’s power state maps like a [machine’s](); the pool follows the [shared order](): | Azure state | Contract state | Reason | | --- | --- | --- | | Scale set gone (a refresh dropped it) | `terminated` | `ScaleSetNotFound` | | Capacity 0 | `running`, healthy | none | | Not listed yet (first apply, new generation) or no instances | `pending` | `NoMembers`; `provider_id_list` is `[]` until the next refresh | | Some instance stopped, deallocated or in an unknown state | the worst of `degraded`, `stopped`, `unknown` | `:` per instance | | Instances starting, or the count off the capacity | `running`, not healthy | `PowerState/starting:` per instance, `ScalingInProgress` | | Every instance running at capacity | `running`, healthy | none | Azure lists an instance until it is deleted, so no instance maps to `terminated`: a deleted instance leaves `provider_id_list`. ## Lifecycle | Change | Effect | | --- | --- | | `bootstrap_data` (a token rotation, about every 7.5 minutes), `node_labels`, `image_id`, `vm_size`, `os_disk_size_gib`, `accelerated_networking`, `ip_forwarding`, `boot_diagnostics`, `encryption_at_host`, the exports, tags | The scale set’s model updates in place; new instances use it, running ones are left alone | | `replicas`, autoscaling disabled | The autoscale setting pins the new capacity; Azure Autoscale applies it within about a minute | | `autoscaling`, `autoscaling_scale_*_cpu_percent` | The autoscale setting gets the new bounds and rules | | `kubernetes_version`, compared verbatim (a `+rke2rN` bump included) | A new scale set at the current capacity, then the old one is deleted | | `failure_domains`, `spot`, `trusted_launch`, `os_disk_storage_account_type` | A new scale set, as for a version change: Azure cannot change these on a scale set | | `cluster_failure_domains` | Nothing: a pool without its own `failure_domains` keeps the cluster’s zones as of its first apply | The provider turns off `roll_instances_when_required` and `reimage_on_manual_upgrade`: with azurerm’s defaults every token rotation would reimage every instance. So nothing rolls by itself, and a version change creates a new scale set with `create_before_destroy`; the pool briefly needs twice its quota. There is no drain: pools have no Machines. Scheduled Events announce every deletion 5 minutes ahead, for a termination handler that drains. ## Bootstrap - **Delivery.** As for the [machine role](): custom data, `cloud-config` in a MIME multipart message after a boot hook, and Ignition unchanged (gzipped Ignition fails a precondition). The boot hook writes `/etc/kubernetes/azure.json` for the worker identity when the file is absent and renders the node labels as the [shared behavior]() describes; Ignition with labels left to render fails a precondition. - **Size limit.** 65,535 bytes of custom data, 87,380 base64 characters; a precondition stops a larger message. - **Who can read it.** As for the machine role: not through the instance metadata service or a read of the scale set; on the node only root; and the pool’s state Secret on the management cluster. ## Limitations > [!WARNING] > > **Keep pools to about 200 instances** > > A scale set holds at most 1,000 instances, but every refresh reads each instance’s NICs with one Azure call; keep pools to about 200 instances, or raise `spec.membershipRefreshIntervalSeconds`. - Workers only: the control plane uses the [machine role](). - Changes reach new instances only; a version change is the way to roll the pool. - Capacity changes go through Azure Autoscale and take about a minute to reach the scale set; `replicas` follows at the next refresh. - The scale set read fails if an instance disappears between listing it and reading its NICs; the next refresh succeeds. - Azure public cloud only. ## Exceptions - **`pool/autoscaling-ignore-changes`**, a `tfcapi-lint` warning, is allowed: the scale set ignores changes to `instances`, its desired count, which the check’s pattern does not know, and Azure Autoscale holds the capacity in both modes. - **Version rolls by generation.** The roll is a new scale set name with `create_before_destroy`, not `terraform_data.kubernetes_version_roll`: a replacement under the same name would collide with the old scale set. - **Other replacements.** `failure_domains`, `spot`, `trusted_launch` and `os_disk_storage_account_type` replace the scale set, because azurerm 5.7.0 cannot update them in place. - **`replicas` is `null` once the scale set is gone:** the capacity is then unknown, and an empty `provider_id_list` with an unknown capacity keeps Cluster API’s guard against deleting every Node engaged. - The `membership_excludes_terminated` test asserts that stopped and deallocated instances stay members, since Azure has no terminated instance state; the `ScaleSetNotFound` reading has no test. ## Example terraformmachinepool.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: demo-pool-0 namespace: team-a labels: cluster.x-k8s.io/cluster-name: demo spec: source: image: ghcr.io/captf-io/azure-machinepool:v0.1.0-opentofu variables: image_id: /communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver} autoscaling_scale_out_cpu_percent: 70 ``` # OCI The OCI modules build a self-managed Kubernetes cluster on Oracle Cloud Infrastructure, on a VCN you bring. The cluster role creates the nodes’ network security groups, a network load balancer for the API, and a dynamic group and policy so the control-plane nodes run the OCI cloud controller manager and CSI controller as instance principals. The machine role creates one compute instance per `Machine`; the machinepool role runs an instance pool at a fixed size or under OCI autoscaling. The code is in [`captf-io/oci-modules`](); read the status note in [Cloud Modules]() before you rely on it. - **Cluster** --- The API endpoint, security rules, node identities and exports of one workload cluster. - **Machine** --- One instance per `Machine`, registered with the API load balancer. - **MachinePool** --- One native scaling group per `MachinePool`. ## Images | Role | Image | Page | | --- | --- | --- | | cluster | `ghcr.io/captf-io/oci-cluster` | [Cluster]() | | machine | `ghcr.io/captf-io/oci-machine` | [Machine]() | | machinepool | `ghcr.io/captf-io/oci-machinepool` | [MachinePool]() | The images pin the `oracle/oci` provider at 9.8.0. ## Prerequisites - **A VCN** with a regional subnet for the control plane, and optionally one for workers and one for the API load balancer, all in one VCN and in the compartment you pass as `network_compartment_id` (default: the cluster’s `compartment_id`). Private nodes need a NAT gateway route for egress, and a service gateway or NAT route to the OCI APIs. A public API load balancer needs a public subnet. Subnets need no DNS label. Security lists stay yours: the modules put their rules in network security groups. - **A defined-tag namespace and key** for `node_identity.defined_tag`. The cluster’s node dynamic group matches only control-plane instances by this tag; without one the cluster role refuses to plan unless you opt in to a compartment-wide group or turn node identity off. Keep one compartment per cluster either way. - **Permissions** for the identity’s user (or instance principal): ```text Allow group to use virtual-network-family in compartment Allow group to manage network-security-groups in compartment Allow group to manage network-load-balancers in compartment Allow group to inspect compartments in tenancy Allow group to manage instance-family in compartment Allow group to use volume-family in compartment Allow group to manage compute-management-family in compartment Allow group to manage auto-scaling-configurations in compartment Allow group to read metrics in compartment # Node identity (node_identity.enabled, the default): Allow group to manage dynamic-groups in tenancy Allow group to manage policies in compartment Allow group to inspect tenancies in tenancy Allow group to use tag-namespaces in tenancy # boot_volume_kms_key_id: Allow service blockstorage to use keys in compartment ``` - **A node image** in the cluster’s region with containerd, the kubelet and kubeadm (or RKE2) of your Kubernetes version and cloud-init, for example from [image-builder](). The modules switch instance metadata v1 off, so cloud-init must read metadata v2, as current releases do. Autoscaled pools also need the Oracle Cloud Agent’s Compute Instance Monitoring plugin. - **In the workload cluster**, once its API server answers: the [OCI cloud controller manager]() with `useInstancePrincipals: true`, the cluster’s compartment and VCN, and `securityListManagementMode: None`, scheduled on control-plane nodes; a CNI; and the OCI CSI driver if you want block volumes. ## Identity Secret The OCI provider reads each setting from `OCI_` environment variables. CAPTF drops `TF_VAR_*`, so the Secret uses the `OCI_*` names, and the API signing key travels as a file key that `OCI_PRIVATE_KEY_PATH` points to. Config-file profiles do not work: the Job has no `~/.oci/config`. See [Identities and Credentials](). | Key | Value | | --- | --- | | `OCI_TENANCY_OCID` | The tenancy OCID; the cluster role also reads it for the dynamic group | | `OCI_USER_OCID` | The OCID of the user the API key belongs to | | `OCI_FINGERPRINT` | The API key’s fingerprint | | `OCI_PRIVATE_KEY_PATH` | `/var/run/captf/credentials/oci_api_key.pem` | | `oci_api_key.pem` | The PEM private key, unencrypted (or add `OCI_PRIVATE_KEY_PASSWORD`) | | `OCI_AUTH` | Optional: `InstancePrincipal` on a management cluster that runs on OCI, instead of the key | identity.yaml ```yaml apiVersion: v1 kind: Secret metadata: name: oci namespace: captf-system type: Opaque stringData: OCI_TENANCY_OCID: ocid1.tenancy.oc1..replace-me OCI_USER_OCID: ocid1.user.oc1..replace-me OCI_FINGERPRINT: "00:11:22:33:44:55:66:77:88:99:aa:bb:cc:dd:ee:ff" OCI_PRIVATE_KEY_PATH: /var/run/captf/credentials/oci_api_key.pem oci_api_key.pem: | -----BEGIN PRIVATE KEY----- replace-me -----END PRIVATE KEY----- ``` The region is never in the identity: each `TerraformCluster` sets `region`. ## Quick start 1. Apply the identity and its Secret from [`examples/identity.yaml`](), then replace the Secret’s placeholders with your key: ```sh export NAMESPACE=team-a clusterctl generate yaml --from examples/identity.yaml | kubectl apply -f - kubectl create secret generic oci -n captf-system --dry-run=client -o yaml \ --from-literal=OCI_TENANCY_OCID=ocid1.tenancy.oc1.. \ --from-literal=OCI_USER_OCID=ocid1.user.oc1.. \ --from-literal=OCI_FINGERPRINT= \ --from-literal=OCI_PRIVATE_KEY_PATH=/var/run/captf/credentials/oci_api_key.pem \ --from-file=oci_api_key.pem= | kubectl apply -f - ``` 2. Generate the cluster from [`examples/cluster-kubeadm.yaml`](): a three-node `KubeadmControlPlane`, a `MachineDeployment`, a `MachinePool` and MachineHealthChecks. ```sh export CLUSTER_NAME=demo KUBERNETES_VERSION=v1.34.1 export OCI_COMPARTMENT_ID=ocid1.compartment.oc1.. OCI_REGION=us-ashburn-1 export OCI_CONTROL_PLANE_SUBNET_ID=ocid1.subnet.oc1.iad. export OCI_WORKER_SUBNET_ID=ocid1.subnet.oc1.iad. export OCI_IMAGE_ID=ocid1.image.oc1.iad. export OCI_NODE_TAG_NAMESPACE= OCI_NODE_TAG_KEY= clusterctl generate yaml --from examples/cluster-kubeadm.yaml | kubectl apply -n "$NAMESPACE" -f - ``` 3. Wait for the `TerraformCluster` to become ready. KCP then creates the control plane, and each control-plane instance joins the load balancer as it boots. 4. Install the cloud controller manager and a CNI in the workload cluster, as in [Prerequisites](<#prerequisites>). Every node registers with `provider-id: oci://{{ v1.instance_id }}` in the example’s `kubeletExtraArgs`: cloud-init’s Oracle datasource sets `instance_id` to the instance OCID, so the Node’s provider ID matches the modules’ `provider_id` from its first registration. ## API endpoint - **Internal by default.** The network load balancer gets a private address in its subnet, reachable from the VCN. The management cluster must reach the VCN (peering, a DRG or VPN), or set `api_load_balancer_public = true` with `api_allowed_cidrs` and a public subnet. Nodes in private subnets reach a public endpoint through the NAT gateway, so its public IP belongs in `api_allowed_cidrs`. - **Who may connect.** The load balancer’s NSG admits the API ports from the VCN’s CIDRs and `api_allowed_cidrs`; the control-plane NSG admits them only from the load balancer’s NSG. - **Hairpin.** The backend sets do not preserve the client address (`is_preserve_source = false`), so a backend sees the load balancer as the source and a control-plane node reaches itself through the endpoint. [Cluster API Provider OCI]() configures its API load balancer the same way. - **Fixed addresses.** `api_load_balancer_private_ip` and `api_load_balancer_reserved_public_ip_id` pin the address. - **The endpoint guard** records `api_load_balancer_public`, the load balancer subnet, `api_load_balancer_private_ip`, `api_load_balancer_reserved_public_ip_id`, `cluster_network.api_server_port` and the address the load balancer got. See [Shared Behavior](). ## Exports Schema `captf.io/oci-cluster/v1`: | Key | Value | | --- | --- | | `schema` | `captf.io/oci-cluster/v1` | | `region` | The cluster’s region | | `compartment_id` | The compartment of every cluster resource | | `vcn_id` | The control-plane subnet’s VCN | | `control_plane_subnet_id`, `worker_subnet_id` | The node subnets | | `control_plane_nsg_id`, `worker_nsg_id` | The node network security groups | | `failure_domains` | Failure domain name to `{ availability_domain, fault_domain }`; `fault_domain` only in fault-domain mode | | `node_defined_tags` | The defined tag workers carry; `{}` without `node_identity.defined_tag` | | `control_plane_defined_tags` | The defined tag control-plane machines carry, which the node dynamic group matches | | `api` | `{ host, port, network_load_balancer_id, kube_apiserver, rke2_supervisor }`, each listener `{ backend_set_name, port }`; `null` with a user-supplied endpoint | ## Tags Every taggable resource carries `captf_tags` as OCI free-form tags. Free-form keys may not contain periods or spaces, so `.` and space become `_` (`captf_io/cluster`); values are unchanged. OCI allows 10 free-form tags per resource and the six captf tags always apply, so `additional_tags` takes at most four, none starting with `captf_io/`. The provider ignores the defined tags `Oracle-Tags.CreatedBy` and `Oracle-Tags.CreatedOn`; list your own tag defaults in `ignore_defined_tags`. Not taggable on OCI: network security group rules, network load balancer backend sets, listeners and backends. > [!NOTE] > > **Design notes** > > - **A network load balancer for the API**, because it forwards TCP without terminating TLS and, with source preservation off, supports hairpin. > - **Node identity for control-plane nodes only.** `read instance-family` lets a principal read any instance’s user data, which on a control-plane instance holds the cluster’s CA keys; workers get no OCI permissions, and the CSI node driver needs none. > - **The tag namespace is yours.** A per-cluster namespace created by the module would make destroy slow (OCI retires and deletes namespaces asynchronously) and block a new cluster of the same name. > - **Failure domains** are availability domains where the region has several, otherwise the three fault domains of its one availability domain, as CAPOCI does. > - **Two pool resources**, fixed and autoscaled, because the provider overwrites `size` from the cloud on every read and OCI pools have no minimum or maximum to pin. A mode switch replaces the pool, so it needs the `autoscaled` variable to agree. > - **Listings, not reads,** for the subnets and VCN, so a destroy finishes after the network is gone. > - **User data inline.** OCI has no store the modules stage bootstrap data in yet, so it goes in instance metadata; the role pages say who can read it. [DESIGN.md]() has the evidence for each decision. > [!WARNING] > > **Not yet verified** > > - Whether OCI accepts the empty free-form tag value of `captf.io/template`. > - IAM writes through the home-region provider; identity domains other than Default; a `/` in a defined-tag value in a matching rule. > - The minimal node policy for load balancer Services and CSI, and the minimal permissions listed above. > - Autoscaling down to 0 instances (refused until verified), and what the autoscaling configuration’s `initial` size does to an existing pool. > - The casing of pool members’ states, and whether terminating members are listed. > - Deleting an instance configuration that running instances came from. > - An empty second plan for shape, source details, defined tags and the autoscaling rules. > - Backend registration time against a fast `kubeadm init`. > - In-transit boot volume encryption on custom images. > - Whether the cloud controller manager and CSI controller need more than the five policy statements. # Cluster The `ghcr.io/captf-io/oci-cluster` image implements the [cluster role]() for a `TerraformCluster` on OCI. It creates the nodes’ network security groups, the API network load balancer and the control-plane nodes’ instance identity on the VCN you bring, and publishes the endpoint, the failure domains and the ids the machine and machinepool roles need. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `oci_core_network_security_group.control_plane_nsg` and its rules | Control-plane nodes: everything from both node NSGs, the API backend ports from the load balancer NSG, ICMP 3/4 (path MTU) from the VCN, SSH from `ssh_allowed_cidrs` | Always | | `oci_core_network_security_group.worker_nsg` and its rules | Workers: everything from both node NSGs, ICMP 3/4 from the VCN, NodePorts 30000-32767 from `nodeport_allowed_cidrs`, SSH from `ssh_allowed_cidrs` | Always | | `oci_network_load_balancer_network_load_balancer.api_load_balancer` | The API endpoint, private unless `api_load_balancer_public` | No `control_plane_endpoint` input | | `oci_core_network_security_group.api_load_balancer_nsg` and its rules | The API ports from the VCN and `api_allowed_cidrs`, forwarded to the control-plane NSG | Same | | `oci_network_load_balancer_backend_set.api_backend_sets`, `oci_network_load_balancer_listener.api_listeners` | `kube-apiserver`, and `rke2-supervisor` on 9345 with RKE2; TCP health checks | Same | | `terraform_data.api_endpoint_guard` | Fails any plan that would move the endpoint | Same | | `oci_identity_dynamic_group.node_dynamic_group` | The control-plane instances, by defined tag, in the tenancy | `node_identity.enabled` (default) | | `oci_identity_policy.node_policy` | What the cloud controller manager and CSI controller need | Same | It lists the subnets and VCNs of the network compartment and picks the cluster’s by OCID, reads the region’s availability domains (and, in fault-domain mode, the fault domains), and reads the tenancy’s region subscriptions to find the home region, where IAM writes go. Every subnet must be found, `AVAILABLE`, in one VCN, and regional or in the control-plane subnet’s availability domain; these checks are preconditions, so a destroy skips them. The node policy grants the dynamic group, in the cluster’s compartment, `read instance-family`, `manage load-balancers`, `manage network-load-balancers` and `manage volume-family`, plus `use virtual-network-family` in each network compartment. A statement about the root compartment reads `in tenancy`. ## Inputs Contract inputs used: `captf_cluster` (names), `captf_tags`, `control_plane_endpoint` (skips the load balancer) and `cluster_network` (`api_server_port`). `captf_contract` is validated; `captf_object`, `kubernetes_version`, `control_plane_initialized` and `captf_cluster_outputs` are declared and unused. User variables, from [variables.tf](): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_tags` | `map(string)` | `{}` | Extra free-form tags for every taggable resource; at most 4. | | `api_allowed_cidrs` | `list(string)` | `[]` | CIDRs that may reach the API endpoint besides the VCN’s own; required with a public load balancer. | | `api_load_balancer_private_ip` | `string` | `null` | Private IPv4 address of the load balancer, from its subnet; `null` lets OCI pick. | | `api_load_balancer_public` | `bool` | `false` | Give the load balancer a public IP; requires `api_allowed_cidrs`. | | `api_load_balancer_reserved_public_ip_id` | `string` | `null` | OCID of a reserved public IP for a public load balancer. | | `api_load_balancer_subnet_id` | `string` | `null` | Subnet of the load balancer; `null` uses `control_plane_subnet_id`. A public load balancer needs a public subnet. | | `compartment_id` | `string` | `null` | Required. Compartment of every cluster resource, machines and pools included. | | `control_plane_subnet_id` | `string` | `null` | Required. Subnet of the control-plane nodes; its VCN is the cluster’s. | | `distribution` | `string` | `"kubeadm"` | `kubeadm` or `rke2`: RKE2 adds the supervisor listener on 9345 and puts the kube-apiserver backends on 6443. | | `failure_domain_mode` | `string` | `"auto"` | `availability_domain`, `fault_domain` (of one availability domain), or `auto`. | | `home_region` | `string` | `null` | The tenancy’s home region for IAM writes; `null` looks it up. | | `ignore_defined_tags` | `list(string)` | `[]` | Tag-default keys (`.`) to leave alone, besides `Oracle-Tags.CreatedBy` and `Oracle-Tags.CreatedOn`. | | `network_compartment_id` | `string` | `null` | Compartment of the VCN and its subnets; `null` uses `compartment_id`. | | `node_identity` | `object({enabled = optional(bool, true), defined_tag = optional(object({namespace = string, key = string})), allow_compartment_wide = optional(bool, false)})` | `{}` | Instance-principal identity for the control-plane nodes. `defined_tag` scopes the dynamic group to this cluster’s control-plane instances; it is required unless `enabled = false` or `allow_compartment_wide = true`. | | `node_policy_compartment_id` | `string` | `null` | Where the node policy is attached; `null` uses `compartment_id`. Set an ancestor of both when the VCN is in another compartment. | | `nodeport_allowed_cidrs` | `list(string)` | `[]` | CIDRs that may reach worker NodePorts (TCP and UDP). | | `region` | `string` | `null` | Required. Region identifier, such as `us-ashburn-1`; machines and pools inherit it. | | `ssh_allowed_cidrs` | `list(string)` | `[]` | CIDRs that may reach the nodes on SSH. | | `tenancy_id` | `string` | `null` | Tenancy of the dynamic group; `null` reads the identity’s `OCI_TENANCY_OCID` file. | | `worker_subnet_id` | `string` | `null` | Subnet of the workers, in the same VCN; `null` uses `control_plane_subnet_id`. | ## Outputs | Output | Value | | --- | --- | | `control_plane_endpoint` | The `control_plane_endpoint` input when given; otherwise the load balancer’s IPv4 address (public when `api_load_balancer_public`, else private) and the API port | | `failure_domains` | One entry per availability domain, or per fault domain of the one availability domain; all accept control-plane machines | | `exports` | Schema `captf.io/oci-cluster/v1`; see [Exports]() | | `health` | From the load balancer’s lifecycle state (below) | | `api_load_balancer_id` | Extra: OCID of the network load balancer; `null` with a user endpoint | | `node_dynamic_group_id` | Extra: OCID of the dynamic group, for your own policy statements | | `node_policy_id` | Extra: OCID of the node policy | An availability domain’s failure domain name is the part after its tenancy prefix (`Uocm:US-ASHBURN-AD-1` becomes `US-ASHBURN-AD-1`), the zone the OCI cloud controller manager reports. Fault domains keep OCI’s names (`FAULT-DOMAIN-1`). An AD-specific control-plane subnet confines the cluster to the fault domains of its availability domain. ## Health | Load balancer state | Contract state | Reason | | --- | --- | --- | | `ACTIVE`, `UPDATING` (each backend change updates it) | `running` | none | | `CREATING` | `pending` | `LoadBalancerCreating` | | `FAILED` | `degraded` | `LoadBalancerFailed` | | `DELETING`, `DELETED`, or gone from state | `terminated` | `LoadBalancerNotFound` | | anything else | `unknown` | `UnknownState` | With a user-supplied endpoint there is no load balancer to observe, and the cluster reports `running`. ## Limitations > [!CAUTION] > > **Control-plane metadata holds the cluster’s CA keys** > > The policy lets control-plane instances read any instance’s metadata in the compartment, so control-plane nodes of another cluster there can read this cluster’s keys. Keep one compartment per cluster, and block pod access to the metadata service on control-plane nodes. - **The endpoint is fixed** once the load balancer exists: the endpoint guard refuses any change to the inputs behind it (see [API endpoint]()). - **The policy covers the compartment.** Control-plane nodes can manage every load balancer and volume in it. Workers get no OCI permissions. - **The cloud controller manager needs no `manage security-lists`** only with `securityListManagementMode: None`; open Service NodePorts with `nodeport_allowed_cidrs`. - **IAM changes propagate within minutes**, so the first nodes may retry their first OCI calls. - **The dynamic group goes in the Default identity domain**, through the classic IAM API. ## Exceptions - Two rules are open by default besides node-to-node traffic and the API port from the load balancer: ICMP type 3 code 4 from the VCN, so path MTU discovery works, and, with a user-supplied endpoint, the API backend ports from the VCN and `api_allowed_cidrs`, since that endpoint reaches the nodes directly. - With the network gone, the network security groups carry a placeholder VCN id, because OpenTofu’s destroy still evaluates required arguments. Any other plan stops at the subnet checks first. ## Example terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo spec: source: image: ghcr.io/captf-io/oci-cluster:v0.1.0-opentofu identityRef: name: oci defaults: identityRef: name: oci variables: compartment_id: ocid1.compartment.oc1.. region: us-ashburn-1 control_plane_subnet_id: ocid1.subnet.oc1.iad. worker_subnet_id: ocid1.subnet.oc1.iad. node_identity: defined_tag: namespace: captf key: cluster ``` # Machine The `ghcr.io/captf-io/oci-machine` image implements the [machine role]() for a `TerraformMachine` on OCI: one compute instance per `Machine`, placed in the Machine’s failure domain on the cluster’s subnet and network security group and, for a control-plane machine, registered in the API load balancer. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `oci_core_instance.node_instance` | The node and its boot volume | Always | | `oci_network_load_balancer_backend.api_backends` | The instance in each API backend set (`kube-apiserver`, and `rke2-supervisor` with RKE2) | `control_plane` and a module-owned endpoint | The instance is named `machine_name`, the first key the OCI cloud controller manager looks a node up by. It runs `image_id` on `VM.Standard.E5.Flex` with 2 OCPUs and 16 GB by default, on a 100 GiB boot volume that is encrypted at rest (with `boot_volume_kms_key_id` when set) and in transit, and that goes with the instance when it terminates. It has no public IP and no SSH key unless you ask, and instance metadata v1 is off. A control-plane machine carries the cluster’s control-plane defined tag, which the node dynamic group matches; a worker carries the worker tag. ## Inputs Contract inputs used: `captf_cluster_outputs` (or `external_cluster_exports`), `captf_tags`, `machine_name`, `bootstrap_data`, `bootstrap_format`, `failure_domain` and `control_plane`. `captf_contract` is validated; `captf_cluster`, `captf_object` and `kubernetes_version` are declared and unused: the image carries the version. User variables, from [variables.tf](): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_nsg_ids` | `list(string)` | `[]` | Extra network security groups for the VNIC, after the cluster’s; at most 4. | | `additional_tags` | `map(string)` | `{}` | Extra free-form tags for the instance and its VNIC; at most 4. | | `boot_volume_kms_key_id` | `string` | `null` | Vault key for the boot volume; `null` uses Oracle-managed keys. | | `boot_volume_size_gib` | `number` | `100` | Boot volume size, 50 to 32768. | | `external_cluster_exports` | `any` | `null` | The cluster’s exports when the `TerraformCluster` is externally managed. | | `ignore_defined_tags` | `list(string)` | `[]` | Tag-default keys (`.`) to leave alone. | | `image_id` | `string` | `null` | Required. Node image OCID in the cluster’s region. | | `memory_gib` | `number` | `16` | Memory of a flexible shape; ignored for fixed shapes. | | `ocpus` | `number` | `2` | OCPUs of a flexible shape (one OCPU is two vCPUs on x86); ignored for fixed shapes. | | `preemptible` | `bool` | `false` | Preemptible capacity; workers only. | | `public_ip` | `bool` | `false` | A public IP on the VNIC; needs a public subnet. | | `pv_encryption_in_transit` | `bool` | `true` | Encrypt boot volume traffic in transit. | | `shape` | `string` | `"VM.Standard.E5.Flex"` | Compute shape. | | `ssh_authorized_keys` | `list(string)` | `[]` | SSH public keys for the image’s default user. | | `subnet_id` | `string` | `null` | Subnet of the VNIC; `null` uses the cluster’s control-plane or worker subnet. | The image’s `io.captf.capacity` label (`{"cpu":"4","memory":"16Gi"}`, `amd64`) describes the default shape, for Cluster Autoscaler scale from zero. A template that changes the shape needs an image built with matching labels. ## Outputs | Output | Value | | --- | --- | | `provider_id` | `oci://`, the format the OCI cloud controller manager writes (`ProviderName() + "://" + InstanceID`, [ccm.go]()) | | `addresses` | `InternalIP` (the primary private IP) and `ExternalIP` (the public IP, when there is one), as the cloud controller manager reports them | | `failure_domain` | The requested failure domain, or the one picked from `machine_name` | | `interruptible` | `true` for a preemptible instance | | `health` | From the instance’s lifecycle state (below) | A requested failure domain must be one of the cluster’s; the instance goes to its availability domain and, in fault-domain mode, its fault domain. Without one, the module picks from the sorted names by the sha256 of `machine_name`, so a new plan never moves the machine. ## Health | Instance state | Contract state | Reason | | --- | --- | --- | | `PROVISIONING`, `STARTING` | `pending` | `InstanceProvisioning`, `InstanceStarting` | | `RUNNING`, `MOVING` (live migration) | `running` | none | | `CREATING_IMAGE` | `unknown` | `InstanceCreatingImage` | | `STOPPING`, `STOPPED` | `stopped` | `InstanceStopping`, `InstanceStopped` | | `TERMINATING`, `TERMINATED`, or gone from state | `terminated` | `InstanceNotFound` | | anything else | `unknown` | `UnknownState` | The provider drops a `TERMINATED` instance from state when it reads it; the instance is counted, so the module still reads that as `terminated`. ## Lifecycle A machine never updates in place. A change to its bootstrap data or SSH keys replaces the instance (the provider forces it), which is what an immutable `Machine` expects; CAPI rolls machines through a new template instead. Destroying a control-plane machine removes its load balancer backends with it. ## Bootstrap - **Delivery.** `metadata.user_data` is `bootstrap_data` unchanged, since OCI takes user data base64-encoded: cloud-config (CABPK’s Jinja header included), Ignition and gzipped cloud-config alike. Gzipped Ignition is refused. - **Size.** Instance metadata, user data and SSH keys together, may hold 32,000 bytes; a precondition stops a larger payload. - **Who can read it.** The instance itself through its metadata service, and any principal that may read instances in the compartment through the OCI API. The cluster’s node policy grants that to control-plane instances only, which hold the cluster’s keys anyway. A control-plane payload holds the cluster’s CA keys; OCI has no store the module stages it in yet. ## Limitations - **Preemptible capacity** is refused for control-plane machines. - **Backend registration** follows the instance once it is `RUNNING`, a network load balancer work request that typically takes a minute or two; the backend takes traffic once its TCP health check passes. That is inside the time `kubeadm init` and `join` take, but not yet verified. ## Exceptions - `display_name` is `machine_name` rather than a hashed name: the cloud controller manager matches nodes by it. - The control-plane payload is not staged in a secret store; it goes in instance metadata, as [Bootstrap](<#bootstrap>) describes. OCI Vault secrets read by the instance principal could serve as the store. ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 spec: template: spec: source: image: ghcr.io/captf-io/oci-machine:v0.1.0-opentofu variables: image_id: ocid1.image.oc1.iad. boot_volume_size_gib: 200 ``` Set `provider-id: oci://{{ v1.instance_id }}` and `cloud-provider: external` in the bootstrap configuration’s `kubeletExtraArgs`, as the repository’s [example]() does. # MachinePool The `ghcr.io/captf-io/oci-machinepool` image implements the [machinepool role]() for a `TerraformMachinePool` on OCI: one instance pool, launched from an instance configuration and spread over the `MachinePool`’s failure domains, at a fixed size or sized by OCI autoscaling. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `oci_core_instance_configuration.pool_instance_configuration` | What every pool instance launches with | Always | | `oci_core_instance_pool.fixed_instance_pool` | The pool, at `replicas` | `autoscaling.enabled` false | | `oci_core_instance_pool.autoscaled_instance_pool` | The pool, sized by the autoscaler | `autoscaling.enabled` | | `oci_autoscaling_auto_scaling_configuration.pool_autoscaling_configuration` | CPU threshold scaling within the annotations’ bounds | `autoscaling.enabled` | | `terraform_data.kubernetes_version_roll` | Replaces the pool when the Kubernetes version changes | Always | The instances look like the [machine role]()’s workers, on the worker subnet and network security group. The module lists the pool’s members on every refresh. ## Inputs Contract inputs used: `captf_cluster` and `machinepool_name` (names), `captf_cluster_outputs` (or `external_cluster_exports`), `captf_tags`, `replicas`, `bootstrap_data`, `bootstrap_format`, `failure_domains`, `cluster_failure_domains`, `kubernetes_version`, `node_labels` and `autoscaling`. `captf_contract` is validated; `captf_object` is declared and unused. User variables, from [variables.tf](): | Name | Type | Default | Description | | --- | --- | --- | --- | | `additional_nsg_ids` | `list(string)` | `[]` | Extra network security groups for the VNICs, after the cluster’s worker NSG; at most 4. | | `additional_tags` | `map(string)` | `{}` | Extra free-form tags for every pool resource; at most 4. | | `autoscaled` | `bool` | `false` | Must equal `autoscaling.enabled`: the deliberate second switch for a mode change, which replaces the pool. | | `autoscaling_cool_down_seconds` | `number` | `300` | Minimum time between two scaling actions; 300 is OCI’s minimum. | | `autoscaling_scale_in_cpu_percent` | `number` | `30` | Remove an instance below this CPU utilization. | | `autoscaling_scale_out_cpu_percent` | `number` | `70` | Add an instance above this CPU utilization. | | `boot_volume_kms_key_id` | `string` | `null` | Vault key for the boot volumes; `null` uses Oracle-managed keys. | | `boot_volume_size_gib` | `number` | `100` | Boot volume size, 50 to 32768. | | `external_cluster_exports` | `any` | `null` | The cluster’s exports when the `TerraformCluster` is externally managed. | | `ignore_defined_tags` | `list(string)` | `[]` | Tag-default keys (`.`) to leave alone. | | `image_id` | `string` | `null` | Required. Node image OCID in the cluster’s region; a new image reaches new instances only. | | `memory_gib` | `number` | `16` | Memory of a flexible shape; ignored for fixed shapes. | | `ocpus` | `number` | `2` | OCPUs of a flexible shape; ignored for fixed shapes. | | `preemptible` | `bool` | `false` | Preemptible capacity. | | `public_ip` | `bool` | `false` | A public IP on every instance; needs a public subnet. | | `pv_encryption_in_transit` | `bool` | `true` | Encrypt boot volume traffic in transit. | | `shape` | `string` | `"VM.Standard.E5.Flex"` | Compute shape. | | `ssh_authorized_keys` | `list(string)` | `[]` | SSH public keys for the image’s default user. | | `subnet_id` | `string` | `null` | Subnet of the instances; `null` uses the cluster’s worker subnet. | ## Outputs | Output | Value | | --- | --- | | `provider_id` | OCID of the instance pool | | `provider_id_list` | `oci://` of every member, whatever its health, sorted; an instance `TERMINATING` or `TERMINATED` is not a member | | `replicas` | The pool’s size as OCI reports it (`actual_size`): `replicas`, or what the autoscaler decided; not the running count | | `instances` | Per member: `provider_id`, `instance_id`, `failure_domain` (mapped back from its availability and fault domain), `state`, and `addresses = []` | | `health` | Below | | `dropped_node_labels` | Extra: `node_labels` keys left out because NodeRestriction forbids a kubelet to set them | ## Health The pool follows the order in [Shared Behavior](). On OCI: | Situation | Contract state | Reason | | --- | --- | --- | | Pool gone from state, `TERMINATING` or `TERMINATED` | `terminated` | `PoolNotFound` | | Member `PROVISIONING`, `STARTING` | `running` while others run; `pending` only with no members | `InstanceProvisioning:`, `InstanceStarting:`; `NoMembers` | | Member `CREATING_IMAGE` | `unknown` | `InstanceCreatingImage:` | | Member `STOPPING`, `STOPPED` | `stopped` | `InstanceStopping:`, `InstanceStopped:` | | Member in any other state | `unknown` | `UnknownState:` | | Member count below the desired size | `running` | `ScalingInProgress` | Members `RUNNING` or `MOVING` are running. ## Lifecycle | Change | Effect | | --- | --- | | `bootstrap_data` (a rotation, about every 7.5 minutes with kubeadm), `node_labels`, the image or another instance setting | A new instance configuration, created before the old one goes; the pool switches to it in place. Running instances stay. | | `kubernetes_version`, compared verbatim (a `+rke2rN` bump included) | The pool is replaced, new before old: the new pool comes up at the current size, then the old one and its instances go. `provider_id` changes. | | `replicas`, autoscaling off | The pool resizes in place. | | `replicas`, autoscaling on | Nothing: the autoscaler owns the size. | | `autoscaling.min` or `max` | The autoscaling configuration is replaced; the pool stays. | | `autoscaling.enabled` on or off | Refused until `autoscaled` matches; then the pool is replaced, every instance at once. | | `failure_domains` | The pool’s placement updates in place; existing instances stay where they are. | ## Bootstrap - **Delivery.** A cloud-config payload becomes the second part of a MIME multipart whose first part is a `text/cloud-boothook` script with the shared node-labels fragment (see [Shared Behavior]()). The payload is an opaque base64 part, `text/plain` so cloud-init detects cloud-config and CABPK’s Jinja header, or `application/x-gzip` for a gzipped payload. An Ignition payload passes through unchanged; with node labels, or gzipped, it is refused. - **Size.** Instance metadata may hold 32,000 bytes, the boothook part included; a precondition stops a larger payload. - **Who can read it.** As for the [machine role](): the instance itself, and principals that may read instances in the compartment. Pool instances are workers, whose join token is the only secret in the payload. ## Limitations > [!WARNING] > > **Scale-in, a version roll or a mode switch does not drain nodes** > > A scale-in, a version roll or a mode switch terminates instances without draining them: MachinePool Machines are out of scope. OCI’s pre-termination lifecycle action only delays termination, so it is not configured. - **No addresses** in `instances`: the pool’s member list carries none. - **Autoscaling down to 0** is refused (`min` must be at least 1) until it is verified. Autoscaling needs the Compute Instance Monitoring plugin on the image. - **Existing members keep** their labels and bootstrap settings; only new members get a new instance configuration. ## Exceptions - `tfcapi-lint` warns `pool/autoscaling-ignore-changes`, allowed in the repository’s Makefile: the check knows desired-count attributes by name, and an OCI pool’s is `size`, which it does not list. The autoscaled pool does ignore `size`. - Switching autoscaling on or off replaces the pool and all its instances: OCI pools have no minimum or maximum to pin, so one pool resource cannot serve both modes. The `autoscaled` variable makes the switch deliberate. ## Example terraformmachinepool.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: demo-pool-0 labels: cluster.x-k8s.io/cluster-name: demo spec: source: image: ghcr.io/captf-io/oci-machinepool:v0.1.0-opentofu variables: image_id: ocid1.image.oc1.iad. ``` For an autoscaled pool, add `autoscaled: true` to the variables and both autoscaler annotations to the `MachinePool`: ```yaml metadata: annotations: cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "2" cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "10" ``` # OpenStack The OpenStack modules provision a cluster on a network and subnet you bring. The cluster role creates the node and control-plane security groups, an Octavia load balancer for the Kubernetes API (internal by default) and a Nova server group that spreads the control plane; the machine role creates one Nova server per `Machine` on a Neutron port of that subnet and adds control-plane servers to the API pools before they boot. The code lives in [`captf-io/openstack-modules`](). The modules are pre-release; see the status note in [Cloud Modules](). OpenStack has no machinepool role: it has no native scaling group, so use a `MachineDeployment` of individual machines. - **Cluster** --- The API endpoint, security rules, node identities and exports of one workload cluster. - **Machine** --- One instance per `Machine`, registered with the API load balancer. ## Images | Role | Image | Page | | --- | --- | --- | | cluster | `ghcr.io/captf-io/openstack-cluster` | [Cluster]() | | machine | `ghcr.io/captf-io/openstack-machine` | [Machine]() | The machine image carries no `io.captf.capacity` or `io.captf.node-info` label: `flavor_name` has no default, so there is no default node size to describe. For Cluster Autoscaler scale-from-zero, set the `capacity.cluster-autoscaler.kubernetes.io/*` annotations on the `MachineDeployment` instead. ## Prerequisites - **Network.** An existing Neutron network and subnet, IPv4 or IPv6, with DHCP and a router, routed to wherever the nodes pull images from. The management cluster must reach the subnet: the API endpoint is an internal VIP on it unless you make it public. A public endpoint also needs an external network to take the floating IP from, and a gateway on the subnet’s router. - **Services.** Nova, Neutron with security groups and resource tags, Glance, and Octavia with the amphora provider. Octavia API 2.12 or later for `api_allowed_cidrs`. - **Permissions.** An application credential of a user with the `member` role in the project. Clouds still on Octavia’s legacy policy also require `load-balancer_member`. The credential needs no unrestricted flag: the modules create no credentials. - **Node images.** A Glance image with cloud-init (or Ignition) and the kubelet, kubeadm or RKE2 and container runtime your bootstrap provider expects, for example one built with [image-builder](). The image must take its hostname from the metadata service or config drive, so that the Node name equals the server name. - **Cloud controller manager.** Install the [OpenStack cloud controller manager]() in the workload cluster, with a `cloud.conf` and a second, dedicated application credential you supply: the modules create no node identity. The repository’s [`examples/cloud-controller-manager.yaml`]() delivers both with a ClusterResourceSet. Keep its `manage-security-groups` off: it would edit the node ports’ security groups behind the machine role. For volumes, install [Cinder CSI]() with the same `cloud.conf`. ## Identity Secret The provider reads its credentials from a `clouds.yaml` file in the identity Secret; the provider block sets nothing but the region. See [Identities and Credentials]() for how the keys reach a Job. | Key | Example | Purpose | | --- | --- | --- | | `OS_CLOUD` | `openstack` | The `clouds.yaml` entry to use | | `OS_CLIENT_CONFIG_FILE` | `/var/run/captf/credentials/clouds.yaml` | Where the provider finds `clouds.yaml` | | `clouds.yaml` | an application credential entry | Mounted as a file at that path | | `cacert.pem` | PEM bundle | Optional: a private CA, named by `cacert` in `clouds.yaml` | The plain `OS_AUTH_URL`, `OS_APPLICATION_CREDENTIAL_ID`, `OS_APPLICATION_CREDENTIAL_SECRET`, `OS_REGION_NAME` and `OS_CACERT` keys work too. identity.yaml ```yaml apiVersion: v1 kind: Secret metadata: name: openstack namespace: captf-system type: Opaque stringData: OS_CLOUD: openstack OS_CLIENT_CONFIG_FILE: /var/run/captf/credentials/clouds.yaml clouds.yaml: | clouds: openstack: auth_type: v3applicationcredential auth: auth_url: https://keystone.example.com:5000/v3 application_credential_id: application_credential_secret: region_name: RegionOne interface: public identity_api_version: 3 --- apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: openstack spec: secretRef: name: openstack namespace: captf-system allowedNamespaces: list: - team-a ``` ## Quick start The repository’s [`examples/`]() holds clusterctl templates for each step. 1. Create the identity from `examples/identity.yaml`: ```sh export NAMESPACE=team-a OPENSTACK_AUTH_URL=https://keystone.example.com:5000/v3 \ OPENSTACK_APPLICATION_CREDENTIAL_ID= OPENSTACK_APPLICATION_CREDENTIAL_SECRET= clusterctl generate yaml --from examples/identity.yaml | kubectl apply -f - ``` 2. Put the cloud controller manager’s manifests in a ConfigMap and apply `examples/cloud-controller-manager.yaml`, with a second application credential; the file’s header has the commands. 3. Generate the cluster from `examples/cluster-kubeadm.yaml`: a `TerraformCluster`, a `KubeadmControlPlane` of three machines and a `MachineDeployment` of two, with health checks for both. ```sh export CLUSTER_NAME=demo KUBERNETES_VERSION=v1.33.1 \ OPENSTACK_SUBNET_ID= OPENSTACK_CONTROL_PLANE_FLAVOR=m1.large \ OPENSTACK_FLAVOR=m1.large OPENSTACK_IMAGE_NAME=ubuntu-2404-kube-v1.33.1 clusterctl generate yaml --from examples/cluster-kubeadm.yaml | kubectl apply -n "$NAMESPACE" -f - ``` 4. Install a CNI once the API server answers. Nodes keep the `node.cloudprovider.kubernetes.io/uninitialized` taint until the cloud controller manager runs. ## API endpoint The [shared behavior]() applies; on OpenStack: - **Internal.** The endpoint is the Octavia load balancer’s VIP on `subnet_id`, one TCP listener per port: kube-apiserver on `Cluster.spec.clusterNetwork.apiServerPort`, and with RKE2 the supervisor on 9345. Listener idle timeouts are an hour, so watches, `kubectl exec` and `logs -f` survive. - **Public.** `api_load_balancer_public = true` puts a floating IP from `floating_ip_pool` on the VIP port and makes it the endpoint host. It needs `api_allowed_cidrs`, which restricts the listeners; the node subnet is always added to them. - **Hairpin.** The amphora provider source-NATs every client to its own address on the node subnet, so a control-plane node reaches itself through the VIP, and the control-plane security group admits the API port from the subnet. Through a public endpoint, nodes reach the floating IP via their router, which source-NATs them to its gateway address: put that address in `api_allowed_cidrs`, or the first control-plane node never reaches itself. The ovn provider rejects `api_allowed_cidrs` (a precondition stops that combination) and its hairpin is not verified. - **Endpoint guard.** `terraform_data.api_endpoint_guard` records `subnet_id`, `api_load_balancer_public`, `floating_ip_pool`, `cluster_network.api_server_port` and the load balancer’s actual provider, and fails a later plan that would change one. ## Exports Schema `captf.io/openstack-cluster/v1`. Machines receive it as `captf_cluster_outputs`. | Key | Value | | --- | --- | | `schema` | `captf.io/openstack-cluster/v1` | | `region` | The cluster’s region, the machines’ provider region | | `network_id` | Network of `subnet_id` | | `subnet_id` | `subnet_id` | | `failure_domains` | One key per availability zone, each an empty object | | `distribution` | `kubeadm` or `rke2` | | `provider_id_format` | `default` or `regional` | | `security_group_ids` | `control_plane` (control-plane and node group) and `worker` (node group) lists | | `control_plane_server_group_id` | The server group control-plane machines join, or `null` | | `node_allowed_address_cidrs` | The allowed address pairs of every node port | | `api` | The endpoint `host` and `port`, and `pools` keyed `kube_apiserver` and `rke2_supervisor`, each `{id, port}` with the backend port; `null` with a supplied endpoint | ## Tags The `captf_tags` keys and `additional_tags` map `/` to `:` in the key (`captf.io:cluster`). Neutron and Octavia resources carry them as `"="` strings in `tags`; the Nova server carries them as metadata, because a Nova tag holds at most 60 characters and no `/`. A pair longer than 255 characters fails a precondition instead of being cut. Not taggable in provider 3.4.0: security group rules, health monitors and server groups. The boot-from-volume root volume is untagged too: Nova creates it from the server’s block device mapping, which takes no metadata. > [!NOTE] > > **Design notes** > > - **Bring your own network.** The modules create no network, subnet or router, so a cluster fits whatever network design the cloud already has. > - **No node identity.** An application credential belongs to the user Terraform runs as, Keystone limits what one credential can create, and the secret would sit in state; so you supply the cloud controller manager’s `cloud.conf`. > - **Amphora by default.** Its source NAT gives the hairpin path the contract requires; `loadbalancer_provider = null` takes the cloud’s default instead. > - **Health from Octavia’s provisioning status.** Cluster health reads the load balancer’s `provisioning_status`, never `operating_status`, which follows the members. > - **No image lookup.** The server takes `image_id` or `image_name` directly, and later image changes are ignored: a Glance data source would fail every refresh and destroy once the image is rotated out, and the provider rebuilds a server in place when its image changes. > - **Register before boot.** The server depends on its pool members, so a control-plane machine is in the API pools before it boots. > - **Bootstrap in user data.** OpenStack has no instance identity to fetch a staged payload with, so the control-plane payload, with the cluster CA keys, is readable from the metadata service; see [Machine](). The evidence for each is in the repository’s [DESIGN.md](). > [!WARNING] > > **Not yet verified** > > - Octavia ovn provider: hairpin from a member to its own VIP, and the `SOURCE_IP_PORT` pool method. > - Octavia tag limits, assumed equal to Neutron’s. > - How often a server refresh, and so a drift Job or destroy, fails while a server is in a transient Nova state (`REBOOT`, a resize, `RESCUE`). > - The cloud controller manager’s `manage-security-groups` editing node ports’ groups. > - Cinder and Nova availability zone names that differ, for boot volumes. > - Octavia returning a listener’s allowed CIDRs in sorted order. > - Nodes reaching a public endpoint source-NATed to the router’s gateway address. > - The kubeadm templates’ Node name equaling the server name; Nova cuts hostnames to 63 characters. > - The cloud controller manager v1.34.1 manifests with the control-plane node selector changed to `""`. > - Neutron’s ML2/OVN driver counting allowed address pairs as security group members, which native pod routing relies on. # Cluster The `ghcr.io/captf-io/openstack-cluster` image implements the [cluster role]() for a `TerraformCluster` on OpenStack. On the subnet you bring, it creates the security groups of the nodes, the Octavia load balancer behind the Kubernetes API, and a server group that spreads the control plane, and publishes them to the machines through its `exports`. ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `openstack_networking_secgroup_v2.node_security_group` | Every node: all traffic from the cluster’s nodes, NodePorts from the subnet, optional SSH, egress | Always | | `openstack_networking_secgroup_v2.control_plane_security_group` | Control-plane nodes, on top of the node group: the API backend ports from the subnet | Always | | `openstack_networking_secgroup_rule_v2.node_ingress_rules` | `any-from-nodes`, `nodeports-tcp`, `nodeports-udp`, one `ssh-from-` per `ssh_allowed_cidrs` entry | Always | | `openstack_networking_secgroup_rule_v2.node_egress_rules` | Egress anywhere, IPv4 and IPv6 | Always | | `openstack_networking_secgroup_rule_v2.control_plane_ingress_rules` | `kube_apiserver-from-subnet`, and `rke2_supervisor-from-subnet` (TCP 9345) | Always; the second with `distribution = "rke2"` | | `openstack_lb_loadbalancer_v2.api_load_balancer` | Octavia load balancer, VIP on `subnet_id` | Without a supplied endpoint | | `openstack_lb_listener_v2.api_listeners` | TCP listeners: kube-apiserver; the RKE2 supervisor on 9345 | With the load balancer; the second with `rke2` | | `openstack_lb_pool_v2.api_pools` | The pools control-plane machines join | One per listener | | `openstack_lb_monitor_v2.api_monitors` | TCP health monitors | One per pool | | `openstack_networking_floatingip_v2.api_floating_ip` | Floating IP on the VIP port, the endpoint host | With `api_load_balancer_public` | | `openstack_compute_servergroup_v2.control_plane_server_group` | Server group of the control-plane machines | Unless `control_plane_server_group_policy` is `null` | | `terraform_data.api_endpoint_guard` | Fails a plan that would move the endpoint | With the load balancer | It reads the subnet (CIDR, address family, network, region) only while a listing of visible subnets still shows it, so a destroy runs after the subnet is gone; every other plan then fails a precondition naming `subnet_id`. It reads Nova’s available zones when `availability_zones` is empty, and the load balancer’s provisioning status for health. ## Inputs Contract inputs used: `captf_contract` (validated), `captf_cluster` (resource names, `captf---`), `captf_tags`, `control_plane_endpoint` (non-null means no load balancer) and `cluster_network` (`api_server_port` sets the API port, `pods` the allowed address pairs). `captf_object`, `kubernetes_version` and `control_plane_initialized` are declared and unused. User variables, set in `TerraformCluster.spec.variables` ([Module Variables](); source: [variables.tf]()): | Variable | Type | Default | Description | | --- | --- | --- | --- | | `additional_tags` | `map(string)` | `{}` | Extra tags on every resource. At most 44; keys follow the Nova metadata key rule and must not start with `captf.io:`; `=` fits in 255 characters | | `api_allowed_cidrs` | `list(string)` | `[]` | Clients the API listeners accept, of the subnet’s address family. Empty: any client that can route to the VIP. Required with `api_load_balancer_public`; refused with the `ovn` provider. The node subnet is always added | | `api_load_balancer_public` | `bool` | `false` | Put a floating IP on the VIP and use it as the endpoint | | `availability_zones` | `list(string)` | `[]` | Nova zones to report as failure domains. Empty: every available zone | | `control_plane_server_group_policy` | `string` | `"soft-anti-affinity"` | `soft-anti-affinity`, `anti-affinity`, or `null` for no server group | | `distribution` | `string` | `"kubeadm"` | `kubeadm` or `rke2`, which adds the supervisor listener on 9345 | | `floating_ip_pool` | `string` | `null` | External network for the floating IP. Required with `api_load_balancer_public` | | `loadbalancer_provider` | `string` | `"amphora"` | Octavia provider; `null` takes the cloud’s default | | `node_allowed_address_cidrs` | `list(string)` | `[]` | Extra source CIDRs every node port may send from, for example a kube-vip address | | `pod_address_pairs` | `bool` | `false` | Add the pod CIDRs to every node port’s allowed address pairs, for CNIs that route pods without encapsulation | | `provider_id_format` | `string` | `"default"` | `default` (`openstack:///`) or `regional` (`openstack:///`); must match the cloud controller manager’s `OS_CCM_REGIONAL` | | `region` | `string` | `null` | OpenStack region; `null` takes the identity’s region | | `ssh_allowed_cidrs` | `list(string)` | `[]` | CIDRs allowed on TCP 22 of every node | | `subnet_id` | `string` | `null` | **Required.** UUID of the subnet for the VIP and every node | ## Outputs | Output | Value | | --- | --- | | `control_plane_endpoint` | The VIP, or the floating IP with `api_load_balancer_public`, on the API port; the supplied endpoint when one is given | | `failure_domains` | One entry per availability zone, all eligible for the control plane, no attributes | | `exports` | See [Exports]() | | `health` | See below | | `api_load_balancer_id` | Not a contract output: the Octavia load balancer’s UUID, `null` without one | ## Health From the load balancer’s Octavia `provisioning_status`, re-read on every refresh; `operating_status` follows the members, which are down during every normal control-plane bring-up, and is not used. | Octavia state | Contract state | Reason | | --- | --- | --- | | `ACTIVE` | `running`, healthy | none | | `PENDING_UPDATE` | `running`, healthy | none: the load balancer keeps serving while a member is added | | `PENDING_CREATE` | `pending` | `LoadBalancerPendingCreate` | | `ERROR` | `degraded` | `LoadBalancerError` | | `PENDING_DELETE` | `terminated` | `LoadBalancerPendingDelete` | | `DELETED` | `terminated` | `LoadBalancerDeleted` | | deleted outside Terraform | `terminated` | `LoadBalancerNotFound` | | anything else | `unknown` | `LoadBalancerStatusUnknown` | With a supplied endpoint the health is `running` and healthy: there is no load balancer to observe. ## Limitations > [!WARNING] > > **Switching to a supplied endpoint plans the load balancer’s destruction** > > On a cluster whose module created the load balancer plans its destruction; only the [destructive-plan guard]() stops it. - **The endpoint is fixed.** `subnet_id`, `api_load_balancer_public`, `floating_ip_pool`, `cluster_network.api_server_port` and `loadbalancer_provider` cannot change once the load balancer exists. - **No hosted control planes.** With no endpoint supplied, the module always creates a load balancer and reports its endpoint, which Cluster API copies first, so a control-plane provider that sets the endpoint itself later is not supported. - **Failure domains follow availability.** Without `availability_zones`, a zone Nova marks unavailable drops out at the next refresh: pin the list for production clusters. - **Rule changes replace rules.** Every security group rule attribute forces a new rule, so a changed CIDR plans a delete and a create, which the destructive-plan guard holds for approval. ## Exceptions - NodePorts are open from the node subnet, beyond the API port: Octavia load balancers for Services of type `LoadBalancer` source-NAT from the subnet, and the alternative, the cloud controller manager’s `manage-security-groups`, edits the node ports’ groups behind the machine role. - No `tfcapi-lint` warning is allowed. ## Example terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo spec: source: image: ghcr.io/captf-io/openstack-cluster:v0.1.0-opentofu identityRef: name: openstack defaults: identityRef: name: openstack variables: subnet_id: 5c1d7a0e-2b4f-4e83-9a61-0d8f3b2c4e71 ``` # Machine The `ghcr.io/captf-io/openstack-machine` image implements the [machine role]() for a `TerraformMachine` on OpenStack. It creates one Nova server on a Neutron port of the cluster’s subnet and boots it with the bootstrap payload; for a control-plane machine it first adds the port’s address to the cluster’s API pools. Everything about the cluster comes from `captf_cluster_outputs`, the [cluster role’s exports](). ## What it creates | Resource | Purpose | When | | --- | --- | --- | | `openstack_networking_port_v2.node_port` | Port on the cluster subnet, with the security groups and allowed address pairs | Always | | `openstack_lb_member_v2.api_members` | Membership in each API pool, removed by this machine’s destroy | One per exported pool, on a control-plane machine | | `openstack_compute_instance_v2.node_instance` | The Nova server, named `machine_name` | Always, after the port and the members | The server waits for the pool members, so a control-plane machine is in the API pools before it boots, as `kubeadm init` and RKE2 joins require (see [Control-plane machines]()). It reads the server’s status for health, except while the server is in `BUILD`. ## Inputs Contract inputs used: - `captf_contract` (validated), `captf_object` (the port name), `captf_tags`; - `captf_cluster_outputs`: network, subnet, groups, server group, pools, zones and region, validated to schema `captf.io/openstack-cluster/v1` or `{}`; - `machine_name`: the server name, member names, the `Hostname` address and the default zone pick; - `bootstrap_data` and `bootstrap_format` (see [Bootstrap](<#bootstrap>)); - `failure_domain`: the server’s availability zone, one of the cluster’s; - `kubernetes_version`: fills `{version}` and `{semver}` in `image_name`, without its `+rke2rN` suffix; - `control_plane`: pool membership, the control-plane group and the server group. `captf_cluster` is declared and unused. User variables, set in the `TerraformMachineTemplate`’s `spec.template.spec.variables` ([Module Variables](); source: [variables.tf]()): | Variable | Type | Default | Description | | --- | --- | --- | --- | | `additional_security_group_ids` | `list(string)` | `[]` | Extra Neutron security group UUIDs for the port | | `additional_tags` | `map(string)` | `{}` | Extra tags, with the cluster role’s rules | | `config_drive` | `bool` | `false` | Attach a config drive with the user data and metadata; it does not turn the metadata service off | | `external_cluster_exports` | `any` | `null` | The exports of an externally managed `TerraformCluster`, schema `captf.io/openstack-cluster/v1` | | `flavor_name` | `string` | `null` | **Required.** Nova flavor | | `image_id` | `string` | `null` | Glance image UUID. Set exactly one of `image_id` and `image_name`; boot from volume needs `image_id` | | `image_name` | `string` | `null` | Glance image name, resolved to an ID by the provider through Glance at create; must match exactly one image. `{version}` (`v1.31.4`) and `{semver}` (`1.31.4`) stand for `kubernetes_version` | | `key_pair` | `string` | `null` | Nova key pair for SSH | | `root_volume_size_gib` | `number` | `null` | Boot from a new Cinder volume of this size, deleted with the server; `null` boots from the flavor’s disk | | `root_volume_type` | `string` | `null` | Cinder volume type of the root volume | ## Outputs | Output | Value | | --- | --- | | `provider_id` | `openstack:///`, or `openstack:///` with the cluster’s `provider_id_format = "regional"`; `null` once the server is gone | | `addresses` | `InternalIP` for each fixed IP of the port, then `Hostname`, `machine_name` | | `failure_domain` | The availability zone: the requested one, or the module’s pick | | `interruptible` | Always `false`: Nova has no spot servers | | `health` | See below | | `api_member_ids` | Not a contract output: Octavia member UUIDs keyed by pool, `{}` on a worker | | `node_port_id` | Not a contract output: the Neutron port’s UUID, `null` once it is gone | `provider_id` is what the OpenStack cloud controller manager writes to the Node: `makeInstanceID` in [cloud-provider-openstack v1.34.1]() returns `openstack:///`, or `openstack:///` when `OS_CCM_REGIONAL=true`. The addresses mirror what it reports for a server without floating IPs. Without a requested failure domain, the module picks a zone from the sha256 of `machine_name`, the same way in every repository. ## Health From the server’s Nova status, re-read on every refresh. | Nova status | Contract state | Reason | | --- | --- | --- | | `ACTIVE`, `MIGRATING`, `PASSWORD` | `running`, healthy | none | | `BUILD` | `pending` | `ServerBuilding` | | `REBOOT`, `HARD_REBOOT` | `pending` | `ServerRebooting` | | `REBUILD` | `pending` | `ServerRebuilding` | | `RESIZE`, `VERIFY_RESIZE`, `REVERT_RESIZE` | `pending` | `ServerResizing` | | `SHUTOFF` | `stopped` | `ServerShutOff` | | `SUSPENDED` | `stopped` | `ServerSuspended` | | `PAUSED` | `stopped` | `ServerPaused` | | `SHELVED`, `SHELVED_OFFLOADED` | `stopped` | `ServerShelved` | | `RESCUE`, `ERROR` | `degraded` | `ServerRescued`, `ServerError` | | `SOFT_DELETED`, `DELETED` | `terminated` | `ServerDeleted` | | deleted outside Terraform | `terminated` | `ServerNotFound` | | `UNKNOWN`, anything else | `unknown` | `ServerStatusUnknown` | Provider 3.4.0 reads only some of these. The server resource reads `ACTIVE`, `BUILD`, `SHUTOFF`, `PAUSED`, `SHELVED`, `SHELVED_OFFLOADED`, `MIGRATING` and `ERROR`; in any other status its refresh fails, and so does every drift Job and destroy until the server leaves that status. The status read is skipped while the server is in `BUILD`, where it would fail, so health then comes from the server itself and a server stuck building can still be destroyed. ## Lifecycle Machines are immutable: the module applies once, then refreshes for health and destroys on delete. Its inputs, `captf_cluster_outputs` included, are pinned at the first apply, so a later change to the cluster’s zones or exports never reaches a running machine. Changes outside Terraform show up in drift reports only: - The server’s image is ignored after creation, so a rotated or deleted image never shows as drift (the provider would rebuild the server in place). - A deleted flavor shows as a planned replacement; a renamed or resized one as an in-place resize. - On destroy, the server goes before its pool members; Cluster API has drained the node by then, and the load balancer’s monitor takes the backend out. ## Bootstrap - `bootstrap_data` goes to Nova as user data unchanged: valid base64 passes through as is, so gzipped cloud-config works, and state keeps only a SHA1 of it. Gzipped Ignition fails a precondition. - Nova accepts at most 65,535 bytes of base64 user data; a larger payload fails a precondition. Compress it (CAPRKE2 `gzipUserData`) if needed. - **Who can read it.** OpenStack has no instance identity to fetch a staged payload with, so there is no `bootstrap_delivery`: the control-plane payload, with the cluster CA keys, sits in user data, and anything that reaches the metadata service (169.254.169.254) on the node can read it. A config drive does not turn the metadata service off. Deny 169.254.169.254/32 to pods with a CNI `NetworkPolicy`. Cluster API Provider OpenStack has the same exposure. ## Limitations > [!WARNING] > > **Keep machine\_name within 63 characters and DNS-safe** > > Nova derives the hostname from the server name, cut to 63 characters. Keep `machine_name` within 63 characters and DNS-safe, or the Node name does not match the server and the cloud controller manager cannot find it. - **No spot.** `interruptible` is always `false`. - **Port security groups.** Keep the cloud controller manager’s `manage-security-groups` off: groups it adds to the port show as drift. ## Exceptions - The server is named `machine_name`, not the conventions’ prefixed name, because the cloud controller manager finds servers by Node name. - The boot-from-volume root volume carries no tags: Nova creates it from the block device mapping, which takes no metadata, and a separately created, tagged volume would need a Cinder availability zone named like the Nova one. - Bootstrap payloads are readable from the metadata service (see [Bootstrap](<#bootstrap>)), a deviation from the contract’s checklist. - There is no `rejects_spot_control_plane` test: Nova has no spot servers. - No capacity labels on the image: there is no default flavor to describe. ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 spec: template: spec: source: image: ghcr.io/captf-io/openstack-machine:v0.1.0-opentofu variables: flavor_name: m1.large image_name: ubuntu-2404-kube-{version} ``` # No-op The no-op modules implement all three roles of the [v1alpha1 module contract]() and provision nothing. Every resource is a `terraform_data` holding the inputs it was given, so the plans, the state and the outputs are real while no cloud is touched and no credentials are read. Use them to try CAPTF, to test a management cluster, or as the smallest starting point for a module of your own. The images are built from [`captf-io/noop-modules`](), which holds the modules and is their only source. - **Cluster** --- A stand-in load balancer, an endpoint that never resolves, one failure domain. - **Machine** --- A stand-in instance per `Machine`, with a stable provider ID. - **MachinePool** --- A stand-in scaling group per `MachinePool`, one provider ID per replica. ## Images | Role | Image | Page | | --- | --- | --- | | cluster | `ghcr.io/captf-io/noop-cluster` | [Cluster]() | | machine | `ghcr.io/captf-io/noop-machine` | [Machine]() | | machinepool | `ghcr.io/captf-io/noop-machinepool` | [MachinePool]() | The images need no Terraform provider: `terraform_data` is built into Terraform and OpenTofu. They are built on the [Terraform and OpenTofu base images](), for `linux/amd64` and `linux/arm64`, with the same tags as every set (see [Images and tags]()). The modules need Terraform 1.5 or later; every OpenTofu release works. ## Prerequisites - **A management cluster with CAPTF installed.** See [Install the provider](). - **An identity.** Every `TerraformCluster` names a `TerraformClusterIdentity`, even one whose modules read no credentials. The Secret behind it can hold anything: the provider’s `templates/identity.yaml` with its placeholder keys left as they are is enough. Nothing else: no network, no node image, no cloud controller manager and no CNI. ## Quick start Follow the [Quick Start](), with the published images, which its third step selects. When you generate the cluster, set: ```sh export TERRAFORM_CLUSTER_IMAGE=ghcr.io/captf-io/noop-cluster:opentofu export TERRAFORM_MACHINE_IMAGE=ghcr.io/captf-io/noop-machine:opentofu ``` The `opentofu` tag is the newest release; pin a release tag such as `vX.Y.Z-opentofu`, or a digest, in anything you keep. The `terraform` tags run the same modules on Terraform. ## What you will see The `TerraformCluster` and each `TerraformMachine` reach `Ready=True`, and the `Cluster` and its `Machine`s reach phase `Provisioned`: the modules return an endpoint and provider IDs. No `Machine` reaches `Running`, and `KubeadmControlPlane` never reports itself initialized, because no node ever boots or registers. The images exercise CAPTF, not Kubernetes. > [!WARNING] > > **Remediation starts after about 30 minutes** > > No `Node` ever registers, so once a `MachineHealthCheck`’s node-startup timeout passes, Cluster API starts remediating the `Machine`s. See [Remediation](). ## Exports The cluster role’s `exports` carry one key, `backend_id`: the id of its stand-in load balancer. The machine and machinepool roles record it in their own state, so a machine’s state holds a value that exists only after the cluster applied, as with a real set. The exports have no `schema` key. ## Tags Each role records `captf_tags` in its resource, so the tags show in the state and the plan. There is nothing to tag. ## Writing your own The no-op modules are the smallest modules that meet the contract, which makes them the starting point of [Your First Module](). Each role page lists the inputs it reads and the outputs it returns. # Cluster The no-op cluster module stands in for what a real cluster module creates around a workload cluster’s nodes. Its image is `ghcr.io/captf-io/noop-cluster`. ## What it creates One `terraform_data` resource, `load_balancer`, the stand-in for a load balancer and a network. It holds the inputs below, so a change to any of them shows as a change in the plan. Its id feeds the `exports`, so machines receive a value that exists only after this module applied. ## Inputs The module declares the contract inputs only; it has no variables of its own, so `spec.variables` has nothing to set. | Input | Use | | --- | --- | | `captf_contract` | Declared, as the contract requires | | `captf_cluster`, `captf_object` | Recorded; the object’s name also names the default endpoint | | `captf_tags` | Recorded | | `control_plane_endpoint` | Returned as the endpoint when set | | `kubernetes_version`, `control_plane_initialized`, `cluster_network` | Recorded | ## Outputs | Output | Value | | --- | --- | | `control_plane_endpoint` | `control_plane_endpoint` when set; otherwise host `noop-.invalid`, port `6443` | | `failure_domains` | One failure domain, `fd1`, eligible for the control plane | | `exports` | `{ backend_id = "noop-backend-" }`, the stand-in load balancer’s id | | `health` | Always running and healthy | The `.invalid` top-level domain never resolves (RFC 2606), so the default endpoint is valid and stable but reaches nothing. ## Health The module always reports `running` and healthy: there is no infrastructure to check. ## Limitations - **The endpoint reaches nothing.** No load balancer exists behind it. - **One failure domain.** Machines and pools all land in `fd1`. ## Example terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo spec: source: image: ghcr.io/captf-io/noop-cluster:opentofu identityRef: name: noop defaults: identityRef: name: noop ``` # Machine The no-op machine module stands in for one instance per `Machine`. Its image is `ghcr.io/captf-io/noop-machine`. ## What it creates One `terraform_data` resource, `instance`, holding the inputs below. It decodes `bootstrap_data` the way a module that passes a plain user-data argument would, which shows the controller’s base64 encoding round-trips; the decoded value stays sensitive in the plan. ## Inputs The module declares the contract inputs only; it has no variables of its own, so `spec.variables` has nothing to set. | Input | Use | | --- | --- | | `captf_contract` | Declared, as the contract requires | | `captf_cluster`, `captf_object`, `captf_tags` | Recorded | | `captf_cluster_outputs` | Its `backend_id` is recorded | | `machine_name` | Names the provider ID | | `bootstrap_data` | Decoded and recorded, as user data would be | | `bootstrap_format`, `kubernetes_version`, `control_plane` | Recorded | | `failure_domain` | Recorded and returned | ## Outputs | Output | Value | | --- | --- | | `provider_id` | `noop:////`, stable per `Machine` | | `addresses` | One `InternalIP`, `10.0.0.1` | | `failure_domain` | The requested failure domain | | `interruptible` | `false` | | `health` | Always running and healthy | No `Node` carries the provider ID, so the `Machine` never gets a `nodeRef`. A real module returns the ID in the format its cloud controller manager or kubelet sets. ## Image labels The image carries the scale-from-zero labels a cluster autoscaler reads through `TerraformMachineTemplate.status`: `io.captf.capacity` (`{"cpu":"2","memory":"4Gi"}`) and `io.captf.node-info` (`{"architecture":"","operatingSystem":"linux"}`, with the image’s own architecture). See [OCI labels](). ## Health The module always reports `running` and healthy. ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 spec: template: spec: source: image: ghcr.io/captf-io/noop-machine:opentofu ``` # MachinePool The no-op machinepool module stands in for one native scaling group per `MachinePool`. Its image is `ghcr.io/captf-io/noop-machinepool`. ## What it creates One `terraform_data` resource, `group`, holding the inputs below. Like the machine role, it decodes `bootstrap_data` as user data would be, and the decoded value stays sensitive. ## Inputs The module declares the contract inputs only; it has no variables of its own, so `spec.variables` has nothing to set. | Input | Use | | --- | --- | | `captf_contract` | Declared, as the contract requires | | `captf_cluster`, `captf_object`, `captf_tags` | Recorded; the object’s name also names the provider IDs | | `captf_cluster_outputs` | Its `backend_id` is recorded | | `machinepool_name` | Recorded | | `replicas` | The group’s size: one provider ID per replica | | `bootstrap_data` | Decoded and recorded, as user data would be | | `bootstrap_format`, `failure_domains`, `cluster_failure_domains`, `kubernetes_version`, `node_labels` | Recorded | | `autoscaling` | Declared, as the contract requires, but not read | ## Outputs | Output | Value | | --- | --- | | `provider_id` | `noop-group:////`, stable per `TerraformMachinePool` | | `provider_id_list` | `noop://///` for each replica ``, sorted | | `replicas` | `replicas` | | `instances` | One `running` instance per provider ID | | `health` | Always running and healthy | ## Autoscaling The module always sets the group’s size from `replicas`, so it has nothing to leave to an autoscaler: it accepts `autoscaling` and ignores it. A real module that hands the size to its cloud’s autoscaler reads it; see [Machine Pools](). ## Health The module always reports `running` and healthy. ## Example terraformmachinepool.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: demo-pool-0 labels: cluster.x-k8s.io/cluster-name: demo spec: source: image: ghcr.io/captf-io/noop-machinepool:opentofu ``` # Module Authors # Your First Module This tutorial writes a small machine-role module from scratch, lints it, packages it as the OCI image CAPTF runs, lints that image, and references it from a `TerraformMachineTemplate`. It ends with a working, lint-clean module image; it does not provision real infrastructure, since the module in this tutorial creates no cloud resources. For the full set of rules a module must follow, see the [module contract]() and the [image contract](). > [!NOTE] > > **Before you begin** > > - `tfcapi-lint`, installed as in [tfcapi-lint](). > - `podman` or `docker`, to build the image. > - A directory to work in. This tutorial calls it `machine/`. Terraform or OpenTofu itself is not required on your machine: the reference Containerfile below downloads it inside the build, and `tfcapi-lint module` parses the module’s files directly. ## 1\. Write the module A machine-role module implements one Terraform/OpenTofu module that CAPTF calls once per `TerraformMachine`. It receives the [common contract inputs]() plus the [machine role’s inputs](), and must return the [machine role’s outputs]() plus the common `health` output. This tutorial’s module returns fixed values instead of calling a cloud provider, which keeps every input and output visible and needs no credentials or provider plugins to build or lint. Swap the fixed values for calls to your cloud’s Terraform/OpenTofu provider once you understand the shape. Create four files in `machine/`. ### `variables.tf` The contract inputs. `captf_contract`, `captf_cluster`, `captf_object` and `captf_tags` apply to every role; `captf_cluster_outputs` applies to the machine and machinepool roles only (the cluster role produces it, and so does not receive it); the rest are specific to the machine role. variables.tf ```hcl # Contract inputs of the machine role, v1alpha1 (https://captf.io/docs/module-author/contract/v1alpha1/common.html # and machine.html). variable "captf_contract" { type = string } variable "captf_cluster" { type = object({ name = string namespace = string }) } variable "captf_object" { type = object({ kind = string name = string namespace = string }) } # The cluster module's exports. The controller always sets it; the default # follows the contract skeleton (machine.md). variable "captf_cluster_outputs" { type = any default = null } variable "captf_tags" { type = map(string) } variable "machine_name" { type = string } # Base64 of the bootstrap Secret's value. variable "bootstrap_data" { type = string sensitive = true } variable "bootstrap_format" { type = string } variable "failure_domain" { type = string default = null } variable "kubernetes_version" { type = string default = null } variable "control_plane" { type = bool } ``` ### `main.tf` The module’s own logic. This tutorial stands in a `terraform_data` resource for a real instance, so the module has something to hold its inputs; a real module replaces this with the resources that create an instance. main.tf ```hcl # No-op machine module: implements the v1alpha1 machine role with no cloud. # The stand-in for an instance. user_data decodes bootstrap_data the way a # module feeding a plain user-data argument would, which proves the # controller's base64 encoding round-trips (the value stays sensitive). resource "terraform_data" "instance" { input = { cluster = var.captf_cluster object = var.captf_object machine_name = var.machine_name tags = var.captf_tags backend_id = try(var.captf_cluster_outputs.backend_id, null) user_data = base64decode(var.bootstrap_data) bootstrap_format = var.bootstrap_format failure_domain = var.failure_domain kubernetes_version = var.kubernetes_version control_plane = var.control_plane } } ``` ### `outputs.tf` The contract outputs. `provider_id`, `addresses` and `failure_domain` are required; `interruptible` must be declared even when it is always `false`; `health` is the common output every role returns. outputs.tf ```hcl # Contract outputs of the machine role, v1alpha1. # Stable per Machine. With no Node behind it, the e2e suite never expects a # nodeRef; a real module emits the CCM's or kubelet's format (machine.md). output "provider_id" { value = "noop:///${var.captf_object.namespace}/${var.machine_name}" } output "addresses" { value = [{ type = "InternalIP", address = "10.0.0.1" }] } output "failure_domain" { value = var.failure_domain } output "interruptible" { value = false } output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } } ``` ### `versions.tf` versions.tf ```hcl terraform { # terraform_data needs Terraform 1.4; the variants' plantimestamp() needs # 1.5. Every OpenTofu release (1.6+) has both. required_version = ">= 1.5" } ``` This module declares no `required_providers`: `terraform_data` ships with Terraform/OpenTofu itself, so there is no provider to install or mirror. A module that calls a cloud provider adds a `required_providers` block here as usual. ## 2\. Lint the module ```sh tfcapi-lint module ./machine --role machine --strict ``` A clean module prints an empty finding list and an all-zero summary. > [!NOTE] > > **`--strict` fails on warnings too** > > `--strict` fails the command on a warning as well as an error, so a lint-clean module here stays lint-clean once you add real inputs, outputs and user variables. See [tfcapi-lint]() for the full command reference and [tfcapi-lint CLI]() for every check ID. ## 3\. Package it as an image CAPTF runs a module as one OCI image that bundles the module’s files and the Terraform or OpenTofu binary; there is no separate module source and no separate runtime image. Save the reference Terraform-based Containerfile into `machine/`: Containerfile.terraform ```dockerfile # Reference source image, Terraform base (docs/book/src/module-author/image-contract.md "Reference: # Terraform base"). Run from your module's root directory: # # podman build -f Containerfile.terraform --build-arg ROLE=cluster \ # --build-arg IMAGE_SOURCE=https://github.com// \ # --build-arg IMAGE_REVISION="$(git rev-parse HEAD)" \ # --build-arg IMAGE_VERSION= -t /: . # # The image tag is the module version. Lint first: # tfcapi-lint module . --role cluster --strict ARG RUNTIME_VERSION=1.16.4 # The org.opencontainers.image.* labels below (image-contract.md "OCI # labels"): leave these unset for a local/test build, or pass them from # your CI pipeline (source repo URL, commit SHA, the image tag). ARG IMAGE_SOURCE="" ARG IMAGE_REVISION="" ARG IMAGE_VERSION="" # Optional but recommended: hermetic provider mirror for the platforms you # publish. Needs registry egress at build time; drop this stage (and the # COPY --from=mirror below) for a non-hermetic image. FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION} AS mirror WORKDIR /src COPY . /src # get: `providers mirror` refuses a module whose nested local modules are not # installed; get installs them (no providers), in this stage only. # mkdir: `providers mirror` does not create the target when the module # requires no providers, and the COPY --from=mirror below needs it. RUN terraform get \ && mkdir -p /captf/providers \ && terraform providers mirror -platform=linux_amd64 -platform=linux_arm64 /captf/providers FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION} ARG ROLE=cluster ARG RUNTIME_VERSION ARG IMAGE_SOURCE ARG IMAGE_REVISION ARG IMAGE_VERSION COPY --from=mirror /captf/providers /captf/providers COPY . /captf/module RUN ln -s /bin/terraform /captf/runtime \ && adduser -D -u 65532 captf \ && chown -R 65532:65532 /captf USER 65532 LABEL io.captf.contract="v1alpha1" \ io.captf.role="${ROLE}" \ io.captf.runtime="terraform" \ io.captf.runtime.version="${RUNTIME_VERSION}" \ org.opencontainers.image.source="${IMAGE_SOURCE}" \ org.opencontainers.image.revision="${IMAGE_REVISION}" \ org.opencontainers.image.version="${IMAGE_VERSION}" ``` > [!TIP] > > **OpenTofu works too** > > An OpenTofu-based equivalent is also available ([`Containerfile.opentofu`]()); either runtime satisfies the contract. Build from inside `machine/`: ```sh export IMAGE=registry.example.com/acme/machine:v1.0.0 podman build -f Containerfile.terraform --build-arg ROLE=machine -t "$IMAGE" . ``` `IMAGE` is the registry, repository and tag you push to; `ROLE=machine` tags the image with the role it implements. See the [image contract]() for the fixed paths, OCI labels and execution environment every image must satisfy. ## 4\. Lint the image `tfcapi-lint image` reads a pushed registry reference or a local OCI layout. Without a registry to push to yet, save the image podman just built to a local OCI directory and lint that: ```sh podman save --format oci-dir -o /tmp/first-module-image "$IMAGE" tfcapi-lint image --role machine "oci:/tmp/first-module-image" ``` Once you push `$IMAGE` to a registry, lint the pushed reference the same way: `tfcapi-lint image --role machine "$IMAGE"`. ## 5\. Reference it from a TerraformMachineTemplate A `TerraformMachineTemplate` is what a `MachineDeployment`, a `KubeadmControlPlane` or a `MachineSet` points at to create `TerraformMachine`s; its `spec.template.spec.source.image` names the image you just built: first-module-template.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: first-module namespace: default spec: template: spec: source: image: registry.example.com/acme/machine:v1.0.0 ``` Apply it: ```sh kubectl apply -f first-module-template.yaml ``` `kubectl get terraformmachinetemplate first-module -n default` confirms the object exists; there is nothing to provision yet, since nothing references this template as a `MachineDeployment`, `MachineSet` or control plane’s `infrastructureRef` does. A `TerraformMachine` created from it needs an identity to run under: either its own `spec.identityRef`, or one inherited from its `TerraformCluster`’s `spec.defaults.identityRef`, as set up in the [quick start](). ## Next steps - Walk through the [quick start]() to wire a machine template like this one into a full cluster. - Read the [module contract]() for every rule a production module must follow, and the [cluster role]() and [machine pool role]() for the other two roles a full deployment needs. - Read [Runtime Environment]() for what a module sees when the runner actually executes it. # Job Inputs This explains what goes into a CAPTF Job’s Terraform inputs — `terraform.tfvars.json` and the generated root `main.tf.json` — and where each value comes from on the Kubernetes side. For each input’s type and whether it is required, see the [module contract](); this page traces the plumbing that fills the contract in, for anyone administering CAPTF or writing a module against it. A `TerraformCluster` runs the module’s `cluster` role, each `TerraformMachine` runs its `machine` role, and each `TerraformMachinePool` runs its `machinepool` role. All three roles are rendered the same way from the same kind of input structure, so this page covers all three side by side. ## 1\. The pipeline ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD subgraph capi[CAPI objects] Cluster["Cluster
spec.topology.version
spec.clusterNetwork
spec.controlPlaneEndpoint
status.initialization.controlPlaneInitialized"] Machine["Machine
spec.version
spec.failureDomain
spec.bootstrap.dataSecretName
labels[cluster.x-k8s.io/control-plane]"] MachinePool["MachinePool
spec.replicas
spec.failureDomains
spec.template.spec.version
spec.template.metadata.labels
autoscaler min/max-size annotations"] Bootstrap["Bootstrap Secret (KubeadmConfig/RKE2Config)
keys: value, format"] end subgraph objs[Terraform* objects] TC["TerraformCluster
spec + status"] TM["TerraformMachine
spec + status"] TMP["TerraformMachinePool
spec + status"] end Exports["TerraformCluster state
outputs: exports, failure_domains"] Cluster --> ClusterInputs TC --> ClusterInputs ClusterInputs[Cluster inputs] --> RenderC[Render, cluster role] RenderC --> DurableC["Durable Secret
captf-inputs-c-name"] RenderC --> RunC["Per-run Secret
captf-run-job
mounted at /captf/config"] RunC --> RunnerC[Runner] --> ModuleC[Module, cluster role] ModuleC --> Exports Machine --> MachineInputs Bootstrap --> MachineInputs Cluster --> MachineInputs TM --> MachineInputs Exports -->|only while not yet provisioned| MachineInputs MachineInputs[Machine inputs] --> RenderM[Render, machine role] RenderM --> DurableM["Durable Secret
captf-inputs-m-name"] RenderM --> RunM["Per-run Secret
captf-run-job
mounted at /captf/config"] RunM --> RunnerM[Runner] --> ModuleM[Module, machine role] MachinePool --> PoolInputs Bootstrap --> PoolInputs Cluster --> PoolInputs TMP --> PoolInputs Exports -->|every reconcile| PoolInputs PoolInputs[Machine pool inputs] --> RenderP[Render, machinepool role] RenderP --> DurableP["Durable Secret
captf-inputs-mp-name"] RenderP --> RunP["Per-run Secret
captf-run-job
mounted at /captf/config"] RunP --> RunnerP[Runner] --> ModuleP[Module, machinepool role] DurableC -.->|destroy always; drift/refresh when gated or current inputs unavailable| RenderC DurableM -.->|destroy and drift/refresh always, machine is immutable| RenderM DurableP -.->|destroy always; drift/refresh when gated or current inputs unavailable| RenderP ``` Why two Secrets exist: - **Durable inputs Secret** (`captf-inputs--`, owned by the Terraform\* object, moves with it across `clusterctl move`): the record of what was rendered when the controller last **started** an apply Job for this object (written at Job start, not gated on the Job’s success). Destroy always prefers it; a **cluster or pool** (both mutable) whose durable Secret is missing falls back to freshly built current inputs, but a **machine** (immutable) never does — with no durable Secret and no live spec to fall back to, destroy has nothing to run against until the Secret is restored. Drift and refresh follow the same asymmetry: a cluster or pool checks against its *current* inputs when they build cleanly, falling back to the durable Secret only when they don’t (a dependency gate, an owner gone); a machine, being immutable, always checks against the durable Secret. - **Per-run Secret** (`captf-run-`, owned by the Job, mounted at `/captf/config`, deleted when the Job finishes): a private copy the Job’s pod reads, so a rewrite of the durable Secret by a later reconcile can never change the files under a running Job. ## 2\. Cluster role inputs For the full type table, see [`contract/v1alpha1/common.md`]() and [`cluster.md`](). The Kubernetes-side sources are: | tfvars key | Type | Source | | --- | --- | --- | | `captf_contract` | `string` | The fixed contract version, `"v1alpha1"` | | `captf_cluster` | `object{name,namespace}` | The owning `Cluster`’s `metadata.name`/`metadata.namespace` | | `captf_object` | `object{kind,name,namespace}` | The `TerraformCluster`’s own kind/name/namespace | | `captf_tags` | `map(string)` | Fixed keys built from the `Cluster` name, the `TerraformCluster` namespace/kind/name, and its `cluster.x-k8s.io/cloned-from-name` annotation (empty string if absent) | | `control_plane_endpoint` | `object{host,port}` or `null` | `Cluster.spec.controlPlaneEndpoint` when valid; else `TerraformCluster.spec.controlPlaneEndpoint` when valid; else `null`. Always `null` once the `captf.io/endpoint-source` annotation is `module` | | `kubernetes_version` | `string` or `null` | `Cluster.spec.topology.version`; `null` without a ClusterClass | | `control_plane_initialized` | `bool` | `Cluster.status.initialization.controlPlaneInitialized`, **latched**: `true` once the durable Secret’s last rendered value was `true`, so it is never rendered `false` again after an apply saw `true` | | `cluster_network` | `object{pods,services,service_domain,api_server_port}` or `null` | `Cluster.spec.clusterNetwork`; `null` only when every attribute is unset. CIDR lists render as `[]`, not `null`, when the network object is present but a list is empty | A trimmed real example: terraform.tfvars.json ```json { "captf_contract": "v1alpha1", "captf_cluster": { "name": "prod", "namespace": "team-a" }, "captf_object": { "kind": "TerraformCluster", "name": "prod", "namespace": "team-a" }, "captf_tags": { "captf.io/cluster": "prod", "captf.io/kind": "TerraformCluster", "captf.io/managed-by": "captf", "captf.io/name": "prod", "captf.io/namespace": "team-a", "captf.io/template": "" }, "control_plane_endpoint": { "host": "api.prod.example.com", "port": 6443 }, "kubernetes_version": "v1.36.2", "control_plane_initialized": true, "cluster_network": { "pods": ["192.168.0.0/16"], "services": ["10.96.0.0/12"], "service_domain": "cluster.local", "api_server_port": 6443 } } ``` The matching root `main.tf.json` declares one Terraform variable per key above (nullable ones get `"default": null`), a `module "role"` call forwarding every variable by name to the module, and a `sensitive = true` re-export of every required cluster output (`control_plane_endpoint`, `failure_domains`, `exports`, `health`). The cluster role never declares `captf_cluster_outputs`: it has no such input at all, not even `null`. ## 3\. Machine role inputs For the full type table, see [`common.md`]() and [`machine.md`](). The Kubernetes-side sources are: | tfvars key | Type | Source | | --- | --- | --- | | `captf_contract`, `captf_cluster`, `captf_object`, `captf_tags` | as above | Same construction as the cluster role, but `captf_object`/`captf_tags` name the `TerraformMachine` | | `captf_cluster_outputs` | `any` | The owning `TerraformCluster`’s `exports` output, verbatim ([section 5](<#5-what-a-cluster-hands-to-its-machines-and-pools>)). `{}` for an externally managed cluster | | `machine_name` | `string` | The owning `Machine`’s `metadata.name` (may differ from the `TerraformMachine`’s own name) | | `bootstrap_data` | `string`, sensitive | **Base64** (standard, padded) of the raw bytes of the bootstrap Secret’s `value` key, named by `Machine.spec.bootstrap.dataSecretName`. Encoded unconditionally for every bootstrap provider, since a gzip payload (CAPRKE2 `gzipUserData: true`) is not valid UTF-8 and cannot be a JSON/HCL string as-is. Contains the cluster CA and service-account keys for control-plane machines — treat it as a secret | | `bootstrap_format` | `string` | The bootstrap Secret’s `format` key (`cloud-config` or `ignition`); `cloud-config` when the key is absent | | `failure_domain` | `string` or `null` | `Machine.spec.failureDomain`; `null` when unset | | `kubernetes_version` | `string` or `null` | `Machine.spec.version`, passed verbatim (may carry a distro suffix such as `+rke2r1`) | | `control_plane` | `bool` | Whether the owning `Machine` carries the `cluster.x-k8s.io/control-plane` label | A trimmed real example: terraform.tfvars.json ```json { "captf_contract": "v1alpha1", "captf_object": { "kind": "TerraformMachine", "name": "prod-md-0-abcde", "namespace": "team-a" }, "captf_cluster_outputs": { "network_id": "net-1", "note": "ad" }, "machine_name": "prod-md-0-abcde-xyz12", "bootstrap_data": "I2Nsb3VkLWNvbmZpZwpydW5jbWQ6...", "bootstrap_format": "cloud-config", "failure_domain": null, "kubernetes_version": "v1.36.2", "control_plane": false } ``` The `note: "ad"` above is deliberate: the machine role’s renderer keeps a module-authored `exports` value’s exact bytes through the durable Secret instead of turning `<`/`&` into their `\u...` escapes ([section 5](<#5-what-a-cluster-hands-to-its-machines-and-pools>) has the one exception to this). The root variables and re-exported outputs (`provider_id`, `addresses`, `failure_domain`, `interruptible`, `health`) follow the same pattern as the cluster role. ## 4\. Machine pool role inputs A `TerraformMachinePool` runs the module’s `machinepool` role. For the full type table and lifecycle, see [`contract/v1alpha1/machinepool.md`]() (this page does not duplicate it). Its inputs follow the same `captf_contract`/`captf_cluster`/`captf_object`/`captf_tags` pattern as the cluster and machine roles, plus `captf_cluster_outputs` (the owning TerraformCluster’s `exports`, as for a machine), and these role-specific inputs: | tfvars key | Type | Source | | --- | --- | --- | | `machinepool_name` | `string` | The owning `MachinePool`’s `metadata.name` | | `replicas` | `number` | The group’s desired capacity. With autoscaling disabled: `MachinePool.spec.replicas` (1 when unset, CAPI’s own default). With autoscaling enabled: `TerraformMachinePool.status.replicas` (the observed count from the last refresh) once one exists, else still `spec.replicas` for the first apply — either way clamped into `[autoscaling.min, autoscaling.max]`. The write-back that syncs the raw observed count to `spec.replicas` is unclamped; see [Machine pools]() | | `bootstrap_data`, `bootstrap_format` | as machine role | From the bootstrap Secret named by `MachinePool.spec.template.spec.bootstrap.dataSecretName`, encoded exactly as for a machine. Because the kubeadm bootstrap provider rewrites this Secret roughly every 7.5 minutes, a pool is re-applied at that cadence for its whole life; see [the machinepool contract]() “Bootstrap rotation” | | `failure_domains` | `list(string)` | `MachinePool.spec.failureDomains`; may be `[]`, meaning the module chooses | | `cluster_failure_domains` | `list(string)` | The names of the owning `TerraformCluster`’s own failure domains — from its state’s `failure_domains` output when it has state, else from `TerraformCluster.status.failureDomains` for an externally managed cluster. Lets a pool module spread across every domain when `failure_domains` is `[]` | | `kubernetes_version` | `string` or `null` | `MachinePool.spec.template.spec.version`; `null` when unset | | `node_labels` | `map(string)` | `MachinePool.spec.template.metadata.labels` verbatim; `{}` when absent. Pool instances have no Machine objects, so core CAPI never syncs labels onto their Nodes; the module renders these into the bootstrap or agent configuration it controls | | `autoscaling` | `object{enabled,min,max}` | Parsed from the owning `MachinePool`’s `cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size`/`-max-size` annotations, not from the TerraformMachinePool spec: a pool has no autoscaling field of its own. `enabled` is `true` only when both annotations are present and parse as `0 <= min <= max`; otherwise `{false, 0, 0}` | Unlike a `TerraformMachine`, a `TerraformMachinePool` is mutable: its inputs are rebuilt, and its inputs hash recomputed, on every reconcile for the whole life of the object, not only until it is provisioned; see [`machinepool.md`]() for exactly which changes re-apply it. > [!NOTE] > > **The pool role’s renderer HTML-escapes string values** > > The pool role’s renderer defaults `failure_domains`, `cluster_failure_domains` and `node_labels` to `[]`/`{}` when unset through a step that, as a side effect, HTML-escapes `<`, `>` and `&` inside every string value it touches — including a module-authored `exports` value carried into `captf_cluster_outputs`. The cluster and machine roles do not escape these characters (section 3’s example). A `\uXXXX` escape and its literal character decode to the same JSON string, so a module never sees a difference; only someone reading the raw Secret bytes does. Still, don’t assume the two roles’ encodings match byte for byte when inspecting one. ## 5\. What a cluster hands to its machines and pools The cluster module’s `exports` output is the only handoff from a `TerraformCluster` to the `TerraformMachine`s and pools of its `Cluster`. It is: 1. Written by the cluster module as any Terraform value (an object is the convention). 2. Stored, like every output, in the cluster’s own Terraform state — nothing else from that state is ever read by the controller for this purpose. 3. Read by each machine’s or pool’s own reconcile whenever it builds its inputs, and rendered into `captf_cluster_outputs` (see section 4 for the pool role’s one encoding difference). **When it is read.** A machine’s inputs (and so `exports`) are rebuilt on every reconcile while it is not yet provisioned — before its status has latched `status.initialization.provisioned = true` — whether or not that reconcile ends up starting a Job. Once provisioned, the reconciler never rebuilds that machine’s inputs again, so it never re-reads `exports` or the bootstrap Secret. Concretely: - **Before a machine is provisioned**, every reconcile re-reads the cluster’s current `exports` and re-renders `captf_cluster_outputs` from it (along with the rest of the machine’s inputs). Only when an apply Job actually starts does the freshly rendered result get written to the durable Secret; a reconcile that only re-renders without starting a Job leaves the durable Secret untouched. - **Once provisioned, a machine keeps exactly the `captf_cluster_outputs` its last apply Job start wrote to its durable Secret**, forever: a later change to the cluster’s `exports` reaches only *new* machines created after that change, never existing ones. This is also why `captf_cluster_outputs` changes are never a trigger for a machine re-apply: machines are immutable, so their inputs hash, though recorded at apply time like any kind’s, is never compared against a fresh render to decide on a re-apply, and the only way a machine applies again is a fresh first apply (for example, retrying one that was interrupted before it recorded an inputs hash). - **A pool re-reads `exports` on every reconcile, for its whole life**, unlike a machine: a `TerraformMachinePool` is mutable and never latches `captf_cluster_outputs`, so a later change to the cluster’s `exports` reaches every existing pool at its next reconcile, and (unlike a machine) a changed `captf_cluster_outputs` is in the pool’s inputs hash and so re-applies it. - **Externally managed cluster** (`cluster.x-k8s.io/managed-by` annotation on the `TerraformCluster`): a machine’s or pool’s `captf_cluster_outputs` is `{}` without reading any state, since there is no cluster module run and so no cluster state to read. - **A malformed `exports` value** — one that cannot be canonicalized for hashing, for example it contains a fractional or exponent number — is an output violation on the cluster side. Every unprovisioned machine, and every pool, waiting on it stays `DependenciesReady=Unknown` with reason `WaitingForClusterExports` until the cluster module’s `exports` output is fixed; this is the one operator-visible way a cluster module can break its machines and pools purely through inputs. For example, a cluster module that creates a shared network and storage pool publishes them in `exports`: outputs.tf ```hcl output "exports" { value = { network = local.network pool = local.pool } } ``` and the machine module’s `captf_cluster_outputs` variable carries a `validation` block that rejects any value missing `network` or `pool`. The machine module then refuses to run against an externally managed cluster (`{}`) or a cluster module that doesn’t publish them, with a clear error instead of an “unsupported attribute” failure deep inside the module. ## 6\. User variables Besides the contract inputs, a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` (and their templates) can pass the module its own variables: instance sizes, CIDRs, SSH public keys, anything the module declares. See [Module Variables]() for how to set them, the merge order between `spec.variables` and `spec.variablesFrom`, and what the module sees. ## 7\. What is (and isn’t) in the inputs hash `captf.io/inputs-hash` covers exactly: the hash scheme, the contract version, the role, `spec.source.image` **as written** (tag or digest — never the resolved digest), the full rendered inputs of the role, and the merged user variables when there are any (their canonical JSON and which of them are sensitive; an object without variables keeps exactly the hash it had before variables existed). Nothing else can change the hash: object metadata (`uid`, labels, annotations, `generation`), the resolved image digest, Job metadata, drift ticks, and `spec.source.imagePullPolicy` are all absent from the hashed value and so cannot trigger anything through it. What a change does, per kind: - **TerraformCluster (mutable).** A hash change re-applies: the controller compares the freshly rendered inputs’ hash against the state’s recorded hash on every reconcile and starts a new apply Job when they differ. A cluster apply that would delete or replace a resource is blocked (`ApplyJobSucceeded=False/DestructivePlanBlocked`) until the `captf.io/approve-destructive-plan` annotation names that exact inputs hash, and with `spec.applyPolicy: Manual` every re-apply also waits for a plan approval; see [Plan Approval](). - **TerraformMachine (immutable).** Nothing, in the sense that matters: a machine’s inputs hash is recorded at apply time like any kind’s, but never compared against a fresh render to decide on a re-apply. Its inputs are still rebuilt on every reconcile up to the first successful apply (section 5), but once that apply has run and the object is provisioned, nothing — not a spec field, not a bootstrap-data change, not an `exports` change — is ever rendered or checked against it again. Roll out a new Machine instead. - **TerraformMachinePool (mutable).** A hash change re-applies, the same way as a TerraformCluster, but with no destructive-plan guard and no `applyPolicy`: an apply just runs. `replicas` is excluded from the hash while autoscaling is enabled, so the native autoscaler’s own scaling never triggers a re-apply on its own; see [`machinepool.md`]() “replicas”. ## 8\. Inspecting the inputs of a live object The durable Secret holds exactly two data keys, `main.tf.json` and `terraform.tfvars.json`, plus three annotations recording the pinned execution context: `captf.io/image`, `captf.io/image-digest`, `captf.io/identity`. Its name follows the pattern `captf-inputs-c-` for a `TerraformCluster`, `captf-inputs-m-` for a `TerraformMachine`, and `captf-inputs-mp-` for a `TerraformMachinePool`. To print the tfvars **without** the sensitive `bootstrap_data` key: ```sh kubectl -n get secret captf-inputs-m- \ -o jsonpath='{.data.terraform\.tfvars\.json}' | base64 -d | jq 'del(.bootstrap_data)' ``` That still prints user variables that came from a Secret. To drop them too, delete every key the generated root declares sensitive: ```sh s=$(kubectl -n get secret captf-inputs-m- -o json) sensitive=$(jq -r '.data["main.tf.json"] | @base64d | fromjson | .variable // {} | to_entries | map(select(.value.sensitive == true) | .key) | @json' <<<"$s") jq -r '.data["terraform.tfvars.json"] | @base64d | fromjson' <<<"$s" \ | jq --argjson drop "$sensitive" 'del(.bootstrap_data) | delpaths([$drop[] | [.]])' ``` > [!CAUTION] > > **Both input Secrets carry bootstrap data in cleartext** > > Both the durable Secret (`captf-inputs--`) and the per-run Secret (`captf-run-`) carry bootstrap data in cleartext by design — for control-plane machines that includes the cluster CA and service-account private keys. Treat both Secrets as sensitive: avoid `kubectl get -o yaml` or `-o json` on them, which print every key including `bootstrap_data` unredacted (`kubectl describe secret` is safe: it prints only key names and byte counts, never values). ## 9\. Limits The rendered root (`main.tf.json` plus `terraform.tfvars.json` combined) is capped at 1,000,000 bytes, comfortably under the \~1 MiB a Kubernetes object can hold, with headroom for the Secret’s own metadata (labels, annotations, owner references). > [!WARNING] > > **Oversized inputs do not start a Job** > > Exceeding it sets `ApplyJobSucceeded=False/InputsTooLarge` and does not start a Job — retrying cannot help until the inputs themselves shrink (fewer or smaller `bootstrap_data`, `exports` or user-variable bytes). The size is reported on the `captf_inputs_bytes` gauge (per object, by kind/namespace/name) even when the render was too large to run; see [`metrics.md`](). ## 10\. Not an input - **Cloud credentials.** They come from the `TerraformClusterIdentity`’s mirrored Secret, injected into the Job as environment variables and as files under a read-only mount. They are never rendered into `terraform.tfvars.json`; a module reads them the way its provider expects (environment variables, a credentials file), not as a contract input. See [Identities](). - **Backend configuration.** The `kubernetes` state backend is a partial stub in the generated `main.tf.json`; the runner completes it at `init` time with backend-config flags built from the object’s identity, never from tfvars. See [`runner-cli.md`](). - **The image reference and pull policy.** `spec.source.image` selects which Job runs, and `spec.source.imagePullPolicy` how the kubelet pulls it; both are Kubernetes Pod fields, not Terraform inputs. Only the image’s resulting module code and its declared variables reach the tfvars. - **`captf.io/inputs-hash` and delete guards.** Passed to the runner as command-line flags used to gate a destructive plan and, under `applyPolicy: Manual`, an approved plan hash; see [`runner-cli.md`](). Not part of the module’s own input surface. Nor is `spec.applyPolicy`: it is not hashed, so switching it re-applies nothing. - **Object metadata and controller-written spec fields.** `uid`, labels, annotations, `generation`, and fields the controller writes back (`spec.providerID`, `spec.providerIDList`, an endpoint copied from the module) are never rendered or hashed, so `clusterctl move`, a label edit, or the controller’s own status-driven writes never trigger a re-apply. - **Module-private variables.** A module may declare extra variables of its own only if they carry a default; the generated root never sets them. > [!NOTE] > > **See also** > > - [`contract/v1alpha1/README.md`]() — the normative type and required-ness of every contract input. > - [Module Variables]() — passing your own variables to a module. > - [Machine Pools]() — autoscaling, replicas write-back and membership refresh. > - [Terraform State]() — where the durable Secrets sit relative to the state backend. # v1alpha1 cluster-api-provider-terraform is a Cluster API infrastructure provider that runs Terraform/OpenTofu **child modules written to its contract**, packaged as OCI images together with the runtime. Existing root modules do not drop in: the provider owns the root module, the state backend and the provider installation. > [!WARNING] > > **Read this first: a module is a child module** > > A module written to this contract is a *child module*. It must not declare a `backend`: the controller generates the root module, configures the state backend and supplies the provider installation (the image’s mirror or the registry), passes exactly the contract inputs and reads the contract outputs back from state. **Status:** Provisional: frozen for implementation, and may still change until the first real module has provisioned a cluster (see [`CHANGELOG.md`]()). This contract’s text was checked against the OpenTofu v1.11.5 docs; that is a documentation-verification snapshot, not a runtime requirement — the reference images ([`image-contract.md`]()) ship a newer OpenTofu. A *module* is a Terraform/OpenTofu module that implements one **role** (`cluster`, `machine`, `machinepool`). It is delivered as an **OCI image that bundles the module code and the runtime binary** ([`image-contract.md`]()): the image is the artifact, the version and the runnable object; there is no separate module source and no separate runtime image. The controller never calls a module directly: it generates a root module inside the Job that (a) configures the `kubernetes` backend, (b) calls the module at `/captf/module` with exactly the contract inputs, and (c) re-exports the contract outputs at the root so they land in state. Everything the controller learns about infrastructure comes from those root outputs. Files: - **Image Contract** --- Fixed paths, optional provider mirror, labels, execution environment, pinning, reference Containerfiles. - **Common** --- Inputs injected into every role; the `health` output every role must produce. - **Cluster Role** --- Cluster role (InfraCluster). - **Machine Role** --- Machine role (InfraMachine). - **MachinePool Role** --- MachinePool role (InfraMachinePool). - **Changelog** --- Changes to the `v1alpha1` contract. - **Job Inputs** --- Where every `terraform.tfvars.json` value comes from in the running controller: field-by-field sources, the durable and per-run Secrets, the inputs hash, and how to inspect a live object’s inputs safely. The `schemas/` directory holds JSON Schemas (draft 2020-12) for the cluster, machine and machinepool inputs and outputs, shared definitions, and valid/invalid examples; they are checked by `hack/verify-schemas.sh`. ## Versioning - The contract version is a string: `v1alpha1`. It is independent of the CRD API version and of the provider release. - **The image tag is the module version.** `spec.source.image` (`registry/repo:tag` or `@sha256:…`) is the only version a CR carries; a new module version is a new image, referenced by a new Terraform\*Template. The controller pins the digest after the first run (see [`image-contract.md`]() “Versioning and pinning”). - A module does not declare its contract version or role in a way the controller acts on; there is no manifest file. The role is implied by which CRD references the image, and the contract version is the one the controller release generates against (injected as `captf_contract`, recorded in object status). The OCI labels `io.captf.contract` and `io.captf.role` are informational (see [`image-contract.md`]()). `tfcapi-lint` takes `--role` and `--contract` on the command line and checks the labels against them. - **Compat policy:** - Within one version: only **additive, optional** changes (new optional inputs with defaults; new optional outputs). Never rename or retype. - Making an input required, removing an output, or changing a type means a new contract version. - The controller supports N and N-1 contract versions concurrently; the generated root differs per version. - `v1alpha*` versions carry no compat guarantee between alphas; this policy starts at `v1beta1`. ## Naming rules - **Reserved prefix `captf_`.** Every controller-injected input and every contract-defined local starts with `captf_`. A module MUST NOT declare a variable or output with this prefix except those defined by the contract. - **Contract outputs** are *not* prefixed (`provider_id`, `addresses`, and so on) because they are part of the module’s public interface. Their names are reserved per role; a module MUST declare every required output of its role. - **Module inputs are the contract inputs plus user variables.** Anything module-specific (instance type, machine image, region) is either fixed inside the module or a user variable: an ordinary module variable the object sets through `spec.variables`/`spec.variablesFrom` (see [`common.md`]() “User variables”). A module SHOULD give every non-contract variable a default, so it plans when no user variable is set. - Extra, non-contract outputs are allowed on any role but are never read by the controller. The only cluster-to-machine/pool handoff is the cluster role’s `exports` output (see [`cluster.md`]() and `captf_cluster_outputs` in [`common.md`]()). - **Tagging is mandatory.** Every role receives `captf_tags` (see [`common.md`]()) and MUST apply it to the cloud resources it creates; the CAPI provider best practices require a mechanism for identifying a provider’s cloud objects. ## Trust boundary > [!CAUTION] > > **Referencing an image grants its publisher access** > > `spec.source.image` is executed, unpoliced, as a Job under the runner ServiceAccount, which can read every Secret in the object’s namespace and runs with the identity’s cloud credentials mounted: referencing an image is equivalent to granting its publisher that access. See [Security Model]() for the full threat model; this contract only requires that module authors and operators treat an image reference accordingly. ## Null semantics - Required outputs must be **declared**; their value may be `null` until known. `null` means “not yet known”, never “empty”. - The controller treats `null` for a required output as “keep waiting” (the object stays unprovisioned, condition `OutputsValid=Unknown/OutputsPending`), unless the role spec says otherwise (for example `control_plane_endpoint` on cluster may stay null forever when CAPI supplies the endpoint). - An **absent** output (not declared in the module) is a hard error: `OutputsValid=False/OutputsMissing`, with no retry until the module changes. - Empty collections are meaningful and distinct from null (`addresses = []` means provisioned with no addresses). ## Type conventions - Types are Terraform type constraints. Object attributes listed as optional use `optional(type)` / `optional(type, default)` (Terraform ≥1.3 / OpenTofu; both in scope; see the [Terraform v1.3.0 changelog]() and [OpenTofu docs: type constraints]()). - `output` blocks take no type constraint: their arguments are `value`, `description`, `sensitive`, `ephemeral`, `depends_on`, `deprecated`, plus `precondition` blocks (see [OpenTofu docs: outputs]()). Output “types” in this contract are the shape the controller decodes, and any `optional(..., default)` on an output is applied by the controller, not by Terraform. - All outputs are read from state JSON; the controller decodes by the contract’s expected shape, not by the `type` recorded in state. Extra attributes on objects are ignored; missing required attributes are an error. - Strings that map to Kubernetes fields must satisfy the target field’s validation: `provider_id` 1–512 chars (see [`infra-machine.md`]() “InfraMachine: provider ID”); failure-domain names 1–256 chars, no pattern (see [`cluster_types.go`]() `FailureDomain.Name`); address types from the CAPI `MachineAddressType` enum (see [`common_types.go`]()). The controller rejects invalid values with `OutputsInvalid` rather than writing them. - `sensitive = true` on a module output marks the value sensitive, and sensitivity propagates to anything derived from a sensitive value (for example `bootstrap_data`). A root `output "x" { value = module.role.x }` referencing a sensitive value fails `plan` with “Output refers to sensitive values” unless the root output itself sets `sensitive = true` (reproduced with OpenTofu v1.11.5; see [OpenTofu docs: outputs]() “sensitive”). Sensitive values are still stored in state in cleartext (same page). The generated root marks every re-exported output `"sensitive": true`, so a module may mark any contract output sensitive without failing `plan`. This has no effect on the controller: the state Secret still stores the value in cleartext, and `sensitive` only changes CLI display. ## Injection mechanics (for spec readers) - The generated root lives at `/captf/work/root` inside the Job and calls the module as `module "role" { source = "../../module" }`, that is, `/captf/module` inside the image. Terraform/OpenTofu local paths must begin with `./` or `../`, so the absolute form `/captf/module` is not used. It is a local path, so `init` fetches no module from anywhere (see [`image-contract.md`]() “Fixed paths”). Providers come from `/captf/providers` when the image ships a mirror, else from the registry. - Contract inputs are supplied via `terraform.tfvars.json` in the generated root (auto-loaded; see [OpenTofu docs: variables]()) and forwarded to the module call by name. Bootstrap data is marked `sensitive = true` on the root variable. - Every contract input is a root variable named identically to the module input; the root forwards `x = var.x`. Nothing else is passed to the module call. - Contract outputs are re-exported at root as `output "x" { value = module.role.x, sensitive = true }`: every re-export is marked sensitive, so a module marking any contract output sensitive never fails `plan` (see Type conventions above). - The rendered `terraform.tfvars.json` of the newest apply is kept in a Secret, written when that apply Job starts (a failed apply may already have created resources from those inputs, so destroy must use them), owned by the Terraform\* object (`captf-inputs--`), together with the resolved image digest (`captf.io/image-digest`) and the identity. **Drift and destroy render from that Secret**, never from the live CAPI objects: destroy must work after the bootstrap Secret, the owning Machine/MachinePool, or the Cluster are gone, and drift on an immutable machine must see exactly the inputs it was applied with. The per-run Secret handed to a Job is deleted when the Job completes. The state Secret stores every value (including `bootstrap_data`, sensitivity notwithstanding) in cleartext for as long as the object exists; modules that can avoid keeping bootstrap data in state (for example by hashing it into a trigger) SHOULD. - State lives in the `kubernetes` backend, in the `default` workspace (the only one CAPTF ever uses), locked with a Lease (see [OpenTofu docs: kubernetes backend]() and [Terraform State]() for Secret naming, locks and backups). - **What triggers a re-apply** (mutable roles): a change in the hash of a canonical struct of the *spec-derived* inputs. Rendered inputs are hashed inputs: object metadata (`uid`, labels, annotations, generation) and fields the controller itself writes back (`spec.providerID`, `spec.providerIDList`, `spec.controlPlaneEndpoint`) are neither rendered nor hashed, so a `clusterctl move`, a label edit or the controller’s own status-driven writes never re-apply anything, and a drift `Remediate` (which renders the current hashed set) cannot apply them either (see [`common.md`]() “What is hashed”). ## Roles and CAPI mapping (summary) | Role | CAPI kind | Immutable spec | Drift | Refresh | Destroy on delete | | --- | --- | --- | --- | --- | --- | | cluster | TerraformCluster | no (re-apply on change) | Report (default) / Remediate | with drift; while `pending` 30s, backing off to 5m | yes, after all machines/pools of the Cluster are gone (`DeletionBlocked/DependentsExist` otherwise) | | machine | TerraformMachine | yes (provider choice; the InfraMachine contract only recommends classifying fields as immutable/mutable, [`infra-machine.md`]() “support for in-place changes”) | Report only; unhealthy → `InfrastructureHealthy=False` → `Ready=False` → MHC | with drift; while `pending` 30s, backing off to 5m | yes; direct deletes refused while the owner Machine is not deleting | | machinepool | TerraformMachinePool | no (re-apply on change) | Report (default) / Remediate; `apply -refresh-only` then plan, same as every kind — relies on the module’s `ignore_changes` when `autoscaling.enabled`, see [`machinepool.md`]() | with drift; membership refresh every `spec.membershipRefreshIntervalSeconds` (default 60), and every 30s while `pending` or not yet converged | yes; destroy from the durable inputs Secret, not the bootstrap Secret | Delivery pipeline for a module author: `tfcapi-lint module

--role `, then build the image (see [`image-contract.md`]() for the reference Containerfiles), then push, then `tfcapi-lint image --role `, then reference the tag from a Terraform\*Template. Lint runs on the source directory *before* the build; the image check confirms the layout after it. Conditions: `health` feeds `InfrastructureHealthy`, which is an explicit input to the `Ready` summary. `Deleting` is also in `Ready`’s `ForConditionTypes` for every kind, declared in `NegativePolarityConditionTypes{Deleting}`, so `Ready=False` while a destroy runs (the CAPI v1beta2 convention). Drift results (`DriftDetected`), drift-Job outcomes, `Paused` and `DeletionBlocked` are not, so `Ready` — mirrored by CAPI into the Cluster’s/Machine’s `InfrastructureReady` — reflects infrastructure only. `Ready` is `Unknown` while an object waits on dependencies, `False/Provisioning` from apply start, and `True` when provisioned (see [`machine.md`]() “Ready timeline”). > [!WARNING] > > **A permanently failed destroy needs manual cleanup** > > A destroy that fails permanently is resolved by cleaning up out of band and removing the finalizer by hand, after backing up or un-owning the `tfstate-*` Secrets (they are garbage-collected with the object). There is no skip-destroy annotation. # Common Applies to every role. Inputs here are injected by the controller; modules MUST declare them (they may ignore them). Outputs here MUST be declared by every role. ## Inputs | Name | Type | Required | Value source | | --- | --- | --- | --- | | `captf_contract` | `string` | yes | Literal `"v1alpha1"`; lets a module assert the contract it was generated for | | `captf_cluster` | `object` (below) | yes | The owning CAPI `Cluster` | | `captf_object` | `object` (below) | yes | The Terraform\* object being reconciled | | `captf_cluster_outputs` | `any` | machine, machinepool only | The cluster role’s `exports` output, read from the cluster state (below) | | `captf_tags` | `map(string)` | yes | Fixed tag keys the controller always sets (below) | variables.tf ```hcl variable "captf_contract" { type = string } variable "captf_cluster" { type = object({ name = string namespace = string }) } variable "captf_object" { type = object({ kind = string # TerraformCluster | TerraformMachine | TerraformMachinePool name = string namespace = string }) } variable "captf_cluster_outputs" { type = any default = null } variable "captf_tags" { type = map(string) } ``` ### `captf_cluster_outputs` The value of the cluster role’s `exports` output ([`cluster.md`]()), read from the cluster state. Nothing else from the cluster state is injected. > [!WARNING] > > **A cluster module must not declare it without a default** > > It is **not passed to the cluster role.** Its generated root declares no such variable and its tfvars carry no key, so a cluster module either leaves it undeclared or declares it with `default = null`, as the skeleton below does. A declaration without a default fails `validate`. - **`{}` when the TerraformCluster is externally managed** (the `cluster.x-k8s.io/managed-by` annotation; the provider then never runs the cluster module, so there is no cluster state to read). - A machine/pool whose owning Cluster’s `spec.infrastructureRef` is not kind `TerraformCluster` is never rendered at all: it sets `DependenciesReady=False/ClusterNotTerraform` and does nothing, since there is no cluster state to read `exports` from. - This is how a machine/pool module gets network ids, security groups, subnet ids and similar cluster-wide values. The cluster module puts them in its `exports` output; the controller reads the cluster state and injects exactly that value. `exports` may be marked `sensitive`: the generated root re-exports every contract output with `"sensitive": true` regardless, so a sensitive `exports` never fails `plan`, and sensitivity has no effect on the controller — the state Secret still stores the value in cleartext; only CLI display is affected (see [`README.md`]() “Type conventions”). Machine/pool apply is gated on the cluster being provisioned and `exports` readable, so the value is complete by then. - Because `captf_cluster_outputs` is `any`, a change in its content is **not** treated as a spec change for machines (immutable) but **is** for pools and clusters (a value change triggers a re-apply). See each role’s Lifecycle section. ### `captf_tags` Fixed keys the controller always sets: `captf.io/cluster` (Cluster name), `captf.io/namespace`, `captf.io/kind` (the Terraform\* kind), `captf.io/name` (the Terraform\* object name), `captf.io/managed-by` = `captf`, and `captf.io/template`. - **`captf.io/template`** is the value of the object’s `cluster.x-k8s.io/cloned-from-name` annotation — the template it was cloned from by `external.GenerateTemplate` ([`util.go`](), [`common_types.go`]()) — when present, else `""`. The key is always present: it is read from the Terraform\* object’s own metadata, so an object created directly gets `""`. It is hashed like the other keys even though annotations in general are not: `GenerateTemplate` writes it once at clone time, but for a TerraformCluster or TerraformMachinePool managed by topology, a ClusterClass rebase onto a differently named template rewrites the annotation through SSA, which re-applies that object once; for immutable machines it is pinned at first apply. > [!WARNING] > > **Modules MUST tag every cloud resource** > > Modules MUST apply these tags to every cloud resource they create that supports tags or labels, so resources can be attributed and garbage-collected out of band. This follows the CAPI provider best practices: a tagging/labeling mechanism for identifying cloud objects “MUST always be provided” ([`best-practices.md`](); [`security-guidelines.md`](), housekeeping). ### What is hashed The controller re-applies mutable kinds when the hash of the rendered inputs changes. **Rendered inputs are hashed inputs**: the hash covers a canonical struct of exactly the *spec-derived* inputs the module receives, and nothing is rendered outside it (the one special case is pool `replicas` while `autoscaling.enabled`, rendered as the module’s own observed desired capacity; see [`machinepool.md`]()). `captf_cluster` and `captf_object` therefore carry only kind, name and namespace: - no `uid` — a `clusterctl move` re-creates objects with new UIDs, and hashing them would re-apply every cluster and pool after a move; - no `generation` — it changes whenever the controller itself writes `spec.providerID`, `providerIDList` or `controlPlaneEndpoint`; - no `annotations` — they toggle around every Job (`clusterctl.cluster.x-k8s.io/block-move`) or are edited by GitOps tools; - no `labels` — a metadata edit must never reconcile live infrastructure. Cloud tags come only from `captf_tags`, whose keys are fixed. Drift for mutable kinds renders the current hashed set; immutable machines render drift and destroy from the durable inputs Secret below. ### Durable inputs The rendered inputs (including `bootstrap_data`) are stored in a Secret owned by the Terraform\* object (`captf-inputs--`), so drift and destroy of immutable machines re-feed exactly the values used at apply, and the Secret moves with the object. The resolved image digest (`captf.io/image-digest`, see [`image-contract.md`]() “Versioning and pinning”) and the identity are pinned there too, so an immutable machine’s drift and destroy always run the exact image that created it. ## User variables Everything a module is parameterized by beyond the contract inputs (instance type, disk size, subnets, a database password) is an ordinary Terraform variable of the module. The object’s owner sets it through `spec.variables` (inline JSON) or `spec.variablesFrom` (ConfigMaps and Secrets labeled `captf.io/variables=true`) on the TerraformCluster, TerraformMachine or their templates; see [Module Variables]() for the operator side. For module authors: - **Declare each one as an ordinary variable, with a type and a `default`.** The generated root passes a user variable as a named argument of `module "role"`, declared in the root **without a type**, so your declaration’s type converts the value: a `number` variable accepts `40` inline and `"40"` from a ConfigMap. A variable nobody sets gets its `default`; without one, every object that does not set it fails `plan` on the missing argument. `tfcapi-lint` reports such a variable as `input/user-variable-default` (a warning: a module may deliberately require it). - **Mark secrets `sensitive = true` in the module.** The root declares a variable `sensitive` only when its value came from a Secret; your declaration makes the value sensitive inside the module whatever its source. A value that arrives sensitive stays sensitive in everything derived from it: an output built from it must be declared `sensitive = true` (or `plan` fails with “Output refers to sensitive values”), and it cannot drive `for_each`. Either way the value is stored in cleartext in the inputs Secrets and in the state, like every input (see [Module Variables]()); credentials the module *runs with* come from the identity, never from a variable. - **Names.** User variables are Terraform identifiers (`^[a-zA-Z_][a-zA-Z0-9_-]*$`). Reserved, and rejected before anything is rendered: every `captf_` name, every contract input of the module’s role (`machine_name` on a machine, `control_plane_initialized` on a cluster, and so on) and the module meta-arguments `source`, `version`, `providers`, `count`, `for_each`, `depends_on`, `lifecycle`, `locals`. Do not rely on a user variable to carry a contract input’s value. - **Unknown names fail the apply.** A variable the object sets and the module does not declare is rejected by Terraform/OpenTofu (`Unsupported argument`; for the JSON root, `Extraneous JSON object property`). The controller does not check names against the image. - **Hashing.** User variables are part of the inputs hash only when the object sets some, so modules and objects without them hash exactly as before. For a TerraformCluster a changed variable re-applies (guarded by the destructive-plan approval, see [`cluster.md`]()); a TerraformMachine reads its variables until it is provisioned, and its later runs use the ones pinned in its durable inputs Secret, like every other input. - The role input schemas (`schemas/*-inputs.json`) describe the contract inputs only; the rendered `terraform.tfvars.json` carries the user variables next to them. ## Outputs | Name | Type | Required | Maps to | | --- | --- | --- | --- | | `health` | `object` (below) | yes | `Ready` condition; machine remediation; drift reporting | outputs.tf ```hcl output "health" { value = { state = "running" # closed enum: pending | running | degraded | stopped | terminated | unknown healthy = true message = null # optional string, surfaced in condition message reasons = [] # optional list(string), machine-readable } } ``` > [!WARNING] > > **`state` is a closed enum** > > Any string other than the listed values is `OutputsValid=False/OutputsInvalid`. **Controller-side semantics.** `health` is written to the `InfrastructureHealthy` condition. `InfrastructureHealthy` is one of the explicitly listed inputs to the `Ready` summary; so is `Deleting` (negative polarity, so `Ready=False` while a destroy runs). Drift results, drift-Job outcomes, `Paused` and `DeletionBlocked` are not, so `Ready` — and therefore the Cluster/Machine `InfrastructureReady` mirror that MachineHealthCheck acts on — reflects infrastructure health, not controller housekeeping. | `state` | `healthy` | `InfrastructureHealthy` | Effect | | --- | --- | --- | --- | | `pending` | any | `False/InstancePending` | Provisioning in progress; object not marked provisioned even if other outputs are set. While pending, the controller refreshes (`apply -refresh-only`) regardless of `drift.intervalSeconds`, so provisioning never waits for a drift tick: for cluster and machine, 30s after the last refresh, then 1m, 2m, 4m and at most every 5m while the readings stay pending (plus the object’s jitter); for a machinepool, every 30s flat, since a pending pool is also not converged (see [`machinepool.md`]() “Membership refresh”) | | `running` | `true` | `True/Healthy` | `Ready=True` once role outputs are satisfied | | `running` | `false` | `False/InstanceUnhealthy` | machine: remediation candidate | | `degraded` | any | `False/InstanceDegraded` | machine: remediation candidate | | `stopped` | any | `False/InstanceStopped` | machine: remediation candidate | | `terminated` | any | `False/InstanceTerminated` | machine: remediation candidate; object stays provisioned and `spec.providerID` is kept | | `unknown` | any | `Unknown/HealthUnknown` | `Ready=Unknown`; no remediation | Once provisioned, the Machine controller treats an empty InfraMachine `spec.providerID` as “waiting” and logs it; it is not a hard error, but clearing it would stall the Machine ([`machine_controller_phases.go`]() `reconcileInfrastructure`), so a `terminated` reading keeps `providerID` rather than clearing it. - `health` is re-read on every refresh (`apply -refresh-only`, which updates state and root output values to match remote objects; see [OpenTofu docs: `cli/commands/plan/`]() “Planning Modes”): the pending refresh loop above, every drift tick, and for pools the membership refresh (see [`machinepool.md`]()). For immutable machines that is the only way it changes after apply. Modules that cannot observe health return `{state="running", healthy=true}` (the no-op module does this) or `unknown`. - **Out-of-band termination.** If, after provisioning, a refresh makes `provider_id` come back `null` or `""` (the resource vanished from state), the first such sample only sets `InfrastructureHealthy=Unknown/ProviderIDMissing`, since a transient read can return no value. A second consecutive missing sample is treated as `terminated`: `InfrastructureHealthy=False/InstanceTerminated`, `spec.providerID` kept, and with `remediation.annotateMachine` the Machine is annotated for remediation. A refresh that itself fails because the instance is gone is treated as `terminated`. Modules SHOULD still report `terminated` explicitly where the cloud API exposes it, because detection through a vanished resource is slower (drift interval plus a Job) and CAPI’s Node-based checks will usually fire first. - A group with zero members (pool role, `replicas == 0`) reports `{state="running", healthy=true}`. - Only the machine role triggers remediation. Cluster/pool `health` feeds conditions only. ## Role-specific inputs `kubernetes_version` is a role input, not a common one: for machines/pools it comes from `Machine.spec.version` / `MachinePool.spec.template.spec.version` (authoritative for nodes); for the cluster role it comes from `Cluster.spec.topology.version` (null without ClusterClass; see [`cluster.md`]()). `cluster_network`, `control_plane_initialized` and the non-module `control_plane_endpoint` (including one set later by a control-plane provider) are cluster-role inputs (see [`cluster.md`]()); machine/pool modules get endpoint-derived values through `exports`. # Cluster Role Implements the CAPI InfraCluster contract ([`infra-cluster.md`]()) for a `TerraformCluster`. One workspace/state per `TerraformCluster`. Typically owns the cluster-wide substrate: network, subnets, security groups, load balancer / control-plane endpoint, shared images/keys. Values that machine/pool modules need are handed over through the `exports` output, injected into them as `captf_cluster_outputs`. ## Inputs Common inputs apply ([`common.md`]()), except `captf_cluster_outputs`: it is not passed to this role. The skeleton below declares it with `default = null`, which is safe; a declaration without a default fails `validate`. | Name | Type | Required | Value source | | --- | --- | --- | --- | | `control_plane_endpoint` | `object({host=string, port=number})` or `null` | yes (may be null) | Any endpoint the module does **not** own (below) | | `kubernetes_version` | `string` or `null` | yes (may be null) | `Cluster.spec.topology.version`; `null` without ClusterClass (below) | | `control_plane_initialized` | `bool` | yes | `Cluster.status.initialization.controlPlaneInitialized`, latched (below) | | `cluster_network` | `object({pods=list(string), services=list(string), service_domain=string, api_server_port=number})` or `null` | yes (may be null) | `Cluster.spec.clusterNetwork` (below) | ### `control_plane_endpoint` (input) Value = `Cluster.spec.controlPlaneEndpoint` whenever the annotation `captf.io/endpoint-source` is not `module` — whether the user set it before the first apply or a control-plane provider (hosted control planes) sets it later ([`cluster_controller_phases.go`]()). While the Cluster field is not yet valid (`APIEndpoint.IsValid`, host and port both set), the input falls back to a valid `TerraformCluster.spec.controlPlaneEndpoint` (the two are equal once CAPI copies ours, [`cluster_controller_phases.go`](), so the copy never changes the hash); else `null`. It is always `null` when `endpoint-source=module`. Non-null means “use this endpoint, don’t create one”: the contract allows the endpoint to come from the user, the control plane provider, or the infra provider ([`infra-cluster.md`]() “InfraCluster: control plane endpoint”). Shape = `APIEndpoint{host 1–512 chars, port 1–65535}` ([`cluster_types.go`]()). **Provenance** (`captf.io/endpoint-source`): `user` is written before the first apply when either field holds a valid endpoint; `module` is written when the controller first copies a non-null module output (from an apply whose rendered input was null) into `TerraformCluster.spec.controlPlaneEndpoint`; otherwise the annotation is absent. Once written it never changes. With `module` the input stays `null` for the life of the object even after CAPI copies the module’s endpoint to `Cluster.spec`: feeding it back would tell the module to drop the load balancer it created, and would change the inputs hash. In the inputs hash: an endpoint set later by a control-plane provider therefore re-applies the cluster once, which is how the module learns it. ### `kubernetes_version` (input) `Cluster.spec.topology.version` ([`cluster_types.go`]()); `null` when the Cluster has no `topology` (no ClusterClass). For modules that provision managed/hosted control planes or version-dependent cluster resources. Passed verbatim (it may carry a distro suffix such as `+rke2r1`; strip it as the machine role does). In the inputs hash, so a topology upgrade re-applies the cluster. ### `control_plane_initialized` (input) `Cluster.status.initialization.controlPlaneInitialized` ([`cluster_types.go`](); nil maps to `false`). In the inputs hash, so the cluster re-applies exactly once when the control plane comes up (CAPI latches the field). The controller latches it too: the rendered value is `true` when the Cluster field is true **or** the last rendered value in the durable inputs Secret (`captf-inputs-c-`, which moves) is true; it is never rendered `false` after an apply with `true`. Cluster status is not moved by `clusterctl move`, so without this a first reconcile on the target could render `false` and destroy every gated resource. Modules that need a live workload API server (in-cluster add-ons, cloud-controller secrets, DNS records pointing at registered nodes) gate those resources on it, for example `count = var.control_plane_initialized ? 1 : 0`. ### `cluster_network` (input) `Cluster.spec.clusterNetwork` (`pods.cidrBlocks`, `services.cidrBlocks`, `serviceDomain`, `apiServerPort`; [`cluster_types.go`]() `ClusterNetwork`, `NetworkRanges`); `null` when the Cluster sets none of them; missing attributes are null or `[]`. > [!WARNING] > > **`api_server_port` is not the kube-apiserver bind port with KCP** > > CABPK/KCP never read `Cluster.spec.clusterNetwork.apiServerPort`; the real bind port is KubeadmConfig `localAPIEndpoint.bindPort` (default 6443, [`workload_cluster.go`]()). By convention the module’s LB backend port is `cluster_network.api_server_port ?? 6443`, and templates must keep `apiServerPort` and `bindPort` equal — the controller has no way to enforce this. ## Outputs Common outputs apply (`health`). | Name | Type | Required | Maps to | | --- | --- | --- | --- | | `control_plane_endpoint` | `object({host=string, port=number})` or `null` | yes (may be null) | `TerraformCluster.spec.controlPlaneEndpoint` → `Cluster.spec.controlPlaneEndpoint` (below) | | `failure_domains` | `list(object({name=string, control_plane=optional(bool, true), attributes=optional(map(string), {})}))` or `null` | yes (may be null or `[]`) | `TerraformCluster.status.failureDomains` (below) | | `exports` | `any` (object recommended) | yes (may be null, treated as `{}`) | Injected verbatim into machine/pool modules as `captf_cluster_outputs` (below) | ### `control_plane_endpoint` (output) Maps to `TerraformCluster.spec.controlPlaneEndpoint` → `Cluster.spec.controlPlaneEndpoint`, surfaced once InfraCluster initialization completes ([`infra-cluster.md`]()). It is written only when all of the following hold: the output is non-null; the field is still empty; and the apply’s rendered `control_plane_endpoint` input was null (the first such write records `captf.io/endpoint-source=module`). It is written only when the output is **valid**: CAPI’s copy-to-`Cluster.spec` path only fires once `APIEndpoint.IsValid()` is true, which requires both `host` and `port` to be set ([`cluster_types.go`]()), so a half-set output (only one of the two present) is treated as `OutputsValid=False/OutputsInvalid`, not partially written. CAPI copies the InfraCluster endpoint only while `Cluster.spec.controlPlaneEndpoint` is not yet valid, so later changes are never propagated ([`cluster_controller_phases.go`]() `reconcileInfrastructure`); CAPI’s Cluster admission webhook has no immutability rule for it ([`cluster.go`]()). **Once emitted, the value MUST be stable for the life of the object.** CAPI never updates `Cluster.spec.controlPlaneEndpoint` after the first valid copy ([`cluster_controller_phases.go`]()), so a module that later replaces its load balancer (a new DNS name or IP) breaks every existing kubeconfig and the apiserver certificate SANs, with no way for CAPI to notice or recover. Keeping it stable is the module author’s responsibility first: the controller blocks, but cannot fix, a planned replacement. Every TerraformCluster apply (a re-apply, a new image tag, a drift remediation) stops before a plan that deletes or replaces any resource, sets `ApplyJobSucceeded=False/DestructivePlanBlocked` with the affected addresses, and applies it only once the `captf.io/approve-destructive-plan` annotation names that apply’s inputs hash; see [Plan Approval]() for the destructive-plan guard. With `spec.applyPolicy: Manual` every re-apply instead waits for the approval of its plan, which covers its deletes; see [Plan Approval]() for plan preview. An operator can still approve a plan that replaces the endpoint’s resource; `lifecycle { prevent_destroy = true }` on the resource behind the endpoint turns such a plan into a failed Job even when approved, instead of a lost cluster. When no user endpoint exists, the module MUST emit this output no later than the apply that makes the cluster `provisioned` (see `EndpointAvailable` below). ### `failure_domains` (output) Maps to `TerraformCluster.status.failureDomains` (v1beta2 list shape: `listType=map`, `listMapKey=name`, `MaxItems=100`; `FailureDomain{name 1–256 chars, controlPlane *bool, attributes map[string]string}`; [`infra-cluster.md`]() “InfraCluster: failure domains”, [`cluster_types.go`]()). CAPI’s `controlPlane` has no default; `optional(bool, true)` is our default, applied by the controller. `null` and `[]` both clear the field. ### `exports` (output) Injected verbatim into machine/pool modules as `captf_cluster_outputs`. May be marked `sensitive`: the generated root re-exports it `"sensitive": true` regardless, so this has no effect on the controller (see [`README.md`]() “Type conventions”). Put secrets in the identity, not here. Convention: publish the load balancer target/backend-pool id(s) here for control-plane machine modules to register against (see [`machine.md`]() “Control-plane machines”). ### Provisioned rule `status.initialization.provisioned` is the v1beta2 contract field: `status.initialization.provisioned = true` when the state Secret carries the inputs-hash annotation of a successful apply, the required outputs are valid, and `health.state != "pending"`. **Latched.** The formula applies only until it first holds; once `provisioned` is true it stays true for the object’s life, whatever later health or outputs say (CAPI never flips `infrastructureProvisioned` back; [`machine_controller_phases.go`]() “should not flip back”). There is **no endpoint gate on `provisioned`**: a null `control_plane_endpoint` output with no endpoint on the Cluster is allowed, because hosted-control-plane providers set the endpoint only after infrastructure is provisioned. Provisioned is derived from state, never from Job history, so it is rebuilt identically after `clusterctl move` (status is not moved; the target evaluates the formula once, then latches). ### Hairpin reachability > [!WARNING] > > **The endpoint MUST be reachable from the first control-plane node** > > The control-plane endpoint MUST be reachable from the first control-plane node itself, not only from the management cluster and other nodes. kubeadm’s `getKubeConfigSpecsBase` (kubernetes/kubernetes release-1.33, `cmd/kubeadm/app/phases/kubeconfig/kubeconfig.go` around lines 582-626) sets the server URL of `admin.conf`, `super-admin.conf` and `kubelet.conf` to `GetControlPlaneEndpoint(cfg.ControlPlaneEndpoint, &cfg.LocalAPIEndpoint)` — that is, the endpoint, not the node’s local address — while only `controller-manager.conf`/`scheduler.conf` use `GetLocalAPIEndpoint`. Because kubeadm’s post-init phases and the first node’s own kubelet talk to the workload cluster through `admin.conf`/`kubelet.conf`, the load balancer or VIP that backs `control_plane_endpoint` must hairpin: the control-plane instance that is itself a backend must be able to reach the frontend it was just registered behind. ### `EndpointAvailable` condition Cluster only, not part of `Ready`: `False/WaitingForEndpoint` once `provisioned` is true and neither the module’s `control_plane_endpoint` output nor `Cluster.spec.controlPlaneEndpoint` is a valid `APIEndpoint` (host and port both set); `True/EndpointAvailable` otherwise. > [!WARNING] > > **A null endpoint output silently stalls the cluster** > > With KCP this case is not cosmetic: KCP creates no Machine at all until `Cluster.spec.controlPlaneEndpoint.IsValid()` ([`kubeadmcontrolplane_controller.go`]() — “Waiting for Cluster spec.controlPlaneEndpoint to be set”), so a module that outputs `null` with no user-provided endpoint silently stalls the cluster forever with no other signal. `tfcapi-lint` warns (`output/endpoint-never-set`) when a module’s `control_plane_endpoint` output is a literal `null` with no conditional path to a value. ## Lifecycle - **Mutable.** Any change to `spec.source.image` (a pull-policy change alone only affects how the next Job pulls, and starts none), a non-module `control_plane_endpoint` (including one set later by a control-plane provider), `cluster_network`, `kubernetes_version` (`Cluster.spec.topology.version`), or `control_plane_initialized` flipping true starts a new apply Job (detected by the hash of the rendered spec-derived inputs; labels, uid, and controller-written fields are neither rendered nor hashed, see [`common.md`]() “What is hashed”). - **Externally managed.** A TerraformCluster carrying `cluster.x-k8s.io/managed-by` is skipped entirely: no Jobs, no status writes ([`infra-cluster.md`]() “Externally managed infrastructure”). Machines and pools of that Cluster receive `captf_cluster_outputs = {}`. - **Drift.** On `spec.drift.intervalSeconds`: `apply -refresh-only` first (so outputs, health and `exports` refresh and the plan compares against current remote values), then `plan -detailed-exitcode` (0 = no changes, 1 = error, 2 = changes; see [OpenTofu docs: `cli/commands/plan/`]()). Exit 2 sets `DriftDetected=True`; with `action: Report` (the default) that is all, and an operator decides; with `action: Remediate` the controller applies, re-rendering the *current* inputs. A changed `exports` value re-applies pools (mutable) but not machines (immutable). Drift results and drift-Job outcomes never feed `Ready`, and after first provisioning neither does a failed re-apply: it is `ApplyJobSucceeded=False` plus `DriftDetected`, never `Ready=False`, because Cluster `InfrastructureReady=False` would suspend every MachineHealthCheck of the cluster ([`machinehealthcheck_targets.go`]()). - **Delete ordering (enforced by the controller).** Destroy is blocked with `DeletionBlocked=True/DependentsExist` (requeue) while any TerraformMachine or TerraformMachinePool labeled `cluster.x-k8s.io/cluster-name=` (`ClusterNameLabel`, [`common_types.go`]()) exists in the namespace, since their modules depend on cluster resources via `exports`. CAPI’s own Cluster deletion already removes Machines/MachinePools first — it requeues while descendants, including MachinePools, exist, then deletes the control plane object, then the InfraCluster ([`cluster_controller.go`]() `reconcileDelete`, `clusterDescendants`) — so this guard only covers out-of-band deletions of the TerraformCluster. Then a `destroy` Job runs; success drops the state Secret(s) and Lease, and removes the finalizer. Module authors may rely on this ordering. - The module MUST be idempotent under repeated apply with unchanged inputs (plan must be empty). `tfcapi-lint` can’t check this, and no automated test does yet; the module author must. ## Control-plane provider requirements What a control-plane provider needs from the cluster module, and how KubeadmControlPlane (KCP) and RKE2ControlPlane (RCP) differ. The KCP column draws on [the KubeadmControlPlane integration guide](); the RCP column draws on [the RKE2ControlPlane integration guide](). | Requirement | KubeadmControlPlane | RKE2ControlPlane | | --- | --- | --- | | Endpoint required before any Machine | Yes: KCP creates no Machine until `Cluster.spec.controlPlaneEndpoint.IsValid()` ([`kubeadmcontrolplane_controller.go`](); see the KubeadmControlPlane guide) | Yes: RCP returns before creating *any* Machine while `!IsValid()` ([`rke2controlplane_controller.go`](); see the RKE2ControlPlane guide) | | LB listeners | `endpoint.port` → CP nodes `:bindPort` (default 6443); one frontend | `endpoint.port` → CP nodes `:6443`, **and** `:9345` → CP nodes `:9345` on the **same host** as the endpoint (the join URL is `https://:9345`, port ignored); two frontends | | Health checks | TCP on the backend port, or HTTPS `/readyz`/`/healthz` without certificate verification; must go green with a single backend | 9345: TLS `GET /v1-rke2/readyz` expecting **403** (unauthenticated), or plain TCP; 6443: TCP, or HTTPS `/healthz` only if anonymous auth is enabled | | Backend membership | Module registers/deregisters CP instances in its own state; instance MUST be in the LB backend before `kubeadm init`/`join` finishes, or `controlPlaneInitialized` never latches ([`infra-machine.md`]()) | Same registration contract; module MUST add a CP node before its Machine is Ready (joins through the LB during bring-up) and MUST tolerate a backend whose supervisor is not yet up | | SG / firewall ports | Backend port (bindPort) from LB, nodes and management; 2379–2380 and 10250 between CP nodes; 10250 from CP to workers | CP↔CP 2379-2381; all→CP 6443 and 9345; all↔all 10250, NodePort range; CNI ports for the chosen `serverConfig.cni` (default canal: 8472/udp, 9099); egress to `get.rke2.io`/GitHub releases unless air-gapped | | `cluster_network.api_server_port` / `service_domain` | `api_server_port` is **not** wired to the kube-apiserver bind port by CABPK/KCP either — the real bind port is KubeadmConfig `localAPIEndpoint.bindPort` (default 6443); backend port is `api_server_port ?? 6443` by convention only | **No effect** under RKE2: the domain comes from `serverConfig.clusterDomain`, and the API server is always 6443 on nodes; RKE2 never reads either field | | `failure_domains[].control_plane` default | KCP filters to entries with `controlPlane == true`; a nil `controlPlane` counts as false, so our `optional(bool, true)` default is required for domains to be visible | Same: RCP only uses entries with `controlPlane: true` (nil counts as false); our default `true` is compatible with both providers | See the [KubeadmControlPlane]() and [RKE2ControlPlane]() guides for the reasoning and exact citations behind each row. ## Minimal skeleton A block written on one line may hold at most one argument ([`OneLineBlock`](); `tofu validate` reports “Invalid single-argument block definition” on a violation), so a variable that needs both `type` and `default` (like `captf_cluster_outputs` below) must use the multi-line block form: main.tf ```hcl variable "captf_contract" { type = string } variable "captf_cluster" { type = object({ name = string, namespace = string }) } variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) } variable "captf_cluster_outputs" { type = any default = null } variable "captf_tags" { type = map(string) } variable "control_plane_endpoint" { type = object({ host = string, port = number }) default = null } variable "kubernetes_version" { type = string default = null } variable "control_plane_initialized" { type = bool } variable "cluster_network" { # Every attribute is always present when cluster_network is non-null # (an unset CIDR list renders as [], an unset scalar as null), so the # type matches the Inputs table exactly: none of the four is optional(). type = object({ pods = list(string) services = list(string) service_domain = string api_server_port = number }) default = null } # Tag every cloud resource this module creates with captf_tags (common.md # "captf_tags"); this stub only has to reference it, not create anything. resource "terraform_data" "tags" { input = var.captf_tags } output "control_plane_endpoint" { value = var.control_plane_endpoint != null ? var.control_plane_endpoint : { host = "...", port = 6443 } } output "failure_domains" { value = [] } output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } } # Handed to machine/pool modules as captf_cluster_outputs: output "exports" { value = { network_id = "...", subnet_ids = ["..."] } } ``` # Machine Role Implements the CAPI InfraMachine contract ([`infra-machine.md`]()) for a `TerraformMachine`. One workspace/state per `TerraformMachine`. The module creates exactly one node instance and reports its provider ID, addresses, and health. The machine is immutable (`spec.source` and `spec.identityRef` cannot change): apply once, refresh/drift for health, destroy on delete. ## Inputs Common inputs apply ([`common.md`]()), including `captf_cluster_outputs`. | Name | Type | Required | Value source | | --- | --- | --- | --- | | `machine_name` | `string` | yes | Owning CAPI `Machine.metadata.name` (the TerraformMachine name may differ) | | `bootstrap_data` | `string`, `sensitive` | yes | Base64 of the bootstrap Secret’s `value` key (below) | | `bootstrap_format` | `string` | yes | `cloud-config` or `ignition`, from the bootstrap Secret’s `format` key (below) | | `failure_domain` | `string` or `null` | yes (may be null) | `Machine.spec.failureDomain` (below) | | `kubernetes_version` | `string` or `null` | yes (may be null) | `Machine.spec.version` (below) | | `control_plane` | `bool` | yes | `true` when the owning Machine carries the control-plane label (below) | ### `bootstrap_data` **Base64** (standard alphabet, padded) of the raw bytes of the `value` key of the Secret named by `Machine.spec.bootstrap.dataSecretName`. The bootstrap contract requires a single key `value` ([`bootstrap-config.md`]() “BootstrapConfig: data secret”; `dataSecretName` in [`machine_types.go`]() `Bootstrap`). The controller base64-encodes `value` when rendering tfvars, **always and for every bootstrap provider**: CAPRKE2 with `gzipUserData: true` writes `value` as raw gzip bytes, which are not valid UTF-8 and so cannot be a JSON/HCL `string` (see [the RKE2ControlPlane guide]()); encoding unconditionally keeps one rule instead of sniffing content. The input stays `sensitive = true`. Modules use one of two patterns: - pass it unchanged to a `user_data_base64`-style argument (for example AWS `aws_instance.user_data_base64` or `aws_launch_template.user_data`, which both take base64); or - `base64decode(var.bootstrap_data)` where the argument wants the plain payload — only valid when the payload is UTF-8, that is not gzipped: `base64decode` fails on non-UTF-8 bytes, so a binary payload must go to a base64-taking argument instead. Apply waits until the Secret exists (the contract workflow also exits while `dataSecretName` is nil; [`infra-machine.md`]() “Typical InfraMachine reconciliation workflow”); `dataSecretName: ""` (the field allows `MinLength=0`, [`machine_types.go`]()) is treated exactly like nil (`WaitingForBootstrapData`). > [!CAUTION] > > **Control-plane bootstrap data embeds cluster private keys** > > For control-plane Machines the payload embeds the cluster CA and service account private keys **uncompressed** ([`controlplane_init.go`](), [`controlplane_join.go`]()), so it is both larger and more sensitive than a worker’s. Size limits (for example AWS’s 16 KiB user-data) and keeping keys out of readable instance metadata are module business: gzip the payload, or stage it in a secret store and pass only a small stub through `bootstrap_data`/user-data. ### `bootstrap_format` `cloud-config` or `ignition`, from the bootstrap Secret’s `format` key; `cloud-config` when absent. `format` is not part of the bootstrap contract, which specifies only `value` ([`bootstrap-config.md`]()). It is written by the kubeadm bootstrap provider ([`kubeadmconfig_controller.go`]() writes `value` and `format`; enum `cloud-config;ignition` in [`kubeadmconfig_types.go`]() `Format`). Other bootstrap providers need not write it. CAPRKE2 also **always** writes both keys, with the same enum ([`rke2config_controller.go`](); see [the RKE2ControlPlane guide]()). `format` describes the decoded payload; with CAPRKE2 `gzipUserData: true` the decoded bytes are gzip of that format. ### `failure_domain` (input) `Machine.spec.failureDomain` (1–256 chars; [`machine_types.go`]()). The InfraMachine MUST be placed in this failure domain ([`infra-machine.md`]() “InfraMachine: failure domain”). ### `kubernetes_version` (input) `Machine.spec.version` (optional, 1–256 chars; [`machine_types.go`]()). The value may carry a control-plane-provider-specific distro suffix: with RKE2 it is `vX.Y.Z+rke2rN` (for example `v1.31.4+rke2r1`), copied verbatim from `RKE2ControlPlane.spec.version` to `Machine.spec.version` (see [the RKE2ControlPlane guide]()). The controller passes this input to the module **verbatim**, unstripped; a module that uses it for an image lookup or a semver comparison MUST strip the `+…` build-metadata suffix itself (see [the RKE2ControlPlane guide]()). ### `control_plane` (input) `true` when the owning Machine carries the `cluster.x-k8s.io/control-plane` label (`MachineControlPlaneLabel`, [`machine_types.go`](); CAPI’s own `util.IsControlPlaneMachine` checks only for this label’s presence, [`util.go`]()). There is no role string in `MachineSpec`. ## Outputs Common outputs apply (`health`). | Name | Type | Required | Maps to | | --- | --- | --- | --- | | `provider_id` | `string` or `null` | yes (may be null until known) | `TerraformMachine.spec.providerID` → `Machine.spec.providerID` (below) | | `addresses` | `list(object({type=string, address=string}))` | yes (may be `[]`) | `TerraformMachine.status.addresses` → `Machine.status.addresses` (below) | | `failure_domain` | `string` or `null` | yes (may be null) | `TerraformMachine.status.failureDomain` → `Machine.status.failureDomain` (below) | | `interruptible` | `bool` or `null` | declared by every machine module; value optional (null = `false`) | `TerraformMachine.status.interruptible` (below) | ### `provider_id` (output) Maps to `TerraformMachine.spec.providerID` → `Machine.spec.providerID`. Must equal the Node’s `spec.providerID` exactly (see “Node providerID matching” below). 1–512 chars ([`infra-machine.md`]() “InfraMachine: provider ID”; `MinLength=1`, `MaxLength=512` markers in [`machine_types.go`]()). `""` is treated as `null`. Written once; a later non-null change is `OutputsValid=False/ProviderIDChanged`. If it turns `null` after provisioning, one sample is not trusted: the first reports `InfrastructureHealthy=Unknown/ProviderIDMissing`, and only a second consecutive missing sample treats the instance as terminated (see [`common.md`]() “Out-of-band termination”). Either way `spec.providerID` is kept. We latch it — never clear it once set — because CAPI does not: the Machine controller recopies `Machine.spec.providerID` from `TerraformMachine.spec.providerID` on every reconcile rather than latching it once ([`machine_controller_phases.go`]()), and clearing our side would blank the Machine’s field and stop address/failure-domain copying from the InfraMachine on that same path. ### `addresses` (output) Maps to `TerraformMachine.status.addresses` → `Machine.status.addresses` ([`infra-machine.md`]() “InfraMachine: addresses”). `type` is one of `Hostname`, `ExternalIP`, `InternalIP`, `ExternalDNS`, `InternalDNS` (`MachineAddressType` enum, [`common_types.go`]()); `address` is 1–256 chars; at most 256 entries (`MachineAddress` and `Machine.status.addresses` markers, [`common_types.go`]() and [`machine_types.go`]()). The controller validates every mapped output against these CRD markers before patching status; a violation is `OutputsValid=False/OutputsInvalid`, and nothing is written. **Canonical order.** Before writing status the controller sorts the module’s list by type precedence `InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname`, then by address. CAPI itself copies the list verbatim ([`machine_controller_phases.go`]()), so an unsorted module output would otherwise reach `Machine.status.addresses` in whatever order the module emitted it. Rationale: RKE2’s `registrationMethod: internal-first` takes the **first** address in list order whose type is `InternalIP` or `ExternalIP` (this is what the code does, not “internal preferred” as its own docs claim); `internal-only-ips` needs an `InternalIP` present, and `external-only-ips` needs an `ExternalIP` present (see [the RKE2ControlPlane guide]()) — a fixed, sorted order makes that first-match deterministic instead of depending on module authoring order. ### `failure_domain` (output) Maps to `TerraformMachine.status.failureDomain` → `Machine.status.failureDomain` (the actual placement; [`infra-machine.md`]() “InfraMachine: failure domain”). When the `failure_domain` input is non-null the output MUST equal it: the contract says the InfraMachine MUST be placed in the requested domain, and the controller enforces it — a mismatch is `OutputsValid=False/FailureDomainMismatch` and `Ready=False`, so KCP/MachineDeployment spreading can never be silently broken. The output adds information only when no domain was requested. ### `interruptible` (output) Maps to `TerraformMachine.status.interruptible` (`bool`); `true` for spot/preemptible instances. CAPI reads `status.interruptible` from the InfraMachine and, when true, sets the Node label `cluster.x-k8s.io/interruptible` (`InterruptibleLabel`; [`machine_controller_noderef.go`]()), which termination handlers and schedulers key on. The output must be **declared** because the generated root re-exports every contract output unconditionally; a module without spot support declares `output "interruptible" { value = false }`. Null is written as `false`. ### Provisioned rule `status.initialization.provisioned = true` when the state Secret carries the inputs-hash annotation of a successful apply, `provider_id` is non-null, and `health.state != "pending"`; derived from state only, never from Job history, so it is rebuilt after `clusterctl move` (evaluated once on the target, then latched). **Latched.** The formula applies only until it first holds; once `provisioned` is true it stays true for the object’s life, whatever later health or outputs say (CAPI never flips `infrastructureProvisioned` back; [`machine_controller_phases.go`]() “should not flip back”). A later null `provider_id` or non-running health is reported on `InfrastructureHealthy` (for example `InstanceTerminated`), never by un-provisioning. While `pending`, the controller refreshes regardless of `drift.intervalSeconds`, 30s after the last refresh and backing off to at most 5m (not a field; see [`common.md`]()). After a successful apply it refreshes once more unless the apply’s own outputs are valid with a `health.state` other than `pending` or `unknown`. `Ready=True` additionally requires `health.healthy` and `state == "running"`. ### Ready timeline What CAPI mirrors into `Machine.status.conditions[InfrastructureReady]`, which MachineHealthCheck can act on: | Phase | `Ready` | Reason | | --- | --- | --- | | Waiting for owner Machine, cluster infrastructure, `exports`, bootstrap data | `Unknown` | `WaitingFor*` (on the `DependenciesReady` condition; `Ready` reason `ReadyUnknown`) | | Apply Job started | `False` | `Provisioning` | | Provisioned, healthy | `True` | `Ready` | | Provisioned, `health` not running/healthy | `False` | `NotReady` (detail on `InfrastructureHealthy`) | | Deleting | `False` | `Deleting` | > [!NOTE] > > **`Unknown` rather than `False` while waiting matters for MHC** > > `unhealthyMachineConditions` timeouts count from the condition’s `lastTransitionTime`, and a reason change does not reset it. Workers created alongside a slow control plane can wait a long time for bootstrap data; starting the `False` window only at apply start means an MHC `InfrastructureReady=False` timeout measures apply time, not control-plane bring-up. After provisioning, `Ready` is computed from `InfrastructureHealthy` and deletion only; drift results, drift-Job failures, identity or RBAC problems never turn it `False` (they are separate conditions), so a credentials outage cannot trigger fleet-wide remediation. ## Lifecycle - **Immutable.** Changes to `spec.source` and `spec.identityRef` after creation are rejected by the webhook, and `spec.providerID` can only be set once, by the controller. Those changes come via MachineDeployment/KCP rollout. `spec.jobs`, `spec.drift` and `spec.remediation` are operational policy and stay mutable: they change how the controller treats the machine (a Job deadline, drift checks, auto-remediation), never what was applied. `captf_cluster_outputs` changes do **not** trigger re-apply: the rendered inputs of the first apply (including `bootstrap_data` and `captf_cluster_outputs`) are stored in the object-owned `captf-inputs--` Secret and re-fed unchanged on drift and destroy; the effective image and identity are pinned there too, so a later change to `TerraformCluster.spec.defaults` (for example a ClusterClass in-place update) never changes what runs against an existing machine. Bootstrap data is read once. The kubeadm bootstrap provider does not rewrite a Machine’s bootstrap Secret after creation — it only extends the join token’s TTL in the workload cluster until the Node joins ([`kubeadmconfig_controller.go`]() `refreshBootstrapTokenIfNeeded`) — so “read once” loses nothing. Pools are different (see [`machinepool.md`]()). - **Gating.** Apply waits for: the owner `Machine` (looked up through the Machine ownerRef, `util.GetOwnerMachine`, then the Cluster through `cluster.x-k8s.io/cluster-name`; a TerraformMachine never has a Cluster ownerRef); `Cluster.status.initialization.infrastructureProvisioned` ([`cluster_types.go`]()); the cluster’s `exports` readable (`WaitingForClusterExports`); and the bootstrap Secret (`WaitingForBootstrapData`; the contract workflow also exits while `dataSecretName` is nil, and `""` counts as nil). The owning Cluster’s `spec.infrastructureRef` must be kind `TerraformCluster`; otherwise `DependenciesReady=False/ClusterNotTerraform` and nothing is done. Once those pass, the `spec.variablesFrom` sources must exist and be labeled (`DependenciesReady=False/VariablesSourceNotFound`) with valid keys and values (`False/VariablesInvalid`); see [`common.md`]() “User variables”. While gated, `Ready=Unknown` (timeline above). - **Retry.** A state Secret that exists without the inputs-hash annotation (a failed or interrupted first apply) means the apply is retried with backoff even though the spec is immutable; each attempt gets a distinct Job name (attempt counter). - **Refresh/drift.** On `spec.drift.intervalSeconds`: `apply -refresh-only` then `plan -detailed-exitcode`. Drift is **reported only** (`DriftDetected=True`, a negative-polarity condition kept out of `Ready`); never auto-applied for machines (`MachineDriftPolicy` has no `action` field: there is nothing to remediate against for an immutable machine). The refresh updates `addresses` and `health`. A failed drift Job sets `DriftJobSucceeded=False` and nothing else. - **Health → remediation.** `health` outside `running`/`healthy` after provisioning sets `InfrastructureHealthy=False`, then `Ready=False/NotReady`, mirrored to `Machine.status.conditions[InfrastructureReady]=False`, which a `MachineHealthCheck` can act on. See [Machine Remediation]() for configuring `spec.remediation`, the unhealthy threshold and sampling interval, and how a MachineHealthCheck reaches a replacement. `health` is purely cloud-side: the controller does not cross-check `Machine.status.nodeRef`/`nodeInfo` (Node existence, [`machine_controller_noderef.go`]()); a missing or NotReady Node is MHC’s `nodeStartupTimeout`/Node-condition checks’ job. - **Delete.** Deletion of a TerraformMachine is normally initiated by the Machine controller after drain and pre-terminate hooks. Ordering is CAPI’s: drain (bounded by `Machine.spec.deletion.nodeDrainTimeoutSeconds`) and volume detach (`nodeVolumeDetachTimeoutSeconds`) finish before the TerraformMachine is deleted, so `destroy` never races a running drain; the Node object is deleted by core-machine only after the TerraformMachine is gone (bounded by `nodeDeletionTimeoutSeconds`; [`machine_types.go`](), [`machine_controller.go`]()). No contract input carries these timeouts. The webhook **allows** a direct delete when the object has **no Machine ownerRef at all** (an orphan: KCP deletes the InfraMachine itself when KubeadmConfig or Machine creation failed, [`helpers.go`]()), when the owning Machine is itself deleting, or when the request carries `clusterctl.cluster.x-k8s.io/delete-for-move`; it **refuses** deletion otherwise, so drain and the KCP etcd-membership hook are never bypassed by deleting the infra object directly. Then a `destroy` Job runs, rendered from the `captf-inputs--` Secret (never from the bootstrap Secret or the Cluster, which may already be gone); on success it drops the state, the Lease and the inputs Secret, and removes the finalizer. If the owning Machine or Cluster is gone the same path runs. If there is no state Secret and no active Job (deleted while still gated), the finalizer is removed without a Job. If destroy fails, the object stays with `ApplyJobSucceeded=False/DestroyFailed` (`Ready=False`) and retries with backoff; a Lease whose holder Job no longer exists is force-unlocked automatically. There is no skip-destroy annotation: a permanently failing destroy is resolved by cleaning up out of band and removing the finalizer by hand, after backing up or un-owning the `tfstate-*` Secrets, which are otherwise garbage-collected with the object (see [the stuck-destroy runbook]()). A stuck control-plane machine destroy blocks KCP remediation, scale and upgrade until resolved ([`remediation.go`]()). - The module MUST accept `bootstrap_data` opaquely: it is base64 of the bootstrap payload (Inputs), passed to a base64-taking user-data argument or `base64decode()`d into a plain one; the module MUST NOT need to parse the payload. - **Taints and labels.** `Machine.spec.taints` and Machine labels are applied to the Node by core-machine ([`machine_types.go`](); [`machine_controller_noderef.go`]()), not by the infrastructure; the machine role has no `node_labels` input (pools do, see [`machinepool.md`]()); no role has a taints input. - **In-place updates** (the CAPI `InPlaceUpdates` feature gate, alpha) are unsupported: no field is mutable, and the webhook rejects the MachineSet’s SSA patch; the rollout path is a new Machine. ## Node providerID matching (module authors) CAPI links a Machine to its Node only when `Node.spec.providerID` equals `Machine.spec.providerID` exactly ([`machine_controller_noderef.go`]()). Our `provider_id` output is copied to the Machine verbatim, so the module MUST emit the same string the Node will carry: - **With a cloud-controller-manager (CCM):** the CCM sets `Node.spec.providerID` in its own format (for example `aws:////`, `azure:///subscriptions/...`, `openstack:///`, `vsphere://`). Emit exactly that format. - **Without a CCM:** the kubelet must be started with `--provider-id=` (via `KubeadmConfig` `nodeRegistration.kubeletExtraArgs`), and the value must be derivable on the instance at boot (an instance-metadata lookup, or a per-machine value baked into the module’s user-data wrapper), because `bootstrap_data` is shared per template and opaque to the module. - **Not supported in v1:** a controller-side patch of `Node.spec.providerID` once the Node registers, as CAPD does ([`machine.go`]()), would need a workload-cluster client and a manager flag; the controller does not implement it. In v1 the CCM or kubelet `--provider-id` must set it. - **RKE2:** CAPRKE2 itself never sets kubelet `provider-id`, `node-ip` or `node-name` (the fields exist in its config but nothing in the repo writes them; see [the RKE2ControlPlane guide]()). Without a CCM, the image MUST set `agentConfig.kubelet.extraArgs: [provider-id=]` from instance metadata so the kubelet-registered value matches our `provider_id` output exactly. This matters more under RKE2 than KCP: every RKE2 join (control plane or worker) requires `RCP.status.availableServerIPs` to be non-empty, which requires at least one **Ready** control-plane Machine, and a Machine only becomes Ready once its Node exists with a matching `providerID` — so a providerID mismatch on the first control-plane Machine blocks every subsequent join, not just that one Machine’s own readiness. Autoscale-from-zero (`InfraMachineTemplate.status.capacity` / `nodeInfo`; [`infra-machine.md`]() “InfraMachineTemplate: support cluster autoscaling from zero”) is supported through the **image**: instance size is fixed inside the module, so the image declares it with the OCI labels `io.captf.capacity` and `io.captf.node-info` (see [`image-contract.md`]() “OCI labels”), and a `TerraformMachineTemplate` reconciler copies them into `status.capacity`/`status.nodeInfo`. An image without the labels leaves both fields unset; the Cluster Autoscaler’s `capacity.cluster-autoscaler.kubernetes.io/*` annotations on the MachineDeployment/MachineSet remain the fallback. ## Control-plane machines When `control_plane = true`, the module is responsible for registering the instance in the control-plane load balancer’s backend or target group. The ids it needs (load balancer, target group, backend pool — whatever the cluster module’s cloud exposes) come from `captf_cluster_outputs` (see [`cluster.md`]() `exports`, which describes the publishing side of this same convention). Registration is done in the machine module’s own Terraform state, not the cluster module’s, so `destroy` on the machine deregisters it as an ordinary part of tearing down that state ([`infra-machine.md`]()). > [!WARNING] > > **Register the instance in the load balancer before kubeadm finishes** > > Ordering matters: the instance MUST be in the load balancer backend before `kubeadm init`/`kubeadm join` finish on it, because the contract requires the endpoint to be reachable through the load balancer during control-plane bring-up ([`status.go`](); see [the KubeadmControlPlane guide]() on reachability) — `Cluster.status.initialization.controlPlaneInitialized` never latches if the first control-plane node can’t be reached at the endpoint it just joined. In practice this means the module’s apply must register-then-boot (or register-then-poll-healthy) rather than boot-then-register as an afterthought. For worker Machines (`control_plane = false`) no load balancer registration applies; `captf_cluster_outputs` is still read for any other cluster-level values the module needs. ## Minimal skeleton A block written on one line may hold at most one argument ([`OneLineBlock`](); `tofu validate` reports “Invalid single-argument block definition” on a violation), so a variable that needs both `type` and `sensitive`, or `type` and `default`, must use the multi-line block form below. main.tf ```hcl variable "captf_contract" { type = string } variable "captf_cluster" { type = object({ name = string, namespace = string }) } variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) } variable "captf_cluster_outputs" { type = any default = null } variable "captf_tags" { type = map(string) } variable "machine_name" { type = string } variable "bootstrap_data" { # base64 of the bootstrap Secret's `value`; e.g. user_data_base64 = var.bootstrap_data type = string sensitive = true } variable "bootstrap_format" { type = string } variable "failure_domain" { type = string default = null } variable "kubernetes_version" { type = string default = null } variable "control_plane" { type = bool } # Tag every cloud resource this module creates with captf_tags (common.md # "captf_tags"); this stub only has to reference it, not create anything. resource "terraform_data" "tags" { input = var.captf_tags } output "provider_id" { value = "noop:///${var.captf_object.namespace}/${var.captf_object.name}" } output "addresses" { value = [{ type = "InternalIP", address = "10.0.0.1" }] } output "failure_domain" { value = var.failure_domain } output "interruptible" { value = false } output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } } ``` # MachinePool Role Implements the CAPI InfraMachinePool contract ([`infra-machinepool.md`](); MachinePool types in [`machinepool_types.go`]()) for a `TerraformMachinePool`. One workspace/state per `TerraformMachinePool`. The module manages a **native scaling group** (an ASG, VMSS, MIG, instance pool, or similar) and reports the group id, the list of provider IDs currently in it, and the desired capacity. Mutable: re-applied on spec or replica changes. The v1 controller implements fixed replicas and native autoscaling (`autoscaling.enabled`). It does not watch the bootstrap Secret: a rotated token reaches the pool at its next reconcile, at the latest one membership-refresh interval later (see “Bootstrap rotation” below). > [!WARNING] > > **An autoscaled pool module MUST ignore its desired count** > > Drift relies entirely on the module’s own `ignore_changes`: the drift Job already runs `apply -refresh-only` then `plan -detailed-exitcode` for every kind, and the controller does no plan-JSON filtering of the result — an autoscaled pool’s module MUST put its desired-count attribute under `lifecycle { ignore_changes = [...] }` (see Lifecycle and the `autoscaling` input) or a cloud-side scale reports as drift. MachinePool Machines (`status.infrastructureMachineKind`; optional, [`infra-machinepool.md`]() “MachinePoolMachines support”) are out of scope. Two consequences follow: > [!WARNING] > > **Pools have no drain and no MachineHealthCheck** > > There is **no drain** on scale-down, replacement or version roll — drain exists only for Machines ([`infra-machinepool.md`]() “MachinePoolMachines support … draining a node before scale down”), so modules SHOULD use native lifecycle hooks or termination handlers instead — and MachineHealthCheck never selects pool instances. ## Inputs Common inputs apply ([`common.md`]()), including `captf_cluster_outputs`. | Name | Type | Required | Value source | | --- | --- | --- | --- | | `machinepool_name` | `string` | yes | Owning CAPI `MachinePool.metadata.name` | | `replicas` | `number` | yes | The desired capacity of the group (below) | | `bootstrap_data` | `string`, `sensitive` | yes | Base64 of the bootstrap Secret’s `value` key (below) | | `bootstrap_format` | `string` | yes | as machine role; the `format` key is kubeadm-specific, not contract (see [`machine.md`]()) | | `failure_domains` | `list(string)` | yes (may be `[]`) | `MachinePool.spec.failureDomains` (below) | | `cluster_failure_domains` | `list(string)` | yes (may be `[]`) | Names from the cluster’s own state (below) | | `kubernetes_version` | `string` or `null` | yes (may be null) | `MachinePool.spec.template.spec.version` (below) | | `node_labels` | `map(string)` | yes (may be `{}`) | `MachinePool.spec.template.metadata.labels` verbatim (below) | | `autoscaling` | `object({enabled=bool, min=number, max=number})` | yes | Parsed from the MachinePool autoscaler annotations (below) | ### `replicas` (input) The **desired capacity** of the group. - With `autoscaling.enabled = false`: `MachinePool.spec.replicas`, authoritative and in the inputs hash (`*int32`, “Defaults to 1”; the pointer distinguishes an explicit 0 from unset, and 0 is allowed). - With `autoscaling.enabled = true`: `TerraformMachinePool.status.replicas` (the *observed* desired capacity from the last refresh) once one exists, else `MachinePool.spec.replicas` for the first apply before any refresh has run — either way **clamped into `[autoscaling.min, autoscaling.max]`** so a render never asks the cloud for an out-of-range desired count. The clamp applies only to this rendered input; the write-back (Lifecycle) always writes the raw observed value, unclamped. `replicas` is excluded from the inputs hash while `autoscaling.enabled`, so a non-replica-triggered apply (bootstrap rotation, module change, `exports` change) never fights the native autoscaler, and an observed-count change alone never triggers `InputsChanged`; a fixed-replica pool hashes `replicas` exactly as before. ### `bootstrap_data` (input) **Base64** of the raw bytes of the `value` key of the bootstrap Secret named by `MachinePool.spec.template.spec.bootstrap.dataSecretName` (`template` is `MachineTemplateSpec`; see [`bootstrap-config.md`]()), encoded by the controller for every bootstrap provider exactly as for the machine role (see [`machine.md`]() `bootstrap_data`: CAPRKE2 `gzipUserData: true` makes `value` raw gzip bytes that cannot be a `string`). Modules pass it to a base64-taking launch-configuration argument (for example `aws_launch_template.user_data`, which expects base64) or `base64decode()` it for a plain-text argument (UTF-8 payloads only). Hashed by content; `dataSecretName: ""` is treated as nil (`WaitingForBootstrapData`). **Rotates.** For MachinePools the kubeadm bootstrap provider re-creates the join token when it is past half its TTL (default TTL 15m, checked every TTL/3) and rewrites the same Secret in place, roughly every 7.5 minutes ([`token.go`]() `refreshBootstrapTokenIfNeeded`/`recreateBootstrapToken`; `kubeadmconfig_controller.go` `storeBootstrapData`). See Lifecycle for what that requires from modules. ### `failure_domains` (input) `MachinePool.spec.failureDomains` (at most 100 items, each 1–256 chars). `[]` means the module chooses. To make that choice informed, the cluster’s failure-domain names are injected as `cluster_failure_domains` (below); the full `attributes` maps are only available through `exports`. ### `cluster_failure_domains` (input) Names read from the cluster’s own state, off the `failure_domains` output (see [`cluster.md`]()) at render time — **not** from `TerraformCluster.status.failureDomains` — so a pool module can spread across all of them when `failure_domains` is `[]` without depending on `exports` conventions. Sourcing it from state rather than status keeps it stable across `clusterctl move` (status is not moved, and is empty until the target’s first cluster reconcile); it stays in the inputs hash. ### `kubernetes_version` (input) `MachinePool.spec.template.spec.version`. A change MUST roll the instances (see Lifecycle). As with the machine role, the value may carry a control-plane-provider-specific distro suffix (RKE2: `vX.Y.Z+rke2rN`, for example `v1.31.4+rke2r1`); the controller passes it to the module **verbatim**, and a module that needs it for an image lookup or a semver comparison MUST strip the `+…` suffix itself (see [`machine.md`]() and [the RKE2ControlPlane guide]()). ### `node_labels` (input) `MachinePool.spec.template.metadata.labels` verbatim (absent maps to `{}`). Pool instances have no Machine objects, so core CAPI never syncs labels onto their Nodes; the module MUST render these into kubelet registration (`--node-labels`) itself. `node_labels` is in the inputs hash, so an edit re-applies: new members pick it up, and existing members keep their registration. **The precise requirement and why it is not simple.** `bootstrap_data` is opaque (see [`machine.md`]() “the module MUST NOT need to parse the payload”) and, depending on the bootstrap provider and its settings, may be plain cloud-config, plain Ignition, or gzip of either. The module cannot edit the kubelet flags *inside* that payload without parsing it, which the contract forbids. The realistic options, in order of how much of `bootstrap_data` they need to understand: - **Cloud-config, uncompressed (`bootstrap_format == "cloud-config"`, CAPRKE2 `gzipUserData` unset or `false`):** wrap `bootstrap_data` and a second, module-generated part in a `multipart/mixed` MIME message (for example Terraform’s `cloudinit_config` data source, or an equivalent built by hand) so cloud-init runs both parts at boot. The second part writes a kubelet drop-in with `--node-labels`: ```hcl data "cloudinit_config" "node" { gzip = false base64_encode = true part { content_type = "text/cloud-config" # bootstrap_data is base64 of the raw payload (machine.md); decode it # only because this part must stay unmodified UTF-8 cloud-config, not # to parse or edit it. content = base64decode(var.bootstrap_data) } part { content_type = "text/cloud-config" content = yamlencode({ write_files = [{ path = "/etc/systemd/system/kubelet.service.d/20-node-labels.conf" content = "[Service]\nEnvironment=\"KUBELET_EXTRA_ARGS=--node-labels=${join(",", [for k, v in var.node_labels : "${k}=${v}"])}\"\n" }] }) } } # data.cloudinit_config.node.rendered is base64 multipart/mixed; pass it # to the same user-data argument bootstrap_data alone would have gone to. ``` RKE2’s agent config has the same shape through a second part that drops a file under `/etc/rancher/rke2/config.yaml.d/*.yaml` with a `node-label:` list, merged by the RKE2 agent at boot alongside the bootstrap part’s own `/etc/rancher/rke2/config.yaml`. - **Ignition (`bootstrap_format == "ignition"`) or a gzipped payload (CAPRKE2 `gzipUserData: true`):** neither is multipart-mixable the same way — Ignition has its own merge/append config mechanism, and a gzipped payload must be decompressed, edited or merged, and recompressed (or passed through a launch-configuration field that decompresses it itself). These need format-aware handling specific to the format; there is no one wrapper that covers them. - A module MAY instead declare, in its own documentation, only the `bootstrap_format`/compression combinations it supports, and fail loudly (a precondition, or an unsupported-value error) on the others, rather than silently dropping `node_labels`. Labels the kubelet may not self-assign (the `node-role.kubernetes.io/*` and other restricted `kubernetes.io`/`k8s.io` prefixes under `NodeRestriction`) are the module’s to filter. There is no taints input: the v1.14.2 MachinePool webhook rejects any `spec.template.spec.taints` (“taints feature for MachinePools is not yet implemented”, [`machinepool.go`](); `MachineTaintPropagation` is off by default, [`feature.go`]()). ### `autoscaling` (input) Parsed from the MachinePool annotations `cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size` and `-max-size` (`AutoscalerMinSizeAnnotation`/`AutoscalerMaxSizeAnnotation`, [`common_types.go`]()). This is the pool’s **only** autoscaling switch; there is no separate drift flag. `enabled=true` with both `min`/`max` parsed if and only if both annotations are present and `0 ≤ min ≤ max`; otherwise `{enabled=false, min=0, max=0}` (no annotations sets `AutoscalingActive=False/AutoscalingDisabled`), plus, when an annotation is present but the pair is incomplete, unparsable or `min > max`, the condition `AutoscalingActive=False/AutoscalingAnnotationsInvalid` (the pool still applies without autoscaling; `AutoscalingActive` never feeds `Ready`). **How this differs from CAPI’s own reading of the same annotations.** CAPI already reads these annotations on MachinePools, but only in a narrower path than ours: the MachinePool admission webhook defaults and clamps `spec.replicas` from them **only when `spec.replicas` is nil** (an explicit value, including 0, is left alone), **requires both annotations to be present**, rejects the create/update outright if either annotation is unparsable, and does **not** reject `min > max` — our controller does, via `AutoscalingAnnotationsInvalid` ([`machinepool.go`](), `calculateMachinePoolReplicas`, called from `Default`) — which is what makes them a sound source for the initial size. Our controller does not rely on the webhook: it reads the annotations itself on every reconcile for `autoscaling.min`/`.max`, and treats a missing or invalid annotation as `enabled=false` plus `AutoscalingActive=False/AutoscalingAnnotationsInvalid` rather than rejecting anything. **`enabled` means the module owns the desired count and its scaling policy:** it sets the group’s min/max to these values, configures whatever native scaling policy it wants (target tracking, scheduled, and so on), and MUST put the group’s desired-count attribute under `lifecycle { ignore_changes = [...] }` so an apply never resets what the cloud autoscaler decided. On the CAPI side the controller then claims `cluster.x-k8s.io/replicas-managed-by: captf` on the MachinePool when the annotation is absent or `"false"` (`ReplicasManagedByAnnotation`; CAPI counts the annotation as set for **any value except the literal string `"false"`**, `hasTruthyAnnotationValue`, [`helpers.go`](), called by `ReplicasManagedByExternalAutoscaler` — a foreign truthy value is left alone, never overwritten), and, on every reconcile pass on which `status.replicas` is known and differs from `MachinePool.spec.replicas`, writes the raw observed desired capacity back to `MachinePool.spec.replicas`. The InfraMachinePool docs make this the provider’s job ([`machine-pool.md`]() “It is the provider’s responsibility to update Cluster API’s Spec.Replicas property to the value observed”); see Lifecycle “Write-back” for the exact rules. With the annotation set, the MachinePool controller reports phase `Scaling` instead of `ScalingUp`/`ScalingDown` ([`machinepool_controller_phases.go`]()). > [!WARNING] > > **The Kubernetes Cluster Autoscaler does not act on these pools** > > Its `clusterapi` provider requires MachinePool Machines ([`README.md`]()), which this role excludes. Running it against a CAPTF pool anyway is unsupported, because its `spec.replicas` patches would be overwritten by the write-back. ## Outputs Common outputs apply (`health`). | Name | Type | Required | Maps to | | --- | --- | --- | --- | | `provider_id` | `string` or `null` | yes (may be null) | `TerraformMachinePool.spec.providerID`, the scaling-group id (below) | | `provider_id_list` | `list(string)` | yes (may be `[]`) | `TerraformMachinePool.spec.providerIDList` → `MachinePool.spec.providerIDList` (below) | | `replicas` | `number` | yes | The group’s observed desired capacity (below) | | `instances` | `list(object({provider_id=string, instance_id=optional(string), addresses=optional(list(object({type=string, address=string})), []), failure_domain=optional(string), state=optional(string)}))` | yes (may be `[]`) | `TerraformMachinePool.status.instances` (below) | ### `provider_id` (output) Maps to `TerraformMachinePool.spec.providerID` — the scaling-group id; optional in the contract and not used by core CAPI, 1–512 chars ([`infra-machinepool.md`]() “InfraMachinePool: providerID”). May stay null for group-less implementations. ### `provider_id_list` (output) Maps to `TerraformMachinePool.spec.providerIDList` → `MachinePool.spec.providerIDList` (mandatory; at most 10000 items, each 1–512 chars; [`infra-machinepool.md`]() “InfraMachinePool: providerIDList”). Rules: 1. It MUST contain **every non-terminated member** of the group regardless of health — pending, starting, standby, rebooting or unhealthy instances included — because CAPI deletes the Node of any providerID that leaves the list ([`machinepool_controller_noderef.go`]() `deleteRetiredNodes`; only an empty list with `status.replicas != 0` is guarded, [`machinepool_controller_phases.go`]()). Listing only “running” or “healthy” instances kills live Nodes. 2. Each entry MUST equal the Node’s `spec.providerID` exactly ([`machinepool_types.go`]() `providerIDList` “must match the provider IDs as seen on the node objects”; see [`machine.md`]() “Node providerID matching”), otherwise Nodes never get a nodeRef, keep the `node.cluster.x-k8s.io/uninitialized:NoSchedule` taint that the kubeadm bootstrap provider applies, and are eventually deleted. 3. Order is irrelevant to CAPI: the controller sorts and de-duplicates before writing, since CAPI compares the list with `reflect.DeepEqual` and any reorder churns status (CAPD sorts too). `tfcapi-lint` still warns (`output/provider-id-list-shape`) when the output’s own expression is not wrapped in `sort()`, because a module is applied and refreshed independently of the controller’s write path: an unsorted expression makes `provider_id_list` reorder between identical `apply -refresh-only` runs, which is a plan diff (and, for a module whose desired-count is not `ignore_changes`d, spurious drift) even though the controller’s own write is stable. Wrap the output’s `value` expression itself — the check inspects that expression directly, not a local it reads from: ```hcl output "provider_id_list" { value = sort(local.raw_ids) # or sort(distinct(...)) } ``` `value = local.sorted_ids` does not satisfy the check even when `sorted_ids` is itself `sort(...)`-derived, and JSON-syntax modules are not scanned at all. Only `sort()` fixes the order, so it must be there; `distinct()` alone does not satisfy the check, but may wrap inside or around `sort()`. The controller sorts and de-duplicates the list regardless. ### `replicas` (output) The group’s **desired capacity** as observed at refresh, not a count of running instances. Maps to `TerraformMachinePool.status.replicas` → `MachinePool.status.replicas` ([`infra-machinepool.md`]() “InfraMachinePool: replicas”; read in [`machinepool_controller_phases.go`]()). Without MachinePool Machines, CAPI sets the MachinePool’s `readyReplicas`/`availableReplicas` equal to this value ([`machinepool_controller_status.go`]()); only the deprecated v1beta1 `readyReplicas` is Node-based. Outside a scaling transition it MUST equal `length(provider_id_list)`; the controller derives `status.replicas` from this output and serializes `0` explicitly (a missing value would make CAPI keep the old count and block scale-to-zero forever, [`machinepool_controller_phases.go`]()). ### `instances` (output) Maps to `TerraformMachinePool.status.instances` — optional, provider-defined shape, not used by core CAPI ([`infra-machinepool.md`]() “InfraMachinePool: instances”; the contract’s own example shape is `{addresses, instanceName, providerID, version, ready}`, ours differs, which is allowed). One Go type, `MachinePoolInstance{ProviderID, InstanceID, Addresses []MachineAddress, FailureDomain, State}`, maps field by field: `provider_id`→`providerID`, `instance_id`→`instanceID`, `addresses`→`addresses` (validated like the machine role’s), `failure_domain`→`failureDomain`, `state`→`state` (the `health.state` enum). **Size.** `providerIDList` alone may reach 10000×512 bytes, and an object must stay under the etcd request limit of roughly 1.5 MiB, so the controller caps `status.instances` at 1000 entries (reported as `OutputsValid=True/InstancesTruncated` when it does), and the documented practical pool size is at most 2000 members. ### Provisioned rule `status.initialization.provisioned = true` — and, for as long as CAPI v1.14 reads it, `status.ready = true`; the MachinePool controller decides provisioning solely from `status.ready` ([`machinepool_controller_phases.go`]() via `external.IsReady`, and CAPD still sets it) — when the state Secret carries the inputs-hash annotation of a successful apply and `health.state != "pending"`. `provider_id` is **not** required, since the contract makes it optional. `provider_id_list` may legitimately be empty when `replicas == 0`, and a 0-member group reports `running`/healthy (see [`common.md`]()). **Latched.** The formula applies only until it first holds; once true, `provisioned` and `status.ready` stay true for the object’s life, whatever later health or outputs say. It is derived from state only, so it is rebuilt after `clusterctl move` (the formula is re-evaluated once on the target, then latched again); CAPI does not latch its copy, so the MachinePool shows `infrastructureProvisioned=false` briefly after a move until the first reconcile on the target. Separately, CAPI copies `TerraformMachinePool.spec.providerIDList`/`status.replicas` to the MachinePool only after `ClusterCache.GetClient` succeeds — that is, only once the workload cluster is reachable ([`machinepool_controller_phases.go`]()); until then the TerraformMachinePool side is correct but the MachinePool does not reflect it. ## Lifecycle - **Mutable.** Re-apply on: a `spec.source` change (image, pull policy); a `replicas` change on the MachinePool (unless `autoscaling.enabled`, where `replicas` is excluded from the inputs hash and a user edit of `MachinePool.spec.replicas` is overwritten by the write-back, surfaced as an event); a `failure_domains` change; a `node_labels` change (applied like `bootstrap_data`: a launch-configuration update, no instance replacement); a `kubernetes_version` change; a `bootstrap_data` change; or a `captf_cluster_outputs` content change. Every apply, including a drift `Remediate`, re-renders the **current** inputs; nothing is replayed from an older run. - **Bootstrap rotation.** Because the kubeadm bootstrap provider rewrites the pool’s bootstrap Secret roughly every 7.5 minutes (`bootstrap_data` above), a pool is re-applied at that cadence for the life of the pool. The controller does not watch the bootstrap Secret — it is not `captf.io/managed`, so it is not in the cache — so a rotated token reaches the pool at its next reconcile, at the latest one membership-refresh interval later, and is then coalesced into the next apply. Module rules that make this cheap and safe: - a `bootstrap_data` change MUST be applied as an in-place update of the group’s launch configuration (a launch template version, instance template, or VMSS model) so that *new* members get the current token, and MUST NOT replace existing instances; - a `kubernetes_version` change MUST roll (replace) the instances, because a ClusterClass upgrade completes only when the Nodes’ kubelet versions match ([`upgrade.go`]() `IsMachinePoolUpgrading`; CAPD’s DevMachinePool rolls on version only and never on bootstrap data, [`dockermachinepool_backend.go`]()); - modules SHOULD expose no other trigger that replaces instances on an apply with unchanged `kubernetes_version`. Each rotation re-applies the pool, so a module that never converges (see below) is applied roughly every 7.5 minutes indefinitely, on top of its 30s refresh cadence. - **Membership refresh.** `provider_id_list`, `replicas`, `instances` and `health` are only as fresh as the last apply or refresh, and a Node joining the group is unschedulable until its providerID appears in `MachinePool.spec.providerIDList` — CAPI removes the kubeadm-applied `node.cluster.x-k8s.io/uninitialized:NoSchedule` taint only then ([`machinepool_controller_noderef.go`]()). The controller therefore runs `apply -refresh-only` on its own **membership-refresh interval**, `spec.membershipRefreshIntervalSeconds` (default 60, 15–86400, may not be 0), and additionally right after every apply and repeatedly (every 30s) while `health.state == "pending"` or `length(provider_id_list) != replicas` — a module that never converges is therefore refreshed every 30s indefinitely. `drift.intervalSeconds: 0` is rejected by the CRD schema (minimum 1) for a pool’s own field, by contrast with a machine’s or the cluster’s `drift.intervalSeconds`, where 0 is a valid setting that disables drift; an inherited `defaults.drift.intervalSeconds: 0` (unset) falls back to the controller’s default drift interval (`--drift-default-interval`, 30m) rather than to 0. Drift may be sparse but membership never is. - **Drift order.** Drift for a pool is the same job every other kind runs — `apply -refresh-only` first, then `plan -detailed-exitcode` ([`internal/runner/plan.go`]()) — and the controller does no plan-JSON filtering of the result: whatever the refreshed plan reports is drift. This is why the desired-count `lifecycle { ignore_changes = [...] }` in the `autoscaling` input is load-bearing, not optional: with `autoscaling.enabled`, the module owns the desired count, and a module that does not `ignore_changes` it reports every cloud-side scale as drift. With `Remediate` a detected diff re-applies the pool; with `autoscaling.enabled = false` that scales the group back to `MachinePool.spec.replicas` — the only mode where the module doesn’t already `ignore_changes` the desired count. - **Write-back (`autoscaling.enabled`).** On every reconcile pass, one patch (`SyncReplicas`): the controller claims `cluster.x-k8s.io/replicas-managed-by: captf` on the MachinePool when the annotation is absent or `"false"` (a foreign truthy value already means another controller manages it, and is left untouched), and when `status.replicas` is known and differs from `MachinePool.spec.replicas`, writes the raw observed value (unclamped by `[min,max]`, unlike the rendered `replicas` input) to `spec.replicas` in the same patch. A `ReplicasWrittenBack` Normal event is emitted on the TerraformMachinePool whenever a write happens. With `autoscaling.enabled = false`, the controller removes `replicas-managed-by` only if it still carries `captf` (never a foreign value), and never writes `spec.replicas`. Without the write-back, CAPI would report `ScalingUp`/`ScalingDown` forever and `upToDateReplicas` would follow the stale spec ([`machinepool_controller_status.go`]()). A ClusterClass-managed MachinePool MUST leave `replicas` unset in the topology for this to work — the topology controller would otherwise reassert it every reconcile, fighting the write-back. **Module note.** Because `lifecycle { ignore_changes }` cannot be conditional on a variable, a module that supports both `autoscaling.enabled = true` and `= false` from the same scaling-group resource should, when disabled, pin the group’s `min_size`/`max_size` to `var.replicas` — not to `var.autoscaling.min`/`.max`, which are `0` when disabled — so the same `ignore_changes` block still lets the group track `replicas` exactly. - **Health.** `health` feeds conditions only; there is no remediation for pools in v1 (MHC selects Machines, and pool replicas only have Machines under MachinePool Machines, which this role excludes; [`infra-machinepool.md`]() “MachinePoolMachines support”). Instances terminated by the cloud simply leave `provider_id_list` at the next refresh, and CAPI deletes their Nodes. - **Delete.** A `destroy` Job renders from the object-owned `captf-inputs--` Secret, not from the bootstrap Secret: CAPI deletes the bootstrap config and the InfraMachinePool in the same pass ([`machinepool_controller_phases.go`]() `reconcileDeleteExternal`), so the bootstrap Secret is usually already gone when destroy runs. Then the controller drops state, the Lease and the inputs Secret, and removes the finalizer. **Node cleanup is partial.** On MachinePool deletion, CAPI deletes only the Nodes whose providerIDs are *absent* from the last-seen `providerIDList` (`reconcileDeleteNodes` → `deleteRetiredNodes`, [`machinepool_controller_noderef.go`]()); Nodes still listed survive and are removed by the cloud-controller-manager once their instances are gone, or remain as orphans without a CCM. A module MAY empty `provider_id_list` on the final refresh before destroy, but the controller does not depend on it. ## Minimal skeleton A block written on one line may hold at most one argument ([`OneLineBlock`](); `tofu validate` reports “Invalid single-argument block definition” on a violation), so a variable that needs both `type` and `sensitive`, or `type` and `default`, must use the multi-line block form below. main.tf ```hcl variable "captf_contract" { type = string } variable "captf_cluster" { type = object({ name = string, namespace = string }) } variable "captf_object" { type = object({ kind = string, name = string, namespace = string }) } variable "captf_cluster_outputs" { type = any default = null } variable "captf_tags" { type = map(string) } variable "machinepool_name" { type = string } variable "replicas" { type = number } variable "bootstrap_data" { # base64 of the bootstrap Secret's `value`; e.g. launch template user_data = var.bootstrap_data type = string sensitive = true } variable "bootstrap_format" { type = string } variable "node_labels" { type = map(string) } variable "failure_domains" { type = list(string) } # always set, may be []; no default (input/default) variable "cluster_failure_domains" { type = list(string) } # always set, may be []; no default (input/default) variable "kubernetes_version" { type = string default = null } variable "autoscaling" { type = object({ enabled = bool, min = number, max = number }) } # Tag every cloud resource this module creates with captf_tags (common.md # "captf_tags"); this stub only has to reference it, not create anything. resource "terraform_data" "tags" { input = var.captf_tags } locals { # Unsorted; provider_id_list must wrap sort() directly in its # own output expression: see "provider_id_list order" below. raw_ids = [for i in range(var.replicas) : "noop:///${var.captf_object.namespace}/${var.captf_object.name}/${i}"] } output "provider_id" { value = "noop-group:///${var.captf_object.namespace}/${var.captf_object.name}" } output "provider_id_list" { value = sort(local.raw_ids) } output "replicas" { value = var.replicas } output "instances" { value = [for id in sort(local.raw_ids) : { provider_id = id, state = "running" }] } output "health" { value = { state = "running", healthy = true, message = null, reasons = [] } } ``` A real module with `autoscaling.enabled` sets the group’s `min_size`/`max_size` from `var.autoscaling`, its `desired_capacity` from `var.replicas`, and declares `lifecycle { ignore_changes = [desired_capacity] }` on the group resource; `output "replicas"` then reads the group’s *current* desired capacity from the resource (refreshed by `apply -refresh-only`), not `var.replicas`. ## Per-instance state `instances[*].state` uses the `health.state` enum (see [`common.md`]()) so per-instance conditions can be derived later. **Deriving group `health` from mixed instance states.** The controller does not aggregate `instances[*].state` into `health` itself: `health` is copied to the `InfrastructureHealthy` condition, and `instances` is copied to `status.instances` (capped, see “Size” above), with no cross-referencing between the two — a module that reports a healthy `health` and a list full of `degraded` instances gets a healthy `InfrastructureHealthy` and a degraded-looking `status.instances`, since nothing else reconciles the difference. A pool module SHOULD derive `health` from the instances it is about to report: - `healthy = true`, `state = "running"` when the group has reached its desired capacity and every counted instance’s own state is `running`; - `healthy = false`, `state` reflecting the worst instance (for example `degraded` if any instance is `degraded`, else `stopped`, and so on down the enum) with `reasons` naming the affected instances, when some instances are not `running`; - `state = "pending"` only while the group itself does not yet exist or has no members to report at all — not merely while some members are still starting, since `length(provider_id_list) != replicas` already forces the controller’s fast (30s) refresh loop on its own (see “Membership refresh” below), so a module does not need `health.state = "pending"` to get that cadence during a scale-up or scale-down. This is a SHOULD, not a MUST: a module that cannot observe per-instance health may still report `{state="running", healthy=true}` as long as the group itself is healthy (see [`common.md`]()). # Image Contract The OCI image is the deliverable a module author ships. This page is the normative image contract for the `v1alpha1` module contract: the fixed paths CAPTF’s runner looks for, the labels it and `tfcapi-lint` read, the user the image should run as, and multi-arch publishing. `tfcapi-lint image` checks an image against this contract. What the module actually sees when the runner executes it — the generated root, the environment, the commands run and their order — is on [Runtime Environment](). One image bundles one role module (`cluster`, `machine` or `machinepool`) *and* the runtime that runs it (`tofu` or `terraform`). There is no separate module source and no separate runtime image: the image tag is the module version, the image is what `spec.source.image` references, and the image is what gets pinned, moved, rolled out and audited. The controller runs it as a Job with its own runner binary as the entrypoint; the image itself never needs a shell. ## Fixed paths The runner discovers a module by fixed path, not by label or registry metadata, and fails the Job at start (`error.kind: image-layout`) if a required path is missing or unusable. | Path | Required | Contents | | --- | --- | --- | | `/captf/module/` | yes | The role module: at least one `.tf`, `.tf.json`, `.tofu` or `.tofu.json` file at its top level, plus any local module it references by a relative `source`. It cannot declare a `terraform { backend … }` or `cloud` block: those are only valid in a root module, and the generated root, not this one, is the root. | | `/captf/runtime` | yes | The `tofu` or `terraform` binary: a regular file, or a symlink to one inside the image, executable by the image’s `USER`. It must support the Terraform 1.x / OpenTofu 1.x CLI surface: `version`, `init`, `validate`, `plan`, `apply`, `destroy`, `force-unlock`, `show`, and `state push`/`state list`. There is no override for this path. | | `/captf/providers/` | no | An optional provider filesystem mirror (below). Without it, `init` needs registry egress. | | `/captf/work/`, `/captf/bin/`, `/captf/config/`, `/var/run/captf/credentials/` | must be empty | The Job mounts an `emptyDir`, the runner binary, the per-run Secret and the identity’s credential files at these paths respectively. Anything the image ships under them is shadowed (or, for `/captf/work`, never used, since the image’s own root filesystem is read-only by default). | Everything else in the image is the author’s business: CA certificates, `git` for `provider` blocks that shell out, or a helper binary a `local-exec` provisioner calls. ### Provider mirror layout Build `/captf/providers` with `terraform providers mirror ` or `tofu providers mirror `, for every platform the image publishes (for example `-platform=linux_amd64 -platform=linux_arm64`). Both runtimes accept either layout it can produce: the *packed* layout (`HOST/NAMESPACE/TYPE/terraform-provider-TYPE_VERSION_TARGET.zip` plus `.json` index files) or the *unpacked* layout (`HOST/NAMESPACE/TYPE/VERSION/TARGET/`). `providers mirror` creates its target directory only when it writes at least one provider, so a module that requires none needs the mirror stage to create `/captf/providers` itself (see the reference Containerfiles below) or the later `COPY --from=mirror` fails. It also fails with “Module not installed” when the module calls local modules that are not yet installed, so run `terraform get` / `tofu get` first, in the same stage; that installs modules only, never providers. Providers are optional: an image without `/captf/providers` still works, but needs registry egress at `init` and is slower and non-hermetic. The reference images in this repository ship a mirror. How the runner uses the mirror at run time is on [Runtime Environment](). ## OCI labels Labels are metadata only: the runner never reads them for behavior. `tfcapi-lint image` checks them for consistency with `--role`/`--contract`. | Label | Value | | --- | --- | | `io.captf.contract` | Contract version, for example `v1alpha1`. | | `io.captf.role` | `cluster`, `machine` or `machinepool`. | | `io.captf.runtime` | `tofu` or `terraform`: what `/captf/runtime` is. | | `io.captf.runtime.version` | For example `1.12.6`. | | `org.opencontainers.image.source`, `.revision`, `.version` | Standard OCI annotations; `.version` should equal the tag. | The reference Containerfiles below take these three as `IMAGE_SOURCE`, `IMAGE_REVISION` and `IMAGE_VERSION` build args (empty by default), so a build pipeline sets them with `--build-arg` from the source repository URL, the commit, and the tag being built. **Capacity labels** (machine role only, optional) let `TerraformMachineTemplate` support Cluster Autoscaler scale-from-zero: the module fixes the instance type, so the image is the only place that knows the node’s size. Set both identically on every platform of a multi-arch index; an image without them leaves `status.capacity`/`status.nodeInfo` unset. | Label | Value | Maps to | | --- | --- | --- | | `io.captf.capacity` | JSON object, resource name to Kubernetes quantity string, for example `{"cpu":"4","memory":"16Gi","nvidia.com/gpu":"1"}`. Each key is a valid Kubernetes resource name; each value parses as a quantity. | `TerraformMachineTemplate.status.capacity` | | `io.captf.node-info` | JSON object `{"architecture":"amd64","operatingSystem":"linux"}`; `architecture` is one of `amd64`, `arm64`, `s390x`, `ppc64le`; at least one key set. | `TerraformMachineTemplate.status.nodeInfo` | A module whose instance type varies needs one image per instance type to use these labels; pool images may carry them, but they are ignored. ## User Any UID works for the runner, but recommend a non-root `USER` (for example `65532`) so the Job can run under a namespace that enforces the Pod Security `restricted` profile. Files and directories under `/captf/module` and, when present, `/captf/providers` must be readable, and directories traversable, by that UID. Credential files are mounted mode `0440`, so a non-root image user reads them through the pod’s `fsGroup`, not through ownership. See [Security Model]() for how the pod’s own security context defaults interact with the image’s `USER`. ## Multi-arch Publish a multi-arch manifest (`linux/amd64`, `linux/arm64`) or pin the management cluster’s node architecture to the platform the image ships: `tfcapi-lint image` checks the `linux/amd64` platform by default, and `--platform`/`--all-platforms` select others. A provider mirror must carry a package for the target platform it is checked against. ## Versioning and pinning The image tag is the module version: `spec.source.image` is `registry/repo:tag` or `registry/repo@sha256:…`, and a new module version is a new tag referenced by a new `Terraform*Template`. The controller pins the digest it actually ran after the first successful apply and, for immutable machines, uses that digest for every later drift and destroy Job; see [Security Model]() for the full mechanics and why it matters. Contract version is not declared in the image: labels are informational. The controller injects the contract it generates against as `captf_contract`, and `tfcapi-lint` takes `--contract` on the command line. ## Trust boundary Referencing an image is equivalent to granting its publisher the runner’s Secret access and the resolved identity’s cloud credentials in that namespace: see [Security Model]() for the full trust boundary and what it means for review and tenancy. ## Building an image Lint the module first, then build, then lint the pushed image: ```sh tfcapi-lint module ./cluster --role cluster --strict podman build -t "$IMAGE" . tfcapi-lint image "$IMAGE" --role cluster --strict ``` The reference images below (a Terraform tab and an OpenTofu tab, after the two base notes) ship as [`examples/Containerfile.terraform`]() and [`examples/Containerfile.opentofu`](), each taking `ARG ROLE` and `ARG RUNTIME_VERSION`. ### Reference: Terraform base `hashicorp/terraform` is Alpine with `git`, `openssh` and CA certificates, its binary at `/bin/terraform`, and `ENTRYPOINT ["/bin/terraform"]`; the runner replaces that entrypoint, so it has no effect. The image has no `USER` (root); the Containerfile below adds one. ### Reference: OpenTofu base `opentofu:*-minimal` is `FROM scratch` with only the `tofu` binary — no CA certificates, no shell, no `git` — which is why the Containerfile below copies it into an Alpine stage instead of using it directly as the final base. The full (non-`minimal`) OpenTofu image refuses to be used as a `FROM` base. Containerfile.terraform ```dockerfile # Reference source image, Terraform base (docs/book/src/module-author/image-contract.md "Reference: # Terraform base"). Run from your module's root directory: # # podman build -f Containerfile.terraform --build-arg ROLE=cluster \ # --build-arg IMAGE_SOURCE=https://github.com// \ # --build-arg IMAGE_REVISION="$(git rev-parse HEAD)" \ # --build-arg IMAGE_VERSION= -t /: . # # The image tag is the module version. Lint first: # tfcapi-lint module . --role cluster --strict ARG RUNTIME_VERSION=1.16.4 # The org.opencontainers.image.* labels below (image-contract.md "OCI # labels"): leave these unset for a local/test build, or pass them from # your CI pipeline (source repo URL, commit SHA, the image tag). ARG IMAGE_SOURCE="" ARG IMAGE_REVISION="" ARG IMAGE_VERSION="" # Optional but recommended: hermetic provider mirror for the platforms you # publish. Needs registry egress at build time; drop this stage (and the # COPY --from=mirror below) for a non-hermetic image. FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION} AS mirror WORKDIR /src COPY . /src # get: `providers mirror` refuses a module whose nested local modules are not # installed; get installs them (no providers), in this stage only. # mkdir: `providers mirror` does not create the target when the module # requires no providers, and the COPY --from=mirror below needs it. RUN terraform get \ && mkdir -p /captf/providers \ && terraform providers mirror -platform=linux_amd64 -platform=linux_arm64 /captf/providers FROM docker.io/hashicorp/terraform:${RUNTIME_VERSION} ARG ROLE=cluster ARG RUNTIME_VERSION ARG IMAGE_SOURCE ARG IMAGE_REVISION ARG IMAGE_VERSION COPY --from=mirror /captf/providers /captf/providers COPY . /captf/module RUN ln -s /bin/terraform /captf/runtime \ && adduser -D -u 65532 captf \ && chown -R 65532:65532 /captf USER 65532 LABEL io.captf.contract="v1alpha1" \ io.captf.role="${ROLE}" \ io.captf.runtime="terraform" \ io.captf.runtime.version="${RUNTIME_VERSION}" \ org.opencontainers.image.source="${IMAGE_SOURCE}" \ org.opencontainers.image.revision="${IMAGE_REVISION}" \ org.opencontainers.image.version="${IMAGE_VERSION}" ``` Containerfile.opentofu ```dockerfile # Reference source image, OpenTofu base (docs/book/src/module-author/image-contract.md "Reference: # OpenTofu base"). Run from your module's root directory: # # podman build -f Containerfile.opentofu --build-arg ROLE=machine \ # --build-arg IMAGE_SOURCE=https://github.com// \ # --build-arg IMAGE_REVISION="$(git rev-parse HEAD)" \ # --build-arg IMAGE_VERSION= -t /: . # # The image tag is the module version. Lint first: # tfcapi-lint module . --role machine --strict # # The full ghcr.io/opentofu/opentofu image refuses to be a FROM base # (ONBUILD RUN exit 1); use the -minimal tag via COPY --from as below. ARG RUNTIME_VERSION=1.12.6 # The org.opencontainers.image.* labels below (image-contract.md "OCI # labels"): leave these unset for a local/test build, or pass them from # your CI pipeline (source repo URL, commit SHA, the image tag). ARG IMAGE_SOURCE="" ARG IMAGE_REVISION="" ARG IMAGE_VERSION="" FROM ghcr.io/opentofu/opentofu:${RUNTIME_VERSION}-minimal AS tofu # Optional but recommended: hermetic provider mirror for the platforms you # publish. Needs registry egress at build time; drop this stage (and the # COPY --from=mirror below) for a non-hermetic image. FROM docker.io/library/alpine:3.22 AS mirror RUN apk add --no-cache ca-certificates COPY --from=tofu /usr/local/bin/tofu /usr/local/bin/tofu WORKDIR /src COPY . /src # get: `providers mirror` refuses a module whose nested local modules are not # installed; get installs them (no providers), in this stage only. # mkdir: `providers mirror` does not create the target when the module # requires no providers, and the COPY --from=mirror below needs it. RUN tofu get \ && mkdir -p /captf/providers \ && tofu providers mirror -platform=linux_amd64 -platform=linux_arm64 /captf/providers FROM docker.io/library/alpine:3.22 ARG ROLE=machine ARG RUNTIME_VERSION ARG IMAGE_SOURCE ARG IMAGE_REVISION ARG IMAGE_VERSION RUN apk add --no-cache ca-certificates \ && adduser -D -u 65532 captf COPY --from=tofu /usr/local/bin/tofu /captf/runtime COPY --from=mirror /captf/providers /captf/providers COPY . /captf/module RUN chown -R 65532:65532 /captf USER 65532 LABEL io.captf.contract="v1alpha1" \ io.captf.role="${ROLE}" \ io.captf.runtime="tofu" \ io.captf.runtime.version="${RUNTIME_VERSION}" \ org.opencontainers.image.source="${IMAGE_SOURCE}" \ org.opencontainers.image.revision="${IMAGE_REVISION}" \ org.opencontainers.image.version="${IMAGE_VERSION}" ``` > [!NOTE] > > **Mirroring providers needs registry egress at build time** > > Drop the `mirror` stage (and its `COPY --from=mirror`) for a non-hermetic image. A `distroless/static:nonroot` final stage also works, and needs no shell, if the module needs no other tool: `tofu` is statically linked. ## Checklist for `tfcapi-lint image` Summarized: `/captf/module` present and lints clean for the role; `/captf/runtime` present and executable; `/captf/providers`, if present, follows the mirror layout and covers every `required_providers` entry for the platform being checked; labels, if present, agree with `--role`/`--contract`; the capacity labels, if present, are valid JSON of the shapes above; no files under the reserved paths; `config.User` is non-root (running as root is a warning, not an error). See [tfcapi-lint CLI: checks]() for every check’s ID, severity and the roles it applies to. > [!NOTE] > > **See also** > > - [Runtime Environment]() > - [Module Contract]() > - [tfcapi-lint]() > - [Security Model]() # Changelog Changes to the module contract, newest first. Within `v1alpha*` there is no compatibility guarantee (see [`README.md`]() “Versioning”); every change is still recorded here. ## Unreleased ### Added - Machinepool role, including native autoscaling (additive within v1alpha1; `machinepool.md`, `README.md` “Roles and CAPI mapping”): the v1 controller now reconciles `TerraformMachinePool`. New API kinds `TerraformMachinePool` and `TerraformMachinePoolTemplate`, and new field `TerraformMachinePoolSpec.membershipRefreshIntervalSeconds` (default 60, 15–86400, rejected below 15 by the CRD schema, not the webhook). New `OutputsValid` reason `InstancesTruncated`, set when `status.instances` is capped at 1000 entries. Manager flag `--terraformmachinepool-concurrency` (default 10). The controller does not watch the bootstrap Secret (a rotated token reaches the pool at its next reconcile, at the latest one membership-refresh interval later) and does not support MachinePool Machines. - Machinepool autoscaling (additive within v1alpha1; `machinepool.md` “Inputs” `autoscaling`/`replicas`, “Lifecycle” “Drift order” and “Write-back”): the `autoscaling` input is parsed from the MachinePool’s `cluster-api-autoscaler-node-group-min-size`/`-max-size` annotations; when enabled, `replicas` renders `status.replicas` (falling back to `MachinePool.spec.replicas` before the first observation) clamped into `[min,max]`, is excluded from the inputs hash, and the controller writes the raw observed value back to `MachinePool.spec.replicas` and claims `cluster.x-k8s.io/replicas-managed-by: captf` (never a foreign truthy value), removing its own claim when autoscaling is disabled again. New condition `AutoscalingActive` (informational, pool-only, never in `Ready`; reasons `ReplicasManagedByModule`, `AutoscalingDisabled`, `AutoscalingAnnotationsInvalid`). New event `ReplicasWrittenBack`. New RBAC: `patch` on `cluster.x-k8s.io` `machinepools` (write-back needs it; `get`/`list`/`watch` already covered reads). Drift for a pool already ran `apply -refresh-only` then `plan` like every other kind; nothing changed there — it is the module’s `lifecycle { ignore_changes }` on the desired-count attribute that keeps a cloud-side scale from reporting as drift, and that responsibility is now load-bearing rather than theoretical. **Upgrade behavior:** a MachinePool that already carries both autoscaler annotations before the upgrade starts autoscaled and re-applies once right after it — its `autoscaling` input flips from `{enabled=false,min=0,max=0}` to the parsed, enabled value, which is hashed, so the inputs hash changes and the pool re-applies. - Plan preview and approval (the module contract is unchanged; see [Plan Approval]()): TerraformCluster and TerraformClusterTemplate gain `spec.applyPolicy` (`Automatic`, the default applied at reconcile, or `Manual`; mutable, not hashed). Under `Manual` every apply but the first (no state yet), including a drift remediation and a retry, first runs a new operation, `plan` (added to the `Operation` enum): a Job that runs `init`, `validate`, `plan -out` and `show -json` under the run lease only and applies nothing. Its plan (counts, up to 50 `"
()"` entries, never values, and a plan hash `p1:` of the sorted `address|actions` of every change) goes to the new `status.plan` (`inputsHash`, `job`, `planHash`, `add`, `change`, `destroy`, `resources`, `truncated`, `createdAt`), and the apply waits until the new `captf.io/approve-plan` annotation names that hash. The approved apply (runner flag `--expect-plan`) plans again and applies only an identical plan; otherwise it stops with the new `RunErrorKind` `plan-changed` and the new plan waits for its own approval. A plan without changes never waits. Approving a plan also approves its deletes; the destructive-plan annotation is not needed on top. The annotation is removed and `status.plan` cleared once the approved apply succeeds. New `ApplyJobSucceeded` Unknown reasons `PlanAwaitingApproval` and `PlanChanged`; events `PlanReady`, `PlanApproved`, `PlanApplied`, `PlanChanged`; metric `captf_plan_approvals_total{kind,result}` (`approved`, `changed`) and `captf_jobs_total` result `plan_changed`; Job annotations `captf.io/approved-plan`, `captf.io/plan-changed` and `captf.io/plan-unreadable`. - State backups and restore (the module contract is unchanged; see [Terraform State]() and [State Restore]()): every new state serial the manager reads is copied verbatim into `captf-state-backup--` Secrets (plus `-part-N` per chunk), owned by the Terraform\* object and labeled for `clusterctl move`; the newest `--state-backups` (default 5, 0 disables) are kept. TerraformCluster and TerraformMachine gain `status.stateBackups` (up to 16 `{serial, takenAt, bytes}`). The `captf.io/restore-state: ""` annotation starts a new operation, `restore` (added to the `Operation` enum of `status.activeJob` and `status.lastRun`): a Job that runs `init`, `state push -force` of the backup and `state list`, under the run lease (and the cluster write lease for a TerraformCluster). It precedes apply, drift and refresh but not deletion; the annotation is removed on success, and a failure is not retried for the same serial. New condition `RestoreJobSucceeded` (never in `Ready`) with reasons `StateRestored`, `RestoreFailed`, `RestoreBackupNotFound` and the three lease waits; events `StateBackedUp`, `StateRestored`, `StateRestoreFailed`; metrics `captf_state_backups_total{kind,result}` and `captf_state_restores_total{kind,result}`; runner steps `state-push` and `state-list`. A TerraformCluster’s restore takes part in the cluster operation gate like its apply. - User variables (see [`common.md`]() “User variables” and [Module Variables]()): TerraformCluster and TerraformMachine (and both templates) gain `spec.variables`, an inline JSON object, and `spec.variablesFrom`, up to 16 ConfigMaps or Secrets labeled `captf.io/variables=true`, each `optional` and read in `format` `String` (default) or `JSON`. Sources merge in list order, a later one winning; inline variables win over all. Each variable becomes a named argument of `module "role"`, declared in the generated root without a type (`sensitive = true` when its value came from a Secret) and valued in `terraform.tfvars.json`; a name the module does not declare fails the apply. Reserved: `captf_` names, the role’s contract inputs and the module meta-arguments. The inputs hash covers the variables only when an object sets some, so existing objects keep their `h2` hash and do not re-apply. On a TerraformCluster a changed variable or source re-applies; on a TerraformMachine both fields are immutable and read only until the machine is provisioned. New `DependenciesReady` False reasons `VariablesSourceNotFound` and `VariablesInvalid`. The manager’s ClusterRole gains `get`, `list` and `watch` on `configmaps`, and it runs a second, label-scoped cache (`captf.io/variables=true`, data stripped) for the source watches. - Run leases (the module contract is unchanged; see [The Reconcile Lifecycle]()): the manager holds a `coordination.k8s.io/v1` Lease `captf-run-` per object before it creates a Job, so a stale Job cache, a leader-election handover or two manager instances with overlapping `--watch-filter` values cannot run two Jobs of one object. With the new `--cluster-operation-gate` (default true) a TerraformCluster’s apply or destroy and its machines’ applies and destroys never run at once (cluster write Lease `captf-cluster-`); refresh and drift are not gated. New `ApplyJobSucceeded` Unknown reasons `WaitingForRunLease`, `WaitingForClusterOperation` and `WaitingForMachineOperations` (`DriftJobSucceeded` gains `WaitingForRunLease`), the Normal events of the same names once per wait, and `captf_lease_waits_total{kind,reason}`. The manager’s ClusterRole gains `create` and `update` on `leases`. - Events for every stage (the module contract is unchanged; see [Observability]() “Events”): the runner now reports its progress as `events.k8s.io/v1` Events on the object that owns the Job (`RunStarted`, `StepStarted`, `StepSucceeded`, `StepFailed`, `PlanSummary`, `ResourcesChanged`, `RunFinished`; reporting controller `captf.io/runner`), best effort and never failing the run. Job pods get two new runner flags, `--event-object` and `--job-name`; the manager’s new `--runner-events` (default true) turns them off. The `captf-runner` ClusterRole gains `events.k8s.io` `events` `create`. The manager emits new reasons once per transition: `JobSucceeded`, `JobInterrupted`, `JobDeadlineExceeded`, `StuckJobDeleted`, `DestructivePlanApprovalConsumed`, `DeletionStarted`, `FinalizerRemoved`, `Paused`, `Resumed`, `ProviderIDSet`, `ControlPlaneEndpointSet`, `FailureDomainsChanged`, `InputsChanged`, `StateAdopted`, `StateLost`, `StateLocked` (both formerly `StateUnreadable`), `DriftResolved`, `DriftRemediationStarted`, `InstanceHealthy`, `InstanceUnhealthy`, `MirrorCreated`, `MirrorRemoved`, `IdentitySecretFound`, `IdentitySecretNotFound`, `CapacityResolved`, and `ConditionChanged` for any other owned condition’s status or reason change. `JobFailed` now also covers failed drift and refresh Jobs; `JobCreated` names the attempt, the image and why the Job started. - Destructive-plan guard: a TerraformCluster apply, including a drift remediation, runs `plan -out`, `show -json` and an apply of the saved plan, and stops before a plan whose actions include `delete` (a removal or a replacement) unless the `captf.io/approve-destructive-plan` annotation names the inputs hash it renders. New `status.lastRun.error.kind: blocked`, reason `ApplyJobSucceeded=False/DestructivePlanBlocked`, event `DestructivePlanBlocked` and `captf_jobs_total{result="blocked"}`. The image’s runtime must support `plan -out=` and `show -json ` (see [Image Contract]()). Machine applies are unchanged. - Controller remediation additions (the module contract is unchanged): - New `remediation.healthCheckIntervalSeconds` (60–86400, default 300): with `annotateMachine`, a provisioned machine is refreshed at that interval to sample health, independent of drift. A health sample is one completed refresh or drift Job, not a new state serial. - New condition reasons: `StateReadable=False/StateLost` (a provisioned object’s state is gone or lost its inputs hash; no Job runs), `StateReadable=False/StateLocked` (the state lock is held by something other than the object’s runner) and `ApplyJobSucceeded=False/InputsTooLarge`. `DriftDetected` starts as `Unknown/DriftNotChecked`. - `captf_job_attempts` records the retry number of a successful Job (failures of that op since its last success, plus one); `status.activeJob.attempt` is the Job’s sequence number. - `tfcapi-lint`’s `input/extra` error is replaced by the `input/user-variable-default` warning (a non-contract variable without a default is legitimate when every object is meant to set it through `spec.variables`/`spec.variablesFrom`). ### Changed - Machine `provider_id` going missing after provisioning takes two samples (`machine.md` “`provider_id` (output)”, `common.md` “Out-of-band termination”): the first refresh that reads `null` or `""` sets `InfrastructureHealthy=Unknown/ProviderIDMissing` (new reason), and only the second consecutive one reports `terminated` (`False/InstanceTerminated`) and, with `remediation.annotateMachine`, annotates the Machine for remediation. `spec.providerID` is still kept. - The plan hash is `p2:` instead of `p1:` ([Plan Approval]()): a keyed HMAC, with a per-object key in the Secret `captf-plankey--`, that binds what each change does (an update’s changed attributes with old and new values, the attributes left unknown until apply, and sensitivity changes; a create’s or read’s planned object; a delete’s address and action only), plus output changes, imports and move sources, not only each change’s address and actions. Attributes a change leaves alone are not hashed. Imports and moves always need approval. A plan that changes only outputs now waits for approval, and the new `status.plan.outputChanges` counts the changed outputs. A plan waiting under a `p1:` hash must be approved again with its new hash after an upgrade; a `status.plan` recorded under `p1:` is re-planned automatically. - The contract pages (`README.md`, `common.md`, `cluster.md`, `machine.md`, `machinepool.md` and this changelog) were reorganized for readability: long paragraphs were split into sections, table cells were shortened with their detail moved into a section below each table, and source citations became links. The contract itself did not change: every MUST/SHOULD/MAY requirement, every schema and every skeleton is unchanged. - Fewer Jobs per bring-up and in steady state (the module contract is unchanged; see [Machine Remediation](), [Plan Approval]() and [The Reconcile Lifecycle]()): a TerraformMachine no longer refreshes after a successful apply whose own outputs are valid with a `health.state` other than `pending` or `unknown`; that reading is the post-apply health sample (`status.lastRefresh` is the apply’s finish, and it counts once toward `status.unhealthySamples`). While health is `pending` the refresh backs off 30s, 1m, 2m, 4m, then every 5m (plus the UID jitter) instead of a fixed 30s; TerraformCluster and TerraformMachine gain `status.pendingRefreshes`, the consecutive pending samples since the last other reading or apply (status only; it restarts after `clusterctl move`). A guarded or approved (`--expect-plan`) apply whose plan has no changes skips the apply step and succeeds right after the plan, with zero changes and no `ResourcesChanged` event, so its leases go as soon as it is bookkept. The example bring-up (one cluster, six machines, one no-change re-apply) drops from 14 Jobs to 8. - Controller remediation behavior (the module contract is unchanged): - The `cluster.x-k8s.io/remediate-machine` annotation CAPTF set (marked `captf.io/remediation-requested`) is removed once the instance reads Healthy and the Machine is not being deleted. - The first drift check runs one interval after the last successful apply; drift and health deadlines get a per-object jitter of up to 10%. - A failed apply is retried even when the inputs equal the state’s hash. An image change clears the pinned digest in the durable inputs Secret. - `TerraformMachine.spec.drift` and `spec.defaults.drift` are a `MachineDriftPolicy` with `intervalSeconds` only: machine drift is always reported. The cluster’s `spec.drift.action` defaults to `Report`, not `Remediate`. - Machine jobs and drift policies are merged field by field over the cluster’s `spec.defaults` (env by name, `imagePullSecrets` as a union), instead of replacing them as a whole. `spec.defaults` no longer applies to the TerraformCluster itself. - `TerraformMachine.spec.jobs`, `drift` and `remediation` are mutable; `source` and `identityRef` stay immutable. - `jobs.lockTimeoutSeconds` is optional as a pointer; `jobs.activeDeadlineSeconds` (at most 86400) and `remediation.unhealthyThreshold` default when unset (0), not through a pointer. A policy that sets both `lockTimeoutSeconds` and `activeDeadlineSeconds` must keep the lock timeout below the deadline. `jobs.securityContext` may not set `privileged`, `allowPrivilegeEscalation` or `capabilities.add`. - `TerraformCluster.spec.identityRef` is required; machines fall back to `spec.defaults.identityRef`, then to `spec.identityRef`. `spec.controlPlaneEndpoint` is immutable once it has a host. - `TerraformClusterIdentity.spec.allowedNamespaces: {}` is rejected; write `selector: {}` to allow every namespace. The identity has a status (`conditions`, `namespaces`), filled by the manager. - `status.lastRun.error.tail` is now `summary` (at most 512 bytes): the runner’s summary of the failure, not raw stderr. - `captf_cluster_outputs` is not passed to the cluster role. Previously the text said it was `null`. The generated cluster root declares no such variable, and its tfvars carry no key. A cluster module may leave it undeclared or declare it with `default = null`, as the skeletons do. A declaration without a default fails `validate`. Updated `common.md`, `cluster.md`, and the `cluster-inputs.json` description; the schema still accepts `null`. - Consequence of the inputs hash: the hash accepts only integer numbers. The cluster role’s `exports` value becomes every machine’s `captf_cluster_outputs` input, which is hashed. So a fractional or exponent number anywhere in `exports` (for example `{"ratio": 0.5}`) makes those machines’ inputs unhashable, and they cannot apply. Module authors should export numbers as integers or strings. The controller reports it as `OutputsValid=False/OutputsInvalid` on `exports`. No `tfcapi-lint` check catches this today. The alternative, a canonical float encoding in the hash, was not chosen, because the design calls for integers only and a float that does not round-trip exactly would change the hash silently. - API field names: the drift and job durations are integer seconds, following the Cluster API v1beta2 convention and the kube-api-linter `nodurations` rule. `spec.drift.interval` is now `spec.drift.intervalSeconds`, `spec.drift.refreshInterval` is `spec.drift.refreshIntervalSeconds`, and `jobs.lockTimeout` is `jobs.lockTimeoutSeconds`. Defaults are unchanged (1800, pool-only 60 with a minimum of 15, and 300 seconds). References in these documents were renamed. ### Removed - `jobs.backoffLimit` (Jobs always get `backoffLimit: 0`) and `jobs.ttlSecondsAfterFinished` (Jobs never get a TTL): the controller owns retries, and backoff, digest pinning and conditions are derived from the Jobs it retains. - `spec.drift.refreshIntervalSeconds` (it did nothing), `spec.source.imagePullSecrets` (use `jobs.imagePullSecrets`, which covers the source and runner images) and `spec.source.command` (the image contract fixes `/captf/runtime`; the image’s `/captf/runtime` is now the only executable the runner starts). The durable inputs Secret no longer carries `captf.io/command`, and the inputs hash no longer covers a command: the scheme is now `h2`, so every provisioned mutable object re-applies once after the upgrade. - The five no-op mutating webhooks; only validating webhooks remain. Unused condition reasons `RBACPending`, `StateReadPending` and `CapacityResolving`. ### Fixed - `machinepool.md`’s “Minimal skeleton” gave `failure_domains` and `cluster_failure_domains` a `default = []`; both are always set (non-null) by the controller, so `tfcapi-lint`’s `input/default` check warned on either one, and `--strict` failed. The defaults are removed; the inputs are documented as always set (may be `[]`), matching `modules/noop/machinepool`. Every role’s “Minimal skeleton” (`cluster.md`, `machine.md`, `machinepool.md`) now also references `captf_tags` (an empty `terraform_data` stub), since the skeletons are now actually run through `tfcapi-lint module --role --strict`, which otherwise warns `input/tags-unused` on all three. - `cluster.md`’s “Minimal skeleton” typed `cluster_network`’s attributes with `optional(...)`, while the Inputs table types them as required and the controller always renders all four (an unset CIDR list as `[]`, an unset scalar as `null`, never omitted). Both forms pass `tfcapi-lint`’s `input/type` check, but the skeleton now matches the table. - The `output/provider-id-list-shape` check registers at both `error` (the output is missing or not a list) and `warning` (the expression is not sorted and deduplicated); `reference/tfcapi-lint-cli.md` listed it as two separate rows with the same description and no way to tell them apart. It is now one row whose Severity cell names both. - `machinepool.md` said `provider_id_list`’s order was irrelevant (true for the controller’s own write, which sorts and de-duplicates) without mentioning that `tfcapi-lint` separately checks the output’s own expression for `sort()`/`distinct()`, and the skeleton’s comment pointed at `modules/noop/machinepool/outputs.tf` instead of explaining why. Both are now explained inline, with the pattern shown directly. - `machinepool.md`’s `node_labels` said only that the module “MUST render these into kubelet registration … through the bootstrap/agent config it controls,” without saying how, while `bootstrap_data` is opaque and may be gzipped or Ignition. The realistic options (a `multipart/mixed` cloud-config extra part, with a worked example; Ignition and gzipped payloads needing format-aware handling; or a module declaring only the formats it supports) are now spelled out. No new MUST or SHOULD is added. - `runtime-environment.md` did not say that a provider’s own debug logging (`TF_LOG`) cannot be turned on for a Job: it is dropped like every other `TF_*` name, and `spec.jobs.env` rejects it outright. The page now says so and points at running the pinned image locally against copies of the rendered inputs (the stuck-destroy runbook’s recipe) with `TF_LOG` set, as the supported way to get provider debug output. - `runtime-environment.md` did not say that a network provider mirror or a private registry cannot be configured at run time: `TF_CLI_CONFIG_FILE` is runner-owned and every other `TF_*` name is dropped. The page now states this and points at a filesystem mirror baked into the image (`/captf/providers`) as the supported path. - `machinepool.md` and `common.md` described `health`, `instances` and their controller-side effects individually, without saying how a pool module should combine mixed `instances[*].state` values into one group `health` — the controller does not aggregate them itself. `machinepool.md` “Per-instance state” now states what the controller does (nothing: both outputs are copied independently) and what a module SHOULD do to derive `health` from `instances` (a new SHOULD, not a MUST). - `image-contract.md` recommended the `org.opencontainers.image.*` labels but the reference Containerfiles (`docs/book/src/module-author/examples/Containerfile.terraform` and `.opentofu`) never set them. Both now take `IMAGE_SOURCE`, `IMAGE_REVISION` and `IMAGE_VERSION` build args (empty by default) and set the three labels from them. ## v1alpha1 (provisional freeze) Frozen for implementation: Go types and the controller are built against this version. It may still change until the first real module has provisioned a cluster (the contract is frozen for good only once one real cloud module has created a cluster); every such change is recorded here. ### Scope - `README.md`, `common.md`, `cluster.md`, `machine.md`, `machinepool.md` and `schemas/` in this directory. - [Image Contract]() and the reference Containerfiles (`docs/book/src/module-author/examples/`, `modules/noop/`). - The control-plane guides in [Control-Plane Integration](). ### Changed at the freeze - **Provider mirror wildcard.** The runner writes `include = ["*/*/*"]` and `exclude = ["*/*/*"]`, not `*/*`. Verified with the pinned runtimes Terraform v1.16.4 and OpenTofu v1.12.6, and separately with OpenTofu v1.11.5: `*/*` matches providers on the default registry host only, so a provider from any other host fell through to `direct`; `*/*/*` matches every host. - **Remediation wording** (`machine.md` “Health → remediation”). Remediation happens for Machines whose owner acts on `MachineOwnerRemediated`: a MachineSet or a control-plane provider that implements remediation (KubeadmControlPlane, RKE2ControlPlane), not only a MachineSet or KubeadmControlPlane. ### Added at the freeze - **Reserved paths.** `/captf/bin` (the injected runner) and `/captf/config` (the per-run Secret) are mount points; an image MUST NOT ship anything there. Added to the image contract’s “Fixed paths” table and the `tfcapi-lint image` checklist. - Machine-checkable JSON Schemas (draft 2020-12) in `schemas/`: `definitions.json`, `cluster-inputs.json`, `cluster-outputs.json`, `machine-inputs.json` and `machine-outputs.json`, with a valid and an invalid example for each. The schemas change nothing in the contract; they encode it. - Inputs are closed at every object level (module inputs are exactly the contract inputs); outputs are open (extra, non-contract outputs are ignored). - Outputs are validated after the controller’s normalization: `""` for an ID such as `provider_id` becomes `null` first, so the schema requires 1–512 characters for a non-null `provider_id`. - Limits come from the CAPI v1.14.2 API markers, including three the role documents do not spell out: `service_domain` 1–253 characters, at most 100 CIDR blocks of 1–43 characters each, `api_server_port` 1–65535. - Not expressible in the schemas and enforced by the controller: unique failure-domain names, the canonical address order, and a machine’s `failure_domain` output equal to its input when one was requested. - The image contract’s reference Containerfiles create `/captf/providers` before running `providers mirror`. Neither `terraform providers mirror` (v1.16.4, in the reference build) nor `tofu providers mirror` (checked with OpenTofu v1.11.5 on the build host) creates the target directory when the module requires no providers, so the previous recipe failed with `lstat /captf/providers: no such file or directory` for a provider-less module (found building the reference images). ### Decisions - **Operator decisions:** native pool autoscaling, runner Secret access accepted and documented (see [Security Model]()), static credential Secrets only in v1. - Every candidate CAPI mapping considered was decided; the adopted ones are in the role documents (e.g. address order, `interruptible`, cluster `kubernetes_version`, `control_plane_initialized`, endpoint provenance, empty `dataSecretName`, `captf.io/template`). - The role documents’ claims were checked against CAPI source and resolved. Resolutions for the items raised in [the RKE2ControlPlane guide](): - binary bootstrap data: resolved by always-base64 `bootstrap_data`; - providerID under RKE2: no contract change; the machine module makes the kubelet `provider-id` match (CCM or kubelet extra args), as the guide and checklist say; - a second listener on 9345: a cluster-module convention, not a contract field, since CAPI has nowhere to put it; - the `InternalIP` guarantee: no contract change; no `tfcapi-lint` check exists for it in `internal/lint`; - the version string: stays verbatim; modules strip `+rke2rN` (as [`machine.md`]() says), because normalizing would change the inputs hash; - remediation wording: fixed above; - CAPRKE2 on CAPI v1.14: needs a compatibility smoke test against CAPI v1.14.2 before RKE2ControlPlane is relied on in production; - a nondeterministic join target: upstream behavior, documented in the guide. - The contract is authoritative on `captf_contract`: every role receives it (`common.md`), and the renderer renders it for every role. - Proposals considered and **not** adopted for `common.md`, `cluster.md`, `machine.md` and `machinepool.md` (extra tags, `captf_identity`, `health.observed_at`, a `resources` inventory output, cluster `tags`/`dns_zone`, machine `image`/`instance_type`/`disk`/`ssh_authorized_keys`/`machine_uid`, `failure_domain_attributes`, `bootstrap_data_hash`, `instance_id`, pool `rolling_update` and `ready_replicas`): none loosens the closed input schemas (`schemas/`) or is needed by a shipped module; several are fixed in the module instead because the CRDs have no free-form per-object variables outside `spec.variables`/`spec.variablesFrom`. # Runtime Environment This page describes what a module sees once CAPTF actually runs it: the working directory, the environment, the provider mirror, the commands the runner runs and in what order, and how a failure is reported. It complements [Image Contract](), which is the static shape of the image itself. For every operation, the runner replaces the image’s `ENTRYPOINT`/`CMD` with its own binary, checks the image’s layout (below), prepares a working directory, then runs `/captf/runtime` (the image’s `tofu` or `terraform`) as a child process, one command at a time, streaming its output to the Job’s log. ## Working directory and the generated root Before running anything, the runner checks that `/captf/module` holds at least one `.tf`, `.tf.json`, `.tofu` or `.tofu.json` file at its top level and that `/captf/runtime` is executable; either failure stops the Job with `error.kind: image-layout` before any command runs. It then builds a **generated root** at `/captf/work/root` from the files mounted read-only at `/captf/config` (the per-run Secret): `main.tf.json`, which declares a partial `kubernetes` backend (`init`’s `-backend-config=` flags complete it) and calls the image’s module as `module "role" { source = "../../module" }` (resolving to `/captf/module`, a local path, so `init` never fetches the module from anywhere), and `terraform.tfvars.json`, the rendered contract and user variables. A restore’s root is the backend block alone: it calls no module and declares no variable, since `state push` and `state list` need neither. Every command in this page runs with `/captf/work/root` as its working directory; the runner never uses `-chdir`. What the module actually receives through these two files — the contract inputs, user variables and where each value comes from — is on [Job Inputs](). > [!NOTE] > > **The runner never invokes a shell** > > It execs `/captf/runtime` directly. A module’s `local-exec` provisioner needs a shell in the image, or its own `interpreter`, to run at all: the reference images on [Image Contract]() ship one, but a `distroless/static` final stage does not. The image’s root filesystem is read-only by default (`readOnlyRootFilesystem: true`); the only writable paths are `/captf/work` (this generated root, the plan files and `TF_DATA_DIR`) and `/tmp`. A provider that writes anywhere else needs the module or the image to relocate it, typically through an environment variable such as a cache directory setting. ## Environment set and dropped The runner starts from the container’s own environment (the identity’s `envFrom`, the image’s `ENV`, and `spec.jobs.env`) and rewrites it before running any command: `HOME` is always forced to `/captf/work` and `TF_DATA_DIR` to `/captf/work/.terraform` (note that the working directory of every command is `/captf/work/root`, one level below); `TMPDIR` is kept if already set (the Job sets it to `/tmp`) and otherwise defaults to `/captf/work/tmp`; and every other `TF_*` and `KUBE_*` variable is dropped — logged by name, never by value — except the three the Job itself sets (`TF_IN_AUTOMATION`, `TF_INPUT`, `KUBE_NAMESPACE`). This stops a leaked `TF_WORKSPACE`, `TF_CLI_ARGS_*`, `TF_VAR_*`, `TF_LOG` or `KUBE_*` value from moving state, rewriting a step’s flags, overriding a rendered input, or logging provider traffic and credentials. See [Job Environment]() for the full variable list, a worked example of what is kept and dropped, and the names `spec.jobs.env` may not set at all. > [!WARNING] > > **There is no way to turn on provider debug logging in a Job** > > `TF_LOG` falls under the dropped `TF_*` names above; `spec.jobs.env` cannot set it either, since any name starting with `TF_` or `KUBE_` is silently left out of the Job rather than passed through. A module author debugging a provider needs to reproduce the run outside a Job: pull the pinned image, and run its `/captf/runtime` binary by hand against a copy of the rendered root and inputs, with `TF_LOG` set in that shell. [The stuck-destroy runbook]() walks through getting those inputs (the durable inputs Secret’s `main.tf.json`/`terraform.tfvars.json`, the image digest, and the identity’s credentials) and invoking the image’s binary directly; the same recipe works for a debug run, add `-e TF_LOG=DEBUG` (or `TRACE`) to the `docker run`/`podman run` invocation. ## Provider mirror and CLI configuration When the image ships `/captf/providers`, the runner writes `/captf/work/cli.tfrc`: cli.tfrc ```hcl provider_installation { filesystem_mirror { path = "/captf/providers" include = ["*/*/*"] } direct { exclude = ["*/*/*"] } } ``` and sets `TF_CLI_CONFIG_FILE` to that path, so `init` installs every provider from the image and never contacts a registry; `init` fails clearly instead of silently downloading a provider missing from the mirror. Without `/captf/providers`, the runner sets neither, and `init` falls back to the runtime’s default direct installation, which needs registry egress. `TF_PLUGIN_CACHE_DIR` is never set: it must not coincide with a filesystem mirror. > [!WARNING] > > **A network mirror or a private registry cannot be configured at run time** > > `TF_CLI_CONFIG_FILE` is runner-owned, written fresh for every Job as shown above (or left unset), and `spec.jobs.env` cannot override it or any other `TF_*` name (see “Environment set and dropped”); there is no field that lets an object or its defaults point `init` at a `network_mirror` block or a private registry host’s credentials. The supported path is a filesystem mirror baked into the image at `/captf/providers` (see [Image Contract]()): build it from whatever upstream, mirror or private registry the image’s build pipeline can reach, and ship the result. There is no run-time equivalent. ## Commands, and their order `version -json` runs first, once per Job, to record the runtime version for the result document; its outcome is informational and never fails the Job. `init`, `plan`, `apply` and `destroy` (including `-refresh-only`) carry `-input=false -no-color -lock-timeout=s` (`init` also carries a `-backend-config=` flag for each value the generated root’s partial `kubernetes` backend needs); `validate` and `show` carry `-json -no-color` instead; `state push` carries only `-force -lock-timeout=s`; and `state list` and `force-unlock` take neither. Every command runs against the generated root in order, and a non-zero exit (outside the codes a step accepts, such as `plan`’s 2 for “changes present”) stops the Job at that step. When the Job carries a stale lock ID to clear, `force-unlock -force ` runs immediately after `init` (force-unlock needs an initialized backend); it accepts an already-unlocked state as success, since a retried pod or another holder may have cleared it already. | Operation | Commands, in order | | --- | --- | | Apply (machine and machine-pool roles) | `init` → `validate -json` → `apply -auto-approve -var-file=` | | Apply (cluster role: always guarded, and re-planned against the approved hash under `applyPolicy: Manual`) | `init` → `validate -json` → `plan -detailed-exitcode -var-file= -out=` → `show -json ` (only if plan exited 2) → `apply ` | | Destroy | `init` → `destroy -auto-approve -var-file=` | | Refresh | `init` → `apply -refresh-only -auto-approve -var-file=` | | Drift check | `init` → `apply -refresh-only -auto-approve -var-file=` → `plan -detailed-exitcode -refresh=false -var-file= -out=` → `show -json ` (only if plan exited 2) | | Plan preview (cluster role, `applyPolicy: Manual`, before approval) | `init` → `validate -json` → `plan -detailed-exitcode -var-file= -out=` → `show -json ` (only if plan exited 2) | | Restore | `init` → `state push -force ` → `state list` | `` is `terraform.tfvars.json`, in the generated root. An apply whose plan exits 0 (no changes at all, not even to outputs) skips its `apply` step entirely: the Job ends successfully without applying anything. `show -json`’s output is never written to the Job’s log, since a plan document carries every input value; every other command’s output streams to the log as it runs. `validate`’s JSON diagnostics are the exception: they are captured for the failure summary below *and* still written to the log, since they carry no input values. `state list`’s output (resource addresses only) is logged too. Approving a destructive plan or a Manual-policy preview is covered in [Plan Approval](); the counts, add/change/ destroy resource lists and the drift check itself are covered in [Drift and Health](). ## Lock timeout and stop timeout `init`, `plan`, `apply`, `destroy` and `state push` all carry `-lock-timeout`, defaulting to 300 seconds and overridable per object with `spec.jobs.lockTimeoutSeconds` ([Tuning Jobs]()). On SIGTERM — a deletion, a drain, or the Job’s `activeDeadlineSeconds` (default 3600 seconds) — the runner sends `/captf/runtime` SIGTERM, never its provider plugins, and gives it up to 570 seconds (the pod’s 600-second termination grace period, less a margin the runner needs to write the result) to finish in-flight provider calls, write state and release the backend lock before SIGKILL. A run stopped this way reports `error.kind: interrupted`, not a module failure. ## What a failed step reports A run’s `status.lastRun.error` never carries raw process output. Its `kind` is one of `image-layout`, `step`, `interrupted`, `blocked` (a guarded apply stopped before a plan that deletes or replaces a resource) or `plan-changed` (an approved apply whose new plan no longer matches the approved hash); `step` names the command that failed, when there is one. A `blocked` or `plan-changed` result changes nothing. A `plan-changed` result also carries a plan summary in `status.plan`; a `blocked` result’s summary appears in the object’s condition message instead. Both are covered in [Plan Approval](). The failure summary itself is built from the failing step’s own output, never a raw stderr dump: for `validate`, from its `-json` diagnostics; for every other step, from the `Error:` diagnostic header lines in its stderr (ANSI escape codes stripped first, since a provider or a `local-exec` child is not bound by `-no-color`). When neither yields a line, it falls back to `step exited ; see the Job's logs`. The summary is capped at 512 bytes; the process exit code is the failing step’s own code, or `1` when the runner itself could not start or was killed. The full stderr always reaches the Job’s own pod log, uncapped, whether or not it contributed to the summary. ## Output size limits | Limit | Value | Applies to | | --- | --- | --- | | Failure summary | 512 bytes | `status.lastRun.error.summary`, above | | Termination message | 4096 bytes | The whole result document; the kubelet truncates a longer one, so the runner drops fields in stages (resource-change counts, then the error tail, then plan and drift resource lists, then step history) to fit, keeping the plan’s hash and counts last | | Plan resources listed | 50 | `status.plan.resources`; a plan with more sets `truncated: true` | | Drift resources listed | 20 | The drift check’s resource address list | | Stderr kept in memory per step | 64 KiB | The tail a step’s failure summary is built from; the log itself is not truncated | > [!NOTE] > > **See also** > > - [Image Contract]() > - [Job Environment]() > - [Job Inputs]() > - [Plan Approval]() > - [Tuning Jobs]() # Module Design Patterns The [module contract]() says what a module must do. This page says how to design modules that behave well under CAPTF: where each kind of infrastructure belongs, how to shape exports, what the health and `provider_id` outputs must do, how pools autoscale, what to expect from `import` and `moved` blocks, how to keep destroys safe, how to handle secrets, and how to test. It links to the contract pages instead of restating them. Read those for the rules; read this for the reasons. ## Put shared and destructive infrastructure in the cluster module A `TerraformCluster` is the **only** kind whose applies can wait for a person. `applyPolicy: Manual` gates every change after the first, and the destructive-plan guard stops any apply, under `Automatic`, whose plan deletes or replaces a resource. A machine’s apply never waits, and a pool’s apply waits only when it renders a changed set of cluster exports and its plan deletes or replaces something. See [Approvals and Gates]() and [What is guarded](). So decide where a resource lives by what it costs to get wrong: | Resource | Put it in | Why | | --- | --- | --- | | Networks, subnets, security groups, load balancers, DNS, the control-plane endpoint | The **cluster** module | Shared by every machine, and losing one is expensive: changes are gated | | One node, its disk and its address | The **machine** module | One instance per machine, and replaced rather than changed | | A scaling group and its launch template | The **pool** module | Reconciled to a count; its own changes are ungated, except a destructive plan after an exports change | | Anything an operator should review before it changes | The **cluster** module | It is the only place a review can happen | If a resource is shared **and** a machine or pool module would need to change it, you have the design wrong: change it in the cluster module and publish what machines need as an export. ## Design exports for stability The cluster’s `exports` output reaches every machine and pool as the input `captf_cluster_outputs`, verbatim. See [`exports`](). What happens when it changes depends on the kind: | Kind | An export change | | --- | --- | | `TerraformMachine` | Nothing. The first apply’s inputs are stored and re-fed on drift and destroy; a later change is ignored | | `TerraformMachinePool` | The inputs hash changes and every pool re-applies. The apply is **guarded**: if its plan deletes or replaces anything, the change is held until approved, and the pool keeps applying the last exports meanwhile | The consequence is the rule: **treat exports as a stable interface.** - **Publish identifiers, not state of the world.** Export a load balancer’s target group id or a subnet id, which change only when you replace the thing. Do not export a timestamp, a count of machines or an address that changes on its own. - **Change an export only when you mean to roll every pool.** Under `Manual` the cluster apply that changes it waits for approval (an output-only change needs one), so the person approving sees the change. Each pool then applies it. A pool apply that only creates or updates in place follows without a prompt; one that deletes or replaces is held until its own approval. - **Export integers, strings, booleans, lists and objects, but not fractions or exponents.** The controller hashes exports, and a number with a fractional part or an exponent is rejected with `exports contains a fractional or exponent number; export numbers as integers or strings`. - **No secrets.** The generated root marks `exports` sensitive, so a secret there does not fail the plan, but secrets belong in the identity. See [Sensitive values](<#sensitive-values>). - **Know what a pool replaces.** A pool module that reads an export and replaces a resource when it changes is held for approval, but rotations that replace resources (an instance configuration, say) matter after a failed or partly applied change, when every apply of the pool is guarded. Put `lifecycle { prevent_destroy = true }` on the resources a bad export could destroy; the apply then fails even if approved. See [Cluster outputs reach pools and machines]() and [Machine pools](). ## Health and `provider_id` A machine reports itself through two outputs, and the controller is deliberately slow to believe the worst. - **`health`** is an object with `state` (`pending`, `running`, `degraded`, `stopped`, `terminated` or `unknown`), `healthy`, an optional `message` and `reasons`. `running` with `healthy: false` is `InstanceUnhealthy`; `pending` is `InstancePending`; the others map one to one. The message and reasons reach the condition’s message, so write them for an operator. The mapping is in [`InfrastructureHealthy`](). - **`provider_id`** is written once and then immutable. Return it as soon as the instance exists, and return **the same value every time**: a later change makes `OutputsValid=False`/`ProviderIDChanged`. It must equal the Node’s `spec.providerID` exactly. - **The two-sample rule.** If `provider_id` becomes `null` after the machine was provisioned, the first reading sets `InfrastructureHealthy=Unknown` with `ProviderIDMissing`. It is reported terminated only if a **later** refresh is also `null`. Reconciling the same output again is not a second sample. The rule exists so that a transient read, such as an eventually consistent cloud API, does not trigger remediation of a healthy machine. Design for it: when an instance is gone, return `null`, and keep returning `null` rather than a placeholder. > [!TIP] > > **Do not invent a healthy state** > > A module that cannot tell should report `unknown`. The health timeline and the provisioned rule are in [Machine role](), and the remediation that acts on unhealthy machines is in [Machine Remediation](). ## Pool autoscaling A pool autoscales through two annotations on the `MachinePool`: `cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size` and `-max-size`. Both must be present, parse as non-negative integers and satisfy `min <= max`. If neither is set, autoscaling is off and `spec.replicas` is the capacity. If one is missing, unparsable or `min > max`, autoscaling is off and `AutoscalingActive=False`/`AutoscalingAnnotationsInvalid` names the problem. When it is on: - The module receives `autoscaling = {enabled, min, max}`, and its `replicas` input is the **observed** capacity, clamped into the range, not `spec.replicas`. - The module must scale **in the cloud**, within `[min, max]`, and **ignore changes** to the desired capacity, or each reconcile fights the cloud autoscaler. `tfcapi-lint` warns about a module that uses `var.autoscaling` without an `ignore_changes` on the desired capacity (`pool/autoscaling-ignore-changes`). - The controller writes the observed `replicas` output back to `MachinePool.spec.replicas` and sets the annotation `cluster.x-k8s.io/replicas-managed-by: captf`. If another controller already owns that annotation with a truthy value, CAPTF leaves `spec.replicas` alone and reports `ReplicasManagedExternally`. Choose one owner. - The Kubernetes Cluster Autoscaler is **not** supported on pools. See [Machine Pools]() and the [`autoscaling` input](). ## `import` and `moved` blocks CAPTF does not block `import` and `moved` blocks. The runner reads a plan’s imports and moves, counts them separately from adds, changes and destroys, shows them as `address (import)` or `address (update, move)`, and folds them into the plan fingerprint. What that means in practice: - **On the cluster under `Manual`, they need approval, except on the first apply.** A plan that only imports or moves is not empty, so it waits for approval like any change. The first apply of an object with no state is not gated, so a wrong id on a first-apply import adopts the wrong resource unattended. A plan with no changes, imports, moves or output changes needs none. - **Under `Automatic`, they are not blocked.** The destructive-plan guard looks only for deletes and replaces, and an import or move is neither. - **On machines and pools they apply unattended**, like everything else there. - **They are not drift.** A drift check ignores imports and moves. > [!WARNING] > > **An `import` block with a wrong id adopts the wrong thing** > > Under `Manual`, read the plan before you approve. Also, no CAPTF documentation or test yet covers `import` blocks inside a module, so treat the pattern as untested. Drive the ids from a variable that defaults to empty, so the block can stay in the module. An object that applied before and lost its state cannot run a module apply at all, so adoption after a state loss needs a recreated object or a rebuilt state: see [Total State Loss and Import](). ## Destroy safety and replaceable machines - **A machine is replaced, never changed.** A `TerraformMachine`’s source, identity and variables are immutable; Cluster API rolls machines by creating a new one and deleting the old. Write a machine module so that creating it needs nothing from a sibling, and deleting it harms nothing shared. Keep per-machine state per machine. - **Destroy runs from stored inputs.** A destroy renders from the durable inputs Secret, pinned to the image of the last apply, so it does not depend on the current Cluster or bootstrap Secret. Do not write a module whose destroy needs a value only the live Cluster has. See [The destroy Job](). - **A destroy that fails retries forever.** A resource that cannot be deleted (a disk still attached, a security group still in use) stalls the deletion. Order dependencies so Terraform deletes in the right order, and see [Stuck Destroy]() for the way out. - **`prevent_destroy` is plain Terraform.** On the load balancer behind the control-plane endpoint it turns an approved replacement into a failed Job instead of a lost cluster. The same setting makes the cluster’s own destroy fail, so deleting a cluster means lifting it first. Use it where the loss would be worse than the stalled delete, and write down how to lift it. - **Do not depend on the controller’s order.** The cluster destroys last, after its machines and pools are gone, but do not rely on timing: a module should tolerate a destroy that finds a resource already gone. ## Sensitive values - **Mark secrets `sensitive = true` in the module.** A variable that arrived from a Secret is declared sensitive in the generated root, but a declaration in your module makes the value sensitive whatever its source. An output that refers to a sensitive value must itself be `sensitive = true`, or `plan` fails. See [Common inputs](). - **Never put credentials in the module.** They come from the identity, as environment variables and files. `tfcapi-lint` warns on a provider block that sets a credential-named argument to a literal (`module/provider-config`). - **What the controller keeps out.** Status and events carry counts and addresses, never values. The runner redacts the values it knows about (credentials, sensitive variables, bootstrap data, sensitive plan values) from the Job’s result and events. It cannot redact a secret a module derives and prints. See [What CAPTF keeps out of status, events and logs](). - **Bootstrap data is sensitive.** `bootstrap_data` carries join tokens. Declare it `sensitive = true`; `tfcapi-lint` warns when you do not (`input/sensitive`). ## Test the module 1. **Start from the noop module.** It creates nothing in a cloud, and shows the contract end to end: [Your First Module]() builds one. The published images and source are in [`captf-io/noop-modules`](). 2. **Lint the module** with `tfcapi-lint`, in strict mode, in CI. It needs neither `terraform` nor `tofu`: ```sh tfcapi-lint module --role cluster --strict ./cluster tfcapi-lint image --role cluster --strict ``` The checks that most often catch design mistakes: `input/required`, `input/type`, `output/required` and `output/health` (the contract shape); `output/provider-id-list-shape` and `pool/autoscaling-ignore-changes` (pools); `output/endpoint-never-set` (a cluster that never sets an endpoint); `module/backend` and `module/cloud` (the controller owns the backend); `module/source-escape` and `module/tofu-shadow` (code the linter would not see); and `module/provider-config` (literal credentials). The full list is in [tfcapi-lint CLI](). 3. **Run it with both runtimes.** A module that validates under Terraform can fail under OpenTofu, or the reverse; `module/tofu-shadow` catches one cause. Initialize and validate under each. 4. **Try the lifecycle**, in a scratch management cluster: the first apply, an input change, a drift check, a machine delete and a destroy, and a state restore. Check each one leaves the object `Ready` and that the destroy leaves nothing behind. See [Disaster Recovery]() for a drill. > [!NOTE] > > **See also** > > - [Module Contract]() and its role pages: [cluster](), [machine]() and [machine pool](). > - [Image Contract](), [Runtime Environment]() and [tfcapi-lint](). > - [Approvals and Gates](). # tfcapi-lint `tfcapi-lint` checks a Terraform or OpenTofu module, and the OCI image built from it, against the CAPTF module contract ([module contract]()) and the [image contract](), without running `init`, `plan` or `apply`. It is for module authors, and for CI pipelines that build module images. It is released alongside the provider, with the same version. ## Install Every release attaches one binary per platform, plus a checksum file: | Asset | Platform | | --- | --- | | `tfcapi-lint-linux-amd64` | Linux x86-64 | | `tfcapi-lint-linux-arm64` | Linux ARM64 | | `tfcapi-lint-darwin-amd64` | macOS Intel | | `tfcapi-lint-darwin-arm64` | macOS Apple silicon | | `tfcapi-lint-windows-amd64.exe` | Windows x86-64 | | `tfcapi-lint-checksums.txt` | SHA-256 of every asset above | ```sh base="https://github.com/captf-io/cluster-api-provider-terraform/releases/download/" curl -fsSLO "${base}/tfcapi-lint--" curl -fsSLO "${base}/tfcapi-lint-checksums.txt" sha256sum --check --ignore-missing tfcapi-lint-checksums.txt install -m 0755 "tfcapi-lint--" /usr/local/bin/tfcapi-lint tfcapi-lint version ``` - `` is the provider release you deploy, for example `v0.1.0`. - `` and `` pick one row of the table above, for example `linux` and `amd64`. On macOS, verify the checksum with `shasum -a 256 -c --ignore-missing tfcapi-lint-checksums.txt` instead. Match the `tfcapi-lint` release to the controller you deploy against: `tfcapi-lint version --json` reports the contract versions it lints against, `["v1alpha1"]`. > [!WARNING] > > **`go install` does not work** > > `go install .../cmd/tfcapi-lint@` fails: the repository is a Go workspace, and `cmd/tfcapi-lint` depends on the `api` module, which only the workspace resolves. Build from source instead (below), or use a release binary. ## Roles Every module implements exactly one role, and `--role` on `module` and `image` is required and takes one of `cluster`, `machine` or `machinepool`, matching the [`TerraformCluster`](), [`TerraformMachine`]() and [`TerraformMachinePool`]() contracts. `tfcapi-lint` checks the module or image against that role’s inputs, outputs and checks only; see [tfcapi-lint CLI]() for which checks apply to which roles. ## Lint a module ```sh tfcapi-lint module --role machine ./machine ``` This reads the `.tf`, `.tf.json`, `.tofu` and `.tofu.json` files under `./machine` directly; it needs neither `terraform` nor `tofu` installed, and never contacts a registry. A clean module prints an empty finding list and an all-zero summary. ## Lint an image ```sh tfcapi-lint image --role machine registry.example.com/acme/machine:v1.0.0 ``` This pulls the image manifest and its layers, and checks the fixed paths and labels the [image contract]() requires, without running the image. `` is a registry reference; `oci:` reads a local OCI image layout instead, such as one written by `podman save --format oci-dir` or `skopeo copy ... oci:`, with no registry or daemon involved. By default it checks the `linux/amd64` platform of a multi-platform image; `--platform os/arch` picks a different one, and `--all-platforms` checks every platform the image publishes. ### Registry credentials `tfcapi-lint image` authenticates the same way `docker` and `podman` do, through go-containerregistry’s default keychain: it reads `~/.docker/config.json`, or `$DOCKER_CONFIG/config.json` when that variable is set; if neither exists, it falls back to a Podman-style config at `$REGISTRY_AUTH_FILE` or `$XDG_RUNTIME_DIR/containers/auth.json`. With none of those present, the pull is anonymous. Log in with `docker login` or `podman login` against the registry before linting a private image; `--insecure` allows a plain-HTTP registry for a local or air-gapped registry that has none. ## Strict mode and allowed warnings `--strict` treats a warning the same as an error for the exit code, so a module or image that is merely clean today does not silently pick up new warnings later. `--allow-warning ` (repeatable, or a comma-separated list) downgrades one check ID’s warnings to informational findings; it never touches errors. Use it for a deliberate, reviewable exception, for example a module whose provider cannot tag anything: `--allow-warning input/tags-unused`. Run with `--strict` by default, and add `--allow-warning` only for checks you have decided not to act on. `--json` prints a report with a `findings` array (`id`, `severity`, `file`, `line`, `message`) and a `summary`, instead of one line of text per finding. See [tfcapi-lint CLI]() for every flag, and [tfcapi-lint CLI: checks]() for every check ID, its severity and the roles it applies to. ## Exit codes `tfcapi-lint` uses its exit code to signal a CI step’s pass or fail; see [tfcapi-lint CLI: exit codes]() for the full list. In short: | Code | Meaning | | --- | --- | | `0` | Clean. | | `1` | At least one error (or, under `--strict`, at least one warning). | | `2` | The module could not be parsed or the image could not be pulled. | | `3` | A usage error. | ## In CI Lint the module before building the image, then lint the built image before pushing it: the source is checked before the build, and the built layout after it. CI step ```sh tfcapi-lint module --role machine --strict ./module podman build -t "$IMAGE" . podman save --format oci-dir -o "$RUNNER_TEMP/image" "$IMAGE" tfcapi-lint image --role machine --strict "oci:$RUNNER_TEMP/image" podman push "$IMAGE" ``` To check exactly what was pushed, for example a multi-platform index built and pushed by a separate step, lint the pushed reference instead of the local layout: ```sh tfcapi-lint image --role machine --strict --all-platforms "$IMAGE" ``` ## Building from source `make release-lint-snapshot` builds all five release assets and the checksum file into `dist/` with GoReleaser, stamped with a snapshot version; `make release-lint` does the same from the current git tag. See [Releasing]() for how these assets reach a GitHub release. > [!NOTE] > > **See also** > > - [tfcapi-lint CLI]() for every flag, every check ID and the exit codes. > - [Module contract]() and [image contract]() for what the checks enforce. > - [Runtime environment]() for what a module sees once CAPTF actually runs it. # Control-Plane Integration CAPTF’s cluster and machine modules are infrastructure only: they create networks, load balancers and instances, and report state back through the [contract](). Bringing up Kubernetes on top of that infrastructure is the job of a Cluster API control-plane provider — KubeadmControlPlane (KCP) or RKE2ControlPlane (RCP) — plus its bootstrap provider. Those providers read specific fields from your modules and write specific fields back, on a specific schedule. Getting the details wrong produces a cluster that hangs rather than one that fails loudly. This guide set distills the verified requirements into what you need as a module author: - **KubeadmControlPlane** --- Everything specific to KubeadmControlPlane / the kubeadm bootstrap provider (CABPK). - **RKE2ControlPlane** --- Everything specific to RKE2ControlPlane / the RKE2 bootstrap provider (CAPRKE2). - **Requirements Checklist** --- Every port, health check and ordering rule from both guides, organized by topic, for a build-time or review-time reference. Both guides cite CAPI/CAPRKE2 source paths and the contract’s own module role pages: [the cluster role](), [the machine role]() and [the common contract](). Where this guide set and the contract state the same requirement, the contract’s wording is authoritative. ## The shared creation sequence Both control-plane providers drive the same shape of sequence against a CAPTF cluster module and one or more CAPTF machine modules. The sequence below is the kubeadm case; RKE2’s differences are called out inline and detailed in [`rke2.md`](), in [what it reads from your modules]() and [lifecycle constraints](). 1. A user (or a ClusterClass topology) creates the `Cluster`, `TerraformCluster`, control-plane object (`KubeadmControlPlane` or `RKE2ControlPlane`) and `TerraformMachineTemplate`. 2. The CAPI Cluster controller reconciles the `infrastructureRef`, setting its owner reference to the `TerraformCluster`. 3. The `TerraformCluster` controller runs the cluster module’s first apply. At this point `control_plane_endpoint` is rendered `null` and `control_plane_initialized` is rendered `false`. The module returns `control_plane_endpoint`, `failure_domains`, `exports` and `health`. 4. The controller copies `control_plane_endpoint` onto `TerraformCluster.spec.controlPlaneEndpoint` and `failure_domains` onto `TerraformCluster.status.failureDomains`; CAPI in turn copies them onto `Cluster.spec.controlPlaneEndpoint` and `Cluster.status.initialization.infrastructureProvisioned = true`. 5. The control-plane provider gates all further work on: a valid `Cluster.spec.controlPlaneEndpoint` **and** `infrastructureProvisioned = true`. Neither KCP nor RCP creates a single Machine before both are true (see [what KubeadmControlPlane reads]() and [what RKE2ControlPlane reads]()). 6. For control-plane machine `i = 1..N`, in strict order — the provider creates machine `i+1` only after machine `i` has a `nodeRef`: 1. The control-plane provider creates the bootstrap config (init for `i=1`, join for `i>1`) and the `Machine` (labeled `cluster.x-k8s.io/control-plane`) plus a `TerraformMachine` cloned from the template. 2. The bootstrap provider writes the bootstrap Secret (`value`, `format`); `Machine.spec.bootstrap.dataSecretName` is set. 3. The `TerraformMachine` controller reconciles machine `i`, gated on: the cluster’s `infrastructureProvisioned`, the cluster’s `exports` being readable, and the bootstrap Secret existing. 4. The machine module’s apply runs once (immutable): it receives `bootstrap_data` (base64), `control_plane = true`, `captf_cluster_outputs` (the cluster’s `exports`), `failure_domain` and `kubernetes_version`. In this same apply the module creates the instance **and** registers it in the control-plane LB backend, or does neither of the latter under [Pattern C](<#pattern-c-kube-vip-style-vip>) below. It returns `provider_id`, `addresses`, `failure_domain`, `interruptible` and `health`. 5. The `TerraformMachine` controller writes `spec.providerID` and `status.addresses` on the `TerraformMachine`, and marks it provisioned and `Ready`. The CAPI Machine controller then copies both onto `Machine.spec.providerID` and `Machine.status.addresses`. 6. cloud-init on the instance runs `kubeadm init` (`i=1`) or `kubeadm join` (`i>1`); post-init phases go through the endpoint (hairpin — see [kubeadm.md’s networking section]()). Under RKE2’s default registration method the join target is the endpoint host itself, but the join cannot proceed until at least one **Ready** control-plane Machine already exists, since `RCP.status.availableServerIPs` stays empty without one (see [what RKE2ControlPlane reads]() and [rke2.md’s lifecycle constraints]()). 7. The Node registers with `spec.providerID` set (by a cloud-controller-manager or `kubelet --provider-id`); the Machine controller matches `nodeRef` by an exact `providerID` match. 8. When `i = 1` and the control plane finishes initializing, `Cluster.status.initialization.controlPlaneInitialized` flips `true`. This triggers exactly one cluster re-apply, for resources gated on a live workload API server. 7. Workers follow the same machine path (`MachineDeployment` → `MachineSet` → `Machine` → `TerraformMachine`, `control_plane = false`), gated on `ControlPlaneInitialized` because the bootstrap provider writes worker join data only after that. A worker fleet backed by a `MachinePool` instead goes straight from `MachinePool` to `TerraformMachinePool` — one infrastructure object for the whole group, with no per-replica `Machine` or `TerraformMachine`; see [Machine Pools](). ## LB-membership patterns The contract requires a control-plane machine module to register (and, on destroy, deregister) its instance with the control-plane load balancer (see [machine.md’s control-plane machines]()), but leaves how up to the module. Three patterns are sanctioned here. ### Pattern A: attachment resource in the machine module’s state The machine module creates the instance and, in the same apply, an explicit attachment/registration resource (for example an LB target-group attachment) pointed at the target-group or backend-pool id it received through `captf_cluster_outputs` (see [common.md’s outputs]()). Because the attachment lives in the machine module’s own state, `destroy` deregisters it as an ordinary part of tearing that state down. This is the pattern the contract documents directly: registration is “done in the machine module’s own Terraform state, not the cluster module’s” ([machine.md’s control-plane machines](), citing `capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645`). The instance MUST be registered before `kubeadm init`/`kubeadm join` finishes (KCP) or before the Machine is Ready (RKE2) — see [kubeadm.md’s module checklist]() and [rke2.md’s module checklist](). Typical fit: clouds whose load balancer exposes an explicit target-group/backend-pool attach API (for example AWS ALB/NLB target-group attachments, Azure Load Balancer backend-pool membership). > [!NOTE] > > **Untested design guidance** > > This is design guidance, not yet tested against a real control plane. ### Pattern B: selector/tag-based backend pools The load balancer derives its backend membership itself, from a tag or label selector query, rather than from an explicit per-instance attach call — “target group by tag/label”. The machine module’s only job is to apply the right tags to the instance (via `captf_tags`, see [common.md’s inputs](), or an additional module-defined tag); the load balancer’s own membership scan does the rest, and removing the instance (destroy) removes it from the pool. The same ordering requirement applies as under Pattern A: the instance MUST be a member of the backend before `kubeadm init`/`join` finishes on it (KCP) or before the Machine is Ready (RKE2) — see [kubeadm.md’s module checklist]() and [rke2.md’s module checklist](). Typical fit: clouds or load balancers whose backend pool is defined by a selector/tag query against an instance group or autoscaling group, rather than an explicit attach call. > [!NOTE] > > **Untested design guidance** > > This is design guidance, not yet tested against a real control plane. ### Pattern C: kube-vip-style VIP No separate load-balancer resource is created by the cluster module at all; a virtual IP is run by the control-plane nodes themselves (implicit via kube-vip or similar), so there is “nothing LB-wise” for the machine module to do. The cluster module still MUST emit a valid `control_plane_endpoint` before any Machine is created — the same gate applies regardless of how the endpoint is realized (see [what KubeadmControlPlane reads]() and [kubeadm.md’s networking section](); [what RKE2ControlPlane reads]() and [rke2.md’s networking section]()). Under this pattern the backend-membership and per-backend health-check rows of [`checklist.md`]() do not apply, because there is no separate LB backend to join. Typical fit: bare-metal, on-premises, or otherwise L2-reachable environments without a managed load balancer in front of the control plane. > [!NOTE] > > **Untested design guidance** > > This is design guidance, not yet tested against a real control plane. > [!NOTE] > > **See also** > > - [The module contract]() — the normative cluster and machine roles this guide set builds on. > - [Security model]() — the trust boundary a control-plane Job runs inside. > - [The Kinds]() — how `TerraformCluster` and `TerraformMachine` map to Cluster API objects. # KubeadmControlPlane What KubeadmControlPlane (KCP) and the kubeadm bootstrap provider (CABPK) require of your CAPTF cluster and machine modules, and why this is the case. > [!NOTE] > > **Citation convention** > > Citations below prefixed `capi/` are files in the CAPI v1.14.2 tree (`https://github.com/kubernetes-sigs/cluster-api/tree/v1.14.2`). See also [the shared creation sequence]() and [the requirements checklist]() for every requirement in one place. ## What KubeadmControlPlane reads from your modules - **`Cluster.spec.controlPlaneEndpoint`.** KCP has no endpoint field of its own; it only reads the Cluster’s. It gates all Machine creation on this field being valid (`host != ""` and `port != 0`) — see [networking and load balancer](<#networking-and-load-balancer>). The endpoint can come from the user, your `TerraformCluster` (copied once while the Cluster field is not yet valid), or a control-plane provider; KCP never supplies one itself (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:316,439`; `capi/core/reconcilers/cluster/cluster_controller_phases.go:210-238`). - **`Machine.status.addresses`.** Not read by KCP or CABPK at all. Only the core Machine controller copies it from the InfraMachine; apiserver certificate SANs come from kubeadm’s own node-IP detection, not from CAPI addresses (`capi/core/reconcilers/machine/machine_controller_phases.go:339-346`). - **`providerID` / `nodeRef`.** The Machine controller links a Node to a Machine only on an exact `providerID` match (see [`../contract/v1alpha1/machine.md`]() “Node providerID matching”). KCP then gates almost everything on the resulting `nodeRef`: etcd-member matching, static-pod health, and scale-up/scale-down preflight all require every existing control-plane Machine to have one (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/preflight.go:193-203`; `capi/controlplane/kubeadm/pkg/workload_cluster_conditions.go:136,390,434`). - **Failure domains.** KCP only sees `Cluster.status.failureDomains` entries with `controlPlane == true`; a nil `controlPlane` counts as `false`. This is why the cluster module’s `failure_domains[].control_plane` default of `true` matters (`capi/controlplane/kubeadm/pkg/control_plane.go:171-184`). - **`cluster_network.api_server_port`.** Not read by CABPK or KCP at all in CAPI v1.14.2. The real kube-apiserver bind port is the KubeadmConfig `localAPIEndpoint.bindPort` (default 6443) (`capi/controlplane/kubeadm/pkg/workload_cluster.go:281-291`). Treat `api_server_port ?? 6443` as a documentation-only convention, not a wired-through value (see [networking and load balancer](<#networking-and-load-balancer>)). - **InfraMachineTemplate rotation.** KCP triggers a rollout when a Machine’s InfraMachine carries a `cluster.x-k8s.io/cloned-from-name`/ `-groupkind` annotation different from the current `machineTemplate.spec.infrastructureRef`. Template **content** is never diffed — CAPTF’s template immutability already matches this expectation (`capi/controlplane/kubeadm/pkg/filters.go:183-227`). ## What KCP and CABPK write - **Bootstrap Secret.** Named after the KubeadmConfig (== the Machine name). Keys `value` and `format` are always both written; `format` defaults to `cloud-config` (Ignition requires the alpha feature gate `KubeadmBootstrapFormatIgnition`, off by default) (`capi/bootstrap/kubeadm/reconcilers/kubeadmconfig/kubeadmconfig_controller.go:1405-1437`; `capi/api/bootstrap/kubeadm/v1beta2/kubeadmconfig_types.go:27-36,117-120`; `capi/feature/feature.go:108`). See [`../contract/v1alpha1/machine.md`]() `bootstrap_format` for how CAPTF surfaces this. - **Payload contents.** Control-plane init and join payloads embed the cluster CA, etcd CA, service-account and front-proxy key material as files, uncompressed (`capi/bootstrap/kubeadm/pkg/cloudinit/controlplane_init.go:63`, `controlplane_join.go:61`). See [module checklist](<#module-checklist>) for what this means for your module. - **Labels and hooks.** Every KCP Machine, InfraMachine and KubeadmConfig gets `cluster.x-k8s.io/cluster-name`, `cluster.x-k8s.io/control-plane: ""` and a `pre-terminate.delete.hook.machine.cluster.x-k8s.io/kcp-cleanup: ""` annotation (`capi/controlplane/kubeadm/pkg/desiredstate/desired_state.go:324-340,110-135`). CAPTF’s own metadata-update path already allows these (label/annotation edits are excluded from TerraformMachine’s immutability checks). - **`KubeadmControlPlane.status.initialization.controlPlaneInitialized`.** Latches `true` once the `kubeadm-config` ConfigMap is readable **through the endpoint** — the reason the endpoint’s hairpin reachability matters (see [networking and load balancer](<#networking-and-load-balancer>)) (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/status.go:176-210`). This mirrors into `Cluster.status.initialization.controlPlaneInitialized` and drives one cluster re-apply (`capi/core/reconcilers/cluster/cluster_controller_status.go:80-81,450-551`). - **Conditions.** KCP writes `Initialized`, `Available`, `CertificatesAvailable`, `EtcdClusterHealthy`, `ControlPlaneComponentsHealthy`, `MachinesReady`, `MachinesUpToDate`, `RollingOut`, `ScalingUp`/`ScalingDown`, `Remediating`, `Deleting`, plus per-Machine pod/etcd-member conditions (`capi/api/controlplane/kubeadm/v1beta2/kubeadm_control_plane_types.go:74-421,504-517`). None of these read anything from your module beyond what is described above. ## Lifecycle constraints - **Init ordering.** KCP creates nothing until the Cluster is both `infrastructureProvisioned` and has a valid endpoint; CABPK waits on the same gate. The first control-plane Machine gets init data; every other Machine (control-plane or worker) waits on `Cluster.status.initialization.controlPlaneInitialized` (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:316,413,439`; `capi/bootstrap/kubeadm/reconcilers/kubeadmconfig/kubeadmconfig_controller.go:289-298,352-368,466-495`). - **Join is serialized.** Scale-up preflight requires certificates available, no Machine deleting, and **every** existing control-plane Machine already healthy with a `nodeRef`. A failed preflight requeues after 15s, so control-plane joins are serialized on your apply time plus kubeadm join time (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/scale.go:60-97`; `preflight.go:61-218`). - **Rollout.** `maxSurge` is 0 or 1 (default 1): with 1, KCP scales up then down, one Machine at a time; with 0 it scales down first (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/update.go:103-120`). Rotation is by creating a new `TerraformMachineTemplate` name; edits to an existing template’s content are never detected. - **Deleting blocks everything.** While any control-plane Machine has a deletion timestamp, KCP pauses scale-up, scale-down, rollout and remediation (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/preflight.go:111-122`). A slow or stuck machine-module destroy therefore freezes the whole control plane, not just that Machine. - **Remediation gates and retries.** Post-initialization remediation needs more than one replica, no Machine still provisioning without a Node, no Machine deleting, and preserved etcd quorum; one remediation runs at a time (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/remediation.go:138-313`). `remediation.maxRetry` defaults to unlimited and `retryPeriodSeconds` defaults to `0` (retry immediately) (`capi/api/controlplane/kubeadm/v1beta2/kubeadm_control_plane_types.go:631-673`). See [module checklist](<#module-checklist>) for the template defaults this motivates. - **No built-in control-plane MHC.** KCP does its own remediation and is not configured through a MachineHealthCheck object. Ship a plain MachineHealthCheck selecting `cluster.x-k8s.io/control-plane: ""` (or a ClusterClass `controlPlane.healthCheck`) if you want MHC-driven remediation in addition to KCP’s own (`capi/docs/book/src/tasks/automated-machine-management/healthchecking.md:76-99`). - **Version skew.** `KubeadmControlPlane.spec.version` must be valid semver with a `v` prefix, and an update may move at most one minor version (admission webhook, `capi/controlplane/kubeadm/webhooks/admission/kubeadmcontrolplane.go:322-328,619-662`). - **Cluster deletion.** Workers and MachinePools are deleted first; once only control-plane Machines remain, KCP deletes **all** of them in parallel, with no drain (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:720-813`). Expect concurrent destroy Jobs on cluster teardown. - **Failed create cleanup.** KCP creates the InfraMachine before the Machine object. If KubeadmConfig or Machine creation then fails, KCP deletes the InfraMachine directly, with no Machine owner reference at all (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/helpers.go:160-219`). CAPTF’s delete webhook already allows a delete with no Machine owner reference for exactly this reason (see [`../contract/v1alpha1/machine.md`]() “Delete”) — nothing further is required of your module. ## Networking and load balancer - **Endpoint before any Machine.** KCP creates no Machine at all until `Cluster.spec.controlPlaneEndpoint.IsValid()` (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:438-457`). - **Frontend port.** `control_plane_endpoint.port`, any value 1-65535. - **Backend port.** The kube-apiserver bind port, i.e. KubeadmConfig `localAPIEndpoint.bindPort` (default 6443). CABPK/KCP never read `Cluster.spec.clusterNetwork.apiServerPort`, so treat `api_server_port ?? 6443` as a convention your templates must keep in sync with `bindPort` yourselves (`capi/controlplane/kubeadm/pkg/workload_cluster.go:281-291`). - **Reachability, including hairpin.** The endpoint must be reachable from the management cluster, from every control-plane and worker node (join discovery, kubelet), and from the first control-plane Machine itself. This last point is the hairpin requirement: kubeadm’s `getKubeConfigSpecsBase` sets the server URL of `admin.conf`, `super-admin.conf` and `kubelet.conf` to the control-plane endpoint, not the node’s own local address, so the node that is itself an LB backend must be able to reach the frontend it was just registered behind (see [`../contract/v1alpha1/cluster.md`]() “Hairpin reachability”, citing kubeadm’s `cmd/kubeadm/app/phases/kubeconfig/kubeconfig.go` \~L582-626, kubernetes/kubernetes release-1.33). - **Backend membership.** Your machine module MUST register (and on destroy deregister) each control-plane instance in the LB backend from its own state (`capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645`; [`../contract/v1alpha1/machine.md`]() “Control-plane machines”). The instance MUST be in the backend before `kubeadm init`/`kubeadm join` finishes on it, or `controlPlaneInitialized` never latches (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/status.go:195-207`). - **Health check.** Either plain TCP, or HTTPS `/readyz`/`/healthz` without certificate verification, works; the check must be able to go green with a single backend during init. CAPD’s reference LB uses haproxy `option httpchk GET /healthz` with `check-ssl verify none` against backend 6443 (`capi/test/infrastructure/docker/internal/loadbalancer/config.go:81-83`). - **Security groups / firewall.** Backend port (bindPort) from the LB, from all nodes, and from the management cluster; TCP 2379-2380 and 10250 between control-plane nodes; 10250 from control-plane nodes to workers. KCP itself reaches etcd only through an apiserver pod port-forward, so the management cluster needs no direct etcd or kubelet port (`capi/controlplane/kubeadm/pkg/etcd_client_generator.go:53`; `capi/controlplane/kubeadm/pkg/proxy/dial.go:72-120`). ## Module checklist ### Cluster module - MUST output a valid `control_plane_endpoint` (host and port) no later than the apply that makes the cluster `provisioned`, when no user endpoint is given — there is no other source with KCP (`capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/kubeadmcontrolplane_controller.go:316,439`; see [what KubeadmControlPlane reads](<#what-kubeadmcontrolplane-reads-from-your-modules>)). `tfcapi-lint`’s `output/endpoint-never-set` check warns on a module that never emits one (see [cluster.md’s `EndpointAvailable` condition]() and the [tfcapi-lint CLI reference]()). - MUST keep the emitted `control_plane_endpoint` stable for the life of the object: CAPI never updates `Cluster.spec.controlPlaneEndpoint` after the first valid copy (`capi/core/reconcilers/cluster/cluster_controller_phases.go:217-228`), so replacing the LB behind it (new DNS name or IP) breaks every kubeconfig and the API server certificate SANs. The controller neither detects nor prevents such a replacement; `lifecycle { prevent_destroy = true }` on the resource behind the endpoint turns it into a failed Job instead of a lost cluster ([`../contract/v1alpha1/cluster.md`]() `control_plane_endpoint` output). - MUST create a stable LB or VIP reachable from the management cluster, all nodes, and the control-plane nodes themselves (hairpin, see [networking and load balancer](<#networking-and-load-balancer>)). - MUST use a frontend on `control_plane_endpoint.port` and a backend on `api_server_port ?? 6443`, kept equal to the KubeadmConfig `bindPort` your templates configure (`capi/controlplane/kubeadm/pkg/workload_cluster.go:281-291`; see [networking and load balancer](<#networking-and-load-balancer>)). - SHOULD publish the LB target/backend-pool id in `exports` for machine modules to register against — a CAPTF convention built on the contract’s `exports` field, not a CAPI-mandated shape (`../contract/v1alpha1/cluster.md` `exports`). - SHOULD leave `failure_domains[].control_plane` at its default `true` (`capi/controlplane/kubeadm/pkg/control_plane.go:171-184`; see [what KubeadmControlPlane reads](<#what-kubeadmcontrolplane-reads-from-your-modules>)); a nil value is invisible to KCP. ### Machine module - MUST register the control-plane instance in the LB target/backend-pool from `captf_cluster_outputs`, in the module’s own state, before `kubeadm init`/`join` finishes on it (`capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645`; `capi/controlplane/kubeadm/reconcilers/kubeadmcontrolplane/status.go:195-207`; see [networking and load balancer](<#networking-and-load-balancer>) and one of the three patterns in [README.md’s LB-membership patterns]()). - MUST accept `bootstrap_data` opaquely as base64 and MUST NOT need to parse it, per [`../contract/v1alpha1/machine.md`]() `bootstrap_data`. - Control-plane payloads embed cluster CA and service-account private key material and are not compressed by CABPK (see [what KCP and CABPK write](<#what-kcp-and-cabpk-write>)). Payload size limits (for example a cloud’s user-data cap) and keeping key material out of readable instance metadata are the module’s responsibility: gzip the payload, or stage it in a secret store and pass only a small stub through `bootstrap_data`/user-data. - MUST emit `provider_id` exactly matching the Node’s `spec.providerID`; KCP gates join preflight, etcd-member matching and remediation on the resulting `nodeRef` (see [what KubeadmControlPlane reads](<#what-kubeadmcontrolplane-reads-from-your-modules>)). - `addresses` are informational only for KCP — no KCP or CABPK code reads them. - MUST make `failure_domain` output equal the `failure_domain` input when the input is non-null (`../contract/v1alpha1/machine.md` `failure_domain` output). ### Templates and health checks - Bring-up time for N control-plane replicas is approximately N × (apply + boot + join), because every join is serialized on the previous Machine’s `nodeRef` (see [lifecycle constraints](<#lifecycle-constraints>)). A rollout adds (apply + join + destroy) per replica. - Ship `KubeadmControlPlane.spec.remediation.maxRetry` (e.g. `3`) and a non-zero `retryPeriodSeconds` in your templates: the unbounded defaults turn a control-plane module whose apply always fails into an endless create/destroy loop of cloud resources (see [lifecycle constraints](<#lifecycle-constraints>)). - Ship a control-plane MachineHealthCheck (selector `cluster.x-k8s.io/control-plane: ""`, or a ClusterClass `controlPlane.healthCheck`). Its `InfrastructureReady=False` timeout MUST exceed apply time plus one drift interval — the same rule as any other machine module (`../contract/v1alpha1/machine.md` “Health → remediation”). ## Open questions - **Bootstrap payload size and secrecy.** The contract leaves the gzip/stub strategy to the module (see [module checklist](<#module-checklist>)); there is no CAPTF-side size or secrecy check today. - **Unbounded remediation retry.** Combined with a deterministically failing module, KCP’s default unlimited `maxRetry` causes cloud-resource churn. The [module checklist](<#module-checklist>)’s template-default recommendation is the only mitigation; there is no CAPTF-side guard against it. > [!NOTE] > > **See also** > > - [`README.md`]() — the shared creation sequence and the LB-membership patterns. > - [`rke2.md`]() — the same requirements under RKE2ControlPlane. > - [`checklist.md`]() — every requirement from both guides in one place. > - [The machine role]() — the normative contract this guide builds on. # RKE2ControlPlane What RKE2ControlPlane (RCP) and the RKE2 bootstrap provider (CAPRKE2) require of your CAPTF cluster and machine modules, and why this is the case. Citations below prefixed `rke2/` are paths in `rancher/cluster-api-provider-rke2` at tag `v0.25.2`; `rke2docs/` is `rancher/rke2-docs` (`main` branch); `capi/` is the CAPI v1.14.2 tree (`https://github.com/kubernetes-sigs/cluster-api/tree/v1.14.2`). See also [the shared creation sequence]() and [the requirements checklist]() for every requirement in one place. ## Version and compatibility CAPRKE2 v0.25.2 is built against `sigs.k8s.io/cluster-api v1.13.5` (`rke2/go.mod:33`) and its v1beta2 API is the contract version CAPTF also targets. Its own e2e suite runs only against CAPI core v1.12.11 and v1.13.5 (`rke2/test/e2e/config/e2e_conf.yaml:22,33`), and its getting-started guide pins CAPI core to v1.13.5 (`rke2/docs/book/src/01_user/01_getting-started.md:56`). > [!WARNING] > > **CAPI v1.14 core is untested by CAPRKE2 v0.25.2** > > CAPTF runs CAPI v1.14.2. The contract shape is unchanged, but run a compatibility smoke test before relying on it in production; see [open questions](<#open-questions>). ## What RKE2ControlPlane reads from your modules - **`Cluster.status.initialization.infrastructureProvisioned`.** RCP creates nothing — not even certificates — until this is `true`; RKE2Config generates no bootstrap data until it is `true` either (`rke2/controlplane/internal/controllers/rke2controlplane_controller.go:376-394`; `rke2/bootstrap/internal/controllers/rke2config_controller.go:171-182`). - **`Cluster.spec.controlPlaneEndpoint`.** A hard gate: RCP creates no Machine at all — including the first — while the endpoint is not valid (host and port both set) (`rke2/controlplane/internal/controllers/rke2controlplane_controller.go:421-440`). Only the endpoint’s **host** feeds the RKE2 `tls-san` list and the init node’s `ServerURL` (`https://:9345`); the endpoint’s **port is ignored** for the supervisor join URL, which is always 9345 (`rke2/pkg/rke2/config.go:350`; `rke2/bootstrap/internal/controllers/rke2config_controller.go:66,455-473`). - **`Cluster.spec.clusterNetwork`.** `pods`/`services.cidrBlocks` map to `cluster-cidr`/`service-cidr` when non-empty (`rke2/pkg/rke2/config.go:209-215`). `serviceDomain` and `api_server_port` are **not read at all**: cluster domain comes from RKE2’s own `serverConfig.clusterDomain`, and the API server is always on 6443 on the node (`rke2/pkg/rke2/config.go:224-225`). Do not derive an LB backend port from `api_server_port` under RKE2 beyond using it as the frontend port if your templates choose to. - **Failure domains.** Same rule as KCP: only entries with `controlPlane: true` are used for scale-up/scale-down placement; a nil `controlPlane` counts as `false` (`rke2/pkg/rke2/control_plane.go:176-190`). - **`Machine.status.addresses`.** Read only for three of the five `registrationMethod` values, and only from Machines whose `Ready` condition is already `True` (the `registrationMethod` table below; `rke2/controlplane/internal/controllers/status.go:80,152,158-177`). The contract’s canonical output order for `addresses` (InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname — see [`../contract/v1alpha1/machine.md`]() `addresses`) is what makes RCP’s first-match logic deterministic; your module’s job is only to **emit the address type(s) each method requires** (below), not to order them — the controller does that. - **`Machine.status.nodeRef`.** Used for etcd leader-move/member-removal ordering and remediation; `HasHealthyMachineStillProvisioning` = a healthy Machine with no Node yet (`rke2/pkg/rke2/workload_cluster_etcd.go:40-45`; `rke2/pkg/rke2/control_plane.go:527-529`). - **`registrationMethod`** (immutable once set, enforced by webhook): | Method | `status.availableServerIPs` = | Address types your module must supply | Join URL | | --- | --- | --- | --- | | `control-plane-endpoint` (default, `""` ≡ this) | `[endpoint.host]` | none | `https://:9345` | | `address` | `[spec.registrationAddress]` | none | `https://:9345` | | `internal-first` | first `InternalIP` or `ExternalIP` per Ready CP Machine, in list order | `InternalIP` and/or `ExternalIP` | `https://:9345` | | `internal-only-ips` | first `InternalIP` per Ready CP Machine | `InternalIP` | `https://:9345` | | `external-only-ips` | first `ExternalIP` per Ready CP Machine | `ExternalIP` | `https://:9345` | (`rke2/pkg/registration/registration.go:46,68-140`; `rke2/controlplane/api/v1beta2/rke2controlplane_types.go:97-100`; `rke2/controlplane/api/v1beta2/rke2controlplane_webhook.go:142-145,211-214`.) A Ready Machine with no address of the required type produces the RCP status error “ready but they have no IP Address available” (`rke2/controlplane/internal/controllers/status.go:175-177`); see [open questions](<#open-questions>). - **`RKE2ControlPlane.spec.version`.** Must match `(v\d\.\d{2}\.\d+\+rke2r\d)|^$` and is copied verbatim to `Machine.spec.version` (`rke2/controlplane/api/v1beta2/rke2controlplane_types.go:79-81`; `rke2/controlplane/internal/controllers/scale.go:566-567,612`). See [`../contract/v1alpha1/machine.md`]() `kubernetes_version` and the [module checklist](<#module-checklist>). ## What RCP and RKE2Config write - **Bootstrap Secret.** Name = RKE2Config name; keys `value` and `format` are **always** both written (there is no case where `format` is absent); `format` defaults to `cloud-config` (`rke2/bootstrap/internal/controllers/rke2config_controller.go:1040-1061`; `rke2/bootstrap/api/v1beta2/rke2config_webhook.go:79-80`). See [`../contract/v1alpha1/machine.md`]() `bootstrap_format`. - **Ignition is supported** end to end (init, CP join, worker), via Butane → Ignition 3.3 (`rke2/bootstrap/internal/controllers/rke2config_controller.go:552-556,798-802,927-931`). - **`gzipUserData: true`.** For cloud-config, `value` becomes raw gzip bytes (not base64, not UTF-8 text); `format` stays `cloud-config`. For Ignition, the gzip is wrapped inside the Ignition config instead (`rke2/bootstrap/internal/controllers/rke2config_controller.go:1015-1037`). See [machine.md’s `bootstrap_data` input]() for how the contract resolves the binary-payload question, and the [module checklist](<#module-checklist>) for what it means for your module. - **Cluster Secrets.** `-ca`, `-cca` (client CA), `-peer-etcd`, `-etcd`, `-kubeconfig`, `-token`. No `-sa`, no `-proxy` (contrast with KCP’s secret set, [kubeadm.md’s what KCP and CABPK write]()) (`rke2/pkg/secret/certificates.go:52-80,168-190,382-384`). - **Machine labels/annotations.** Template labels plus forced `cluster.x-k8s.io/cluster-name`, `cluster.x-k8s.io/control-plane: ""`; a `pre-terminate.delete.hook.machine.cluster.x-k8s.io/rke2-cleanup` annotation on every control-plane Machine (`rke2/controlplane/internal/controllers/scale.go:650-666,577-578`). - **RCP status (v1beta2).** `initialization.controlPlaneInitialized` (true once the workload `kube-system/rke2-serving` Secret exists, or once every owned Machine is Ready); `availableServerIPs`; conditions including `EtcdClusterHealthy`, `ControlPlaneComponentsHealthy`, `Remediating` (`rke2/controlplane/internal/controllers/status.go:98-194`; `rke2/controlplane/api/v1beta2/rke2controlplane_types.go:262-316`). - **Node identity.** CAPRKE2 **never sets** kubelet `provider-id`, `node-ip` or `node-name` itself; the fields exist in its config but nothing in the repo writes them (`rke2/pkg/rke2/config.go:501-503`). See the [module checklist](<#module-checklist>) and [machine.md’s node providerID matching]() — RKE2. ## Lifecycle constraints - **Init.** Exactly one control-plane Machine gets init data, guarded by an init-lock ConfigMap. The init node has no `server:` configured — it creates the cluster standalone and, notably, needs **no LB/9345 reachability to itself** during init, unlike the hairpin requirement that applies to every later join (`rke2/controlplane/internal/controllers/scale.go:51-99`; `rke2/pkg/rke2/config.go:650-676`). Don’t read this as weakening the general endpoint-reachability requirement in [networking and load balancer](<#networking-and-load-balancer>): it applies to every Machine that joins, i.e. every control-plane Machine after the first, and to every worker. - **Join (control-plane and worker), in order.** Requires the Cluster’s `ControlPlaneInitialized` condition; the `-token` Secret; and `RCP.status.availableServerIPs` non-empty, which itself needs at least one **Ready** control-plane Machine and a reachable workload API server via the `-kubeconfig` Secret (`rke2/bootstrap/internal/controllers/rke2config_controller.go:230,699-703,852-856`; `rke2/controlplane/internal/controllers/status.go:114-146,152`). Both CP joins and worker joins use `availableServerIPs[0]` (`rke2/bootstrap/internal/controllers/rke2config_controller.go:717,867`). - **Rollout.** `maxSurge` 0 or 1, default 1; scale-up then scale-down of the oldest outdated Machine in the most-populated failure domain (`rke2/controlplane/api/v1beta2/rke2controlplane_types.go:548-580`; `rke2/controlplane/internal/controllers/scale.go:285-334`). - **Preflight (scale up/down, in-place).** No Machine deleting; every control-plane Machine must have `AgentHealthy` and `EtcdMemberHealthy` both `True` (`Unknown` or missing blocks); a missing Node blocks by extension (`rke2/controlplane/internal/controllers/scale.go:200-266`). - **Scale-down / deletion.** Etcd leadership is forwarded to the newest Machine, then the Machine is deleted; the `rke2-cleanup` pre-terminate hook re-forwards leadership, annotates the Node for etcd removal, and waits for confirmation before releasing — InfraMachine deletion happens only after that (`rke2/controlplane/internal/controllers/scale.go:141-197`). - **Remediation.** RCP remediates control-plane Machines with `HealthCheckSucceeded=False` + `OwnerRemediated=False` (MHC-driven). Pre-init remediation is allowed directly; post-init remediation additionally requires more than one replica, no Machine still provisioning without a Node, no Machine deleting, and preserved etcd quorum (`rke2/controlplane/internal/controllers/remediation.go:97-330,366-420`). [Machine.md’s health → remediation section]() already reflects this: remediation happens for Machines whose owner acts on the MachineHealthCheck’s `MachineOwnerRemediated` condition, and RCP is such an owner for its own control-plane Machines, not only MachineSet and KubeadmControlPlane. - **In-place updates.** Behind the CAPI `InPlaceUpdates` feature gate (alpha) plus exactly one registered `CanUpdateMachine` extension; otherwise falls back to delete/recreate. Not something a CAPTF module needs to support, since `TerraformMachine.spec.source` is immutable regardless (`rke2/controlplane/main.go:287`). - **Payload size / air-gap.** Init control-plane user-data embeds four CA key pairs plus config and optional manifest/registry/audit files. Non-air-gapped installs `curl -sfL https://get.rke2.io` at boot, which means **node internet egress**; `airGapped: true` expects pre-baked artifacts in the image instead (`rke2/bootstrap/internal/cloudinit/controlplane_init.go:34-36`). `gzipUserData` exists for size-limited clouds; see [machine.md’s `bootstrap_data` input]() for how the contract handles the resulting binary payload. ## Networking and load balancer | Path | Port | Needed when | | --- | --- | --- | | LB/VIP → CP nodes | 6443/TCP (endpoint port → node 6443; RKE2 ignores `apiServerPort`) | always | | LB/VIP → CP nodes | 9345/TCP supervisor, **same host** as the endpoint | `control-plane-endpoint` (default) and `address` registration methods | | All nodes → CP nodes | 6443, 9345 TCP direct (after registration) | always | | CP ↔ CP | 2379, 2380, 2381 TCP | embedded etcd (not with `externalDatastoreSecret`) | | all ↔ all | 10250 TCP; NodePort range (default 30000-32767) | always | | CNI (default canal) | 8472/UDP VXLAN, 9099/TCP | `cni: canal` or unset | (`rke2/controlplane/internal/controllers/rke2controlplane_controller.go:757`; `rke2docs/docs/install/ha.md:42`; `rke2docs/docs/install/requirements.md:128,142-148,158-161`.) - **No LB listener on 2379 is needed.** The controller reaches etcd only by port-forwarding through the API server; an LB entry for 2379 is not required by anything in RCP (`rke2/pkg/proxy/dial.go:99-106`; `rke2/pkg/etcd/client_generator.go:34`). - **Health checks.** 9345: TLS `GET /v1-rke2/readyz` expecting **403** (unauthenticated) — or plain TCP. 6443: TCP, or HTTPS `/healthz` only if anonymous auth is enabled (`rke2/examples/templates/docker/cluster-template.yaml:60,176-194`). - **DNS endpoints are supported.** RKE2 HA supports a DNS name / round-robin DNS as the fixed registration address; the endpoint host is added to `tls-san` automatically. `registrationAddress` is **not** added to `tls-san` automatically — add it via `serverConfig.tlsSan` if it differs from the endpoint host (`rke2docs/docs/install/ha.md:36-37`; `rke2/pkg/rke2/config.go:350`). - **Backend membership.** The module owns LB membership — nothing in RCP or CAPI registers a backend for you. The instance MUST be added before its Machine is Ready (joins happen through the LB during bring-up), and the LB MUST tolerate a backend whose supervisor is not yet up (see [lifecycle constraints](<#lifecycle-constraints>), “Join”). ## Module checklist ### Cluster module - MUST output a valid `control_plane_endpoint` (host and port) no later than the apply that makes the cluster `provisioned`, unless the user sets one — RCP creates no Machine at all until `Cluster.spec.controlPlaneEndpoint.IsValid()` (`rke2/controlplane/internal/controllers/rke2controlplane_controller.go:421-440`; see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>)). This matches the contract’s own endpoint-timing rule in [`../contract/v1alpha1/cluster.md`]() `control_plane_endpoint` output. - MUST keep the emitted `control_plane_endpoint` stable for the life of the object: CAPI never updates `Cluster.spec.controlPlaneEndpoint` after the first valid copy (`capi/core/reconcilers/cluster/cluster_controller_phases.go:217-228`), so replacing the LB behind it (new DNS name or IP) breaks every kubeconfig and the API server certificate SANs. The controller neither detects nor prevents such a replacement; `lifecycle { prevent_destroy = true }` on the resource behind the endpoint turns it into a failed Job instead of a lost cluster ([`../contract/v1alpha1/cluster.md`]() `control_plane_endpoint` output). - MUST provision **two** LB listeners on the same host: `endpoint.port` → CP nodes `:6443`, and `:9345` → CP nodes `:9345` — not just one, as with KCP (`rke2docs/docs/install/ha.md:42`; see [networking and load balancer](<#networking-and-load-balancer>)). - MUST NOT rely on `cluster_network.api_server_port` or `service_domain` to configure RKE2; RKE2 reads neither (`rke2/pkg/rke2/config.go:224-225`; see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>)). - Backend membership is the module’s job (not CAPI’s): a target group by tag/label, or an explicit attach resource — see the three patterns in [README.md’s LB-membership patterns](). - SHOULD leave `failure_domains[].control_plane` at its default `true` (`rke2/pkg/rke2/control_plane.go:176-190`; see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>)), same as under KCP. ### Machine module - MUST accept both `bootstrap_format` values RCP can write, `cloud-config` and `ignition` (`rke2/bootstrap/internal/controllers/rke2config_controller.go:1040-1061`; see [what RCP and RKE2Config write](<#what-rcp-and-rke2config-write>)). - MUST register the control-plane instance in both the 6443 and 9345 LB target sets, and open/attach the corresponding security rules, before the Machine becomes Ready (`capi/docs/book/src/developer/providers/contracts/infra-machine.md:633,645`; see [networking and load balancer](<#networking-and-load-balancer>)). - MUST supply `addresses` of the type(s) the cluster’s `registrationMethod` needs (the `registrationMethod` table in [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>)) when that method is anything other than `control-plane-endpoint`/`address`; the controller, not the module, is responsible for output ordering (`../contract/v1alpha1/machine.md` `addresses`). - MUST emit `provider_id` matching exactly what the Node’s kubelet registers — either via a cloud-controller-manager, or via `agentConfig.kubelet.extraArgs: [provider-id=]` set from instance metadata, since CAPRKE2 sets none of this itself (see [what RCP and RKE2Config write](<#what-rcp-and-rke2config-write>)). Under RKE2 a mismatch on the **first** control-plane Machine blocks every subsequent join, not just that Machine’s own readiness, because every join needs `availableServerIPs` non-empty, which needs a Ready CP Machine (see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>) and [lifecycle constraints](<#lifecycle-constraints>)). - `kubernetes_version` arrives as `vX.Y.Z+rke2rN`; strip the `+rke2rN` suffix before using it for an image lookup or a semver comparison (see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>)). - Init/CP-join payloads embed four CA key pairs; on size-limited clouds both `cloud-config` + `gzipUserData: true` and `ignition` + `gzipUserData: true` work for reducing payload size (see [what RCP and RKE2Config write](<#what-rcp-and-rke2config-write>)). Either way the module MUST route `bootstrap_data` to a base64-taking argument (e.g. `user_data_base64`); it MUST NOT `base64decode()` it, because a gzipped payload is not valid UTF-8. ## Differences from KubeadmControlPlane | Aspect | KubeadmControlPlane | RKE2ControlPlane v0.25.2 | | --- | --- | --- | | LB ports | API port only | API (6443) **and** supervisor 9345, same host | | Join target | CP endpoint, backend port | `availableServerIPs[0]`:9345 — endpoint host, `registrationAddress`, or a CP Machine IP | | Uses `Machine.status.addresses` | no | yes, for 3 of 5 `registrationMethod` values | | Join gate | Cluster `ControlPlaneInitialized` | that, **plus** ≥1 Ready CP Machine (needs Node ⇒ providerID match) | | `Machine.spec.version` | `vX.Y.Z` | `vX.Y.Z+rke2rN` | | `clusterNetwork.apiServerPort` | not wired to bindPort either, but the field exists as a convention (see [kubeadm.md’s what it reads]()) | ignored entirely; 6443 is fixed | | `clusterNetwork.serviceDomain` | honored by CABPK | ignored (`serverConfig.clusterDomain`) | | Bootstrap `format` | always present (`cloud-config` or `ignition`) | always present | | Secrets | `-ca`, `-etcd`, `-sa`, `-proxy`, `-kubeconfig` | `-ca`, `-cca`, `-etcd`, `-peer-etcd`, `-kubeconfig`, `-token` | | etcd removal | KCP pre-terminate hook, etcd client | `rke2-cleanup` pre-terminate hook, Node annotation `etcd.rke2.cattle.io/remove` | | Install | kubeadm/kubelet pre-installed in image | `curl get.rke2.io` at boot unless air-gapped artifacts are baked in | | Contract | v1beta2, built on CAPI v1.14 | v1beta2, built on CAPI v1.13.5 | (Individual rows cited in the corresponding sections above and in `kubeadm.md`.) ## Open questions - **No `InternalIP` guarantee.** `addresses` may legally be `[]` (`../contract/v1alpha1/machine.md` `addresses`). For `internal-only-ips`/`external-only-ips`/`internal-first`, a Ready Machine without the right address type blocks every subsequent join with an RCP status error (see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>)). No `tfcapi-lint` rule warns when a module never emits an `InternalIP`. - **CAPI v1.14 compatibility.** CAPRKE2 v0.25.2 is tested only up to CAPI core v1.13.5 (see [version and compatibility](<#version-and-compatibility>)). CAPTF runs v1.14.2. Run a compatibility smoke test against v1.14.2 before relying on RKE2ControlPlane in production; watch for a CAPRKE2 release built on v1.14. - **Nondeterministic join target.** Under IP-based `registrationMethod` values, `availableServerIPs` is built by iterating a Go map, so which control-plane node a joiner targets is nondeterministic, and a stale or just-removed node’s address can be chosen until the next status refresh (see [what RKE2ControlPlane reads](<#what-rke2controlplane-reads-from-your-modules>); `rke2/bootstrap/internal/controllers/rke2config_controller.go:717,867`). A module cannot fix this directly; a Machine that leaves `Ready` (for example through a MachineHealthCheck) drops out of `availableServerIPs` on the next status refresh. > [!NOTE] > > **See also** > > - [`README.md`]() — the shared creation sequence and the LB-membership patterns. > - [`kubeadm.md`]() — the same requirements under KubeadmControlPlane. > - [`checklist.md`]() — every requirement from both guides in one place. > - [The machine role]() — the normative contract this guide builds on. # Requirements Checklist Every port, health check and ordering rule a KubeadmControlPlane (KCP) or RKE2ControlPlane (RCP) cluster needs from a CAPTF cluster or machine module, grouped by topic. Details and citations are in [`kubeadm.md`]() and [`rke2.md`](); the shared creation sequence and LB-membership patterns are in [`README.md`](). `n/a` means the row does not apply to that provider or module role. “Verified by” is one of: | Value | Meaning | | --- | --- | | `lint` | `tfcapi-lint` can check it today. | | `e2e` | Only an end-to-end run with a real control-plane provider catches a violation. | | `operator` | Checked by whoever reviews or writes the module; no automated check exists. | | `none` | Not checked anywhere today. | Some `lint`-worthy rules have no check today; those are called out in [kubeadm.md’s open questions]() and [rke2.md’s open questions](). > [!WARNING] > > **No `e2e` row has been verified** > > The project has no end-to-end suite today, so no `e2e` row below has been verified against a real control plane. ## Endpoint and load balancer | Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by | | --- | --- | --- | --- | --- | --- | | Endpoint valid before any Machine | KCP creates no Machine until `Cluster.spec.controlPlaneEndpoint.IsValid()` (host and port both set) | RCP returns before creating any Machine, including the first, while `!IsValid()` | MUST output a valid `control_plane_endpoint` no later than the apply that makes the cluster `provisioned`, unless the user sets one | n/a | e2e | | LB frontend → CP backend port | `endpoint.port` → CP `:bindPort` (default `6443`); one frontend | `endpoint.port` → CP `:6443` (`apiServerPort` ignored); one of two frontends | MUST create a frontend on `endpoint.port` and a backend on `api_server_port ?? 6443`, kept equal to the module’s `bindPort`/RKE2’s fixed 6443 | n/a | operator | | Second LB listener on 9345 | not applicable | `:9345` → CP `:9345`, **same host** as the endpoint; required for `control-plane-endpoint` (default) and `address` registration methods | MUST create this second listener for RKE2 clusters | MUST register the instance in the 9345 target set alongside 6443 | operator | | Hairpin reachability | The endpoint MUST be reachable from the first CP node itself — `admin.conf`/`super-admin.conf`/`kubelet.conf` all point at the endpoint, not the node’s local address | Same MUST applies to every RKE2 join. The init node itself has no `server:` configured and so needs no LB/9345 reachability at init time ([rke2.md’s lifecycle constraints]()) — this does not weaken the MUST for every later join | MUST make the LB/VIP allow a backend to reach its own frontend | n/a | none | | Endpoint stability | Once emitted, `control_plane_endpoint` MUST be stable for the life of the object — CAPI never updates `Cluster.spec.controlPlaneEndpoint` after the first valid copy | Same rule; RCP also builds `tls-san` and the kubeconfig Secret from the first copy | MUST NOT replace the LB behind an already-emitted endpoint (new DNS name or IP); consider `lifecycle { prevent_destroy = true }` on that resource | n/a | none | ## Health checks and backend membership | Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by | | --- | --- | --- | --- | --- | --- | | Health check — API port | TCP, or HTTPS `/readyz`/`/healthz` without certificate verification; MUST go green with a single backend during init | TCP, or HTTPS `/healthz` only if anonymous auth is enabled | MUST configure a check matching one of these | n/a | operator | | Health check — RKE2 supervisor (9345) | not applicable | TLS `GET /v1-rke2/readyz` expecting **403** (unauthenticated), or plain TCP | MUST configure this check for RKE2 clusters | n/a | operator | | Backend membership timing | Instance MUST be in the LB backend before `kubeadm init`/`kubeadm join` finishes on it, or `controlPlaneInitialized` never latches | Instance MUST be added before its Machine becomes Ready; the LB MUST tolerate a backend whose supervisor is not yet up | n/a (registration is the machine module’s job) | MUST register/deregister the instance in its own Terraform state, in the same apply as instance creation | e2e | | SG / firewall — control plane | Backend port (`bindPort`) from LB, nodes and management; TCP 2379-2380 and 10250 between CP nodes; 10250 from CP to workers | CP↔CP TCP 2379-2381; all→CP TCP 6443 and 9345; all↔all TCP 10250 and the NodePort range (default 30000-32767); CNI ports for the chosen CNI (default canal: 8472/udp, 9099/tcp); egress to `get.rke2.io`/GitHub releases unless air-gapped | MUST create these security groups/firewall rules | MUST attach the instance to them | operator | ## Node identity and addresses | Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by | | --- | --- | --- | --- | --- | --- | | `provider_id` == Node `providerID` | MUST match exactly — the Machine controller links Node to Machine only on an exact match; KCP gates join preflight, etcd matching and remediation on the resulting `nodeRef` | MUST match exactly; a mismatch on the first CP Machine blocks every later join (no Ready CP Machine ⇒ `availableServerIPs` stays empty) | n/a | MUST emit the same value the Node’s kubelet/CCM will register (CCM sets it, or `kubeletExtraArgs`/`agentConfig.kubelet.extraArgs: [provider-id=...]`) | e2e | | `addresses` types and order | Not read by KCP or CABPK at all | Required address types by `registrationMethod`: `control-plane-endpoint`/`address` — none; `internal-first` — `InternalIP` and/or `ExternalIP`; `internal-only-ips` — `InternalIP`; `external-only-ips` — `ExternalIP`. Order is fixed by the controller (`InternalIP, InternalDNS, ExternalIP, ExternalDNS, Hostname`), not the module | n/a | MUST emit the required type(s) for the cluster’s `registrationMethod`; MUST NOT rely on emission order | none | | `kubernetes_version` suffix | Arrives as plain `vX.Y.Z` | Arrives as `vX.Y.Z+rke2rN`, copied verbatim from `RKE2ControlPlane.spec.version` | n/a | Under RKE2, MUST strip `+rke2rN` before using the value for an image lookup or a semver comparison | none | ## Bootstrap payload | Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by | | --- | --- | --- | --- | --- | --- | | `bootstrap_format` always set | Always written by CABPK (`cloud-config` default, `ignition` behind the alpha feature gate) | Always written by CAPRKE2 (`cloud-config` default, `ignition` fully supported) | n/a | MUST accept both `cloud-config` and `ignition` | e2e | | `bootstrap_data` is base64, including the gzip case | Payload is always UTF-8 text (`cloud-config`/`ignition`), base64-encoded by the CAPTF controller like every other bootstrap provider | With `gzipUserData: true`, the decoded payload is raw (non-UTF-8) gzip bytes — still base64-encoded by the CAPTF controller, same as any other payload | n/a | MUST pass `bootstrap_data` to a base64-taking argument, or `base64decode()` it only when the content is known to be UTF-8 (never for a gzipped payload) | none | | CP bootstrap payload size and secrecy | Init/join payload embeds cluster CA, etcd CA, service-account and front-proxy key material, uncompressed | Init payload embeds four CA key pairs, uncompressed unless `gzipUserData: true` | n/a | MUST gzip the payload or stage it in a secret store with a small stub, and MUST keep key material out of readable instance metadata | none | ## Templates, timing and compatibility | Requirement | KubeadmControlPlane | RKE2ControlPlane | Cluster module | Machine module | Verified by | | --- | --- | --- | --- | --- | --- | | `failure_domains[].control_plane` default | Only entries with `controlPlane == true` are visible to KCP; nil counts as `false` | Same rule for RCP | SHOULD leave the field at its default `true` | n/a | none | | Control-plane bring-up timing | N replicas ≈ N × (apply + boot + join); joins strictly serialized on the previous Machine’s `nodeRef` | Same shape; a join additionally requires ≥1 already-Ready CP Machine (Node with matching `providerID`) | n/a | n/a (informs MHC timeout sizing) | none | | Control-plane template defaults | Templates SHOULD set `remediation.maxRetry` (e.g. `3`) and a non-zero `retryPeriodSeconds` (defaults are unlimited retries, immediate retry); the control-plane MachineHealthCheck’s `InfrastructureReady=False` timeout MUST exceed apply time plus one drift interval (no fixed number specified here) | RCP has its own `remediationStrategy{maxRetry, retryPeriod, minHealthyPeriod}`; no CAPTF-recommended value is given | n/a | n/a | operator | | Control-plane provider / CAPI compatibility | Built for and tested against CAPI v1.14.2 (this provider’s target) | CAPRKE2 v0.25.2 is tested by its own e2e suite only against CAPI core v1.12.11 and v1.13.5; CAPTF runs v1.14.2. Run a compatibility smoke test before relying on RKE2ControlPlane in production | n/a | n/a | none | > [!NOTE] > > **See also** > > - [`README.md`]() — the shared creation sequence and the LB-membership patterns. > - [`kubeadm.md`]() and [`rke2.md`]() — the full guidance and citations behind each row above. > - [The module contract]() — the normative cluster and machine roles both guides build on. # Operations # Installation This page covers installing the CAPTF provider into a management cluster with `clusterctl`: what to have ready first, how to register the provider, what the install creates, how the runner image is set, the optional components, and how to confirm the install worked. It is for whoever administers the management cluster, not for cluster tenants. > [!NOTE] > > **Before you begin** > > - A management cluster with Cluster API’s core, bootstrap and control-plane providers already initialized, and a `kubectl` context pointing at it. > - `clusterctl`, matching the version documented for the Cluster API release you run. > - `cert-manager`, with the `cert-manager.io/v1` API available. `clusterctl init` installs a compatible `cert-manager` release itself when one is not already present, so a separate install step is only needed to pin a specific `cert-manager` version or to install it ahead of time. > - Cluster API installed at a release that implements the same contract CAPTF does. CAPTF’s own `metadata.yaml` lists one release series so far, contract `v1beta2`; CAPTF is built and tested against Cluster API v1.14.2. ## Register the provider CAPTF is not one of `clusterctl`’s built-in providers, so `clusterctl` needs a config file naming its release manifest. The config entry’s `name` is `terraform`: clusterctl.yaml ```yaml providers: - name: terraform type: InfrastructureProvider url: https://github.com/captf-io/cluster-api-provider-terraform/releases/latest/infrastructure-components.yaml ``` `url` can also name a specific tag instead of `latest`, or a `file://` path into a local repository built from a release’s assets; see [Installing from a local repository]() for the local repository layout and for pinning a version. > [!WARNING] > > **No release is published yet** > > CAPTF has not published a release yet, so the URL above does not resolve to anything: install from a local repository until one exists. ## Install with clusterctl init ```sh clusterctl init --config clusterctl.yaml --infrastructure terraform ``` `clusterctl init --infrastructure terraform:vX.Y.Z` pins a specific released version instead of the newest one clusterctl can see. If you plan to use the ClusterClass flavor, enable the `ClusterTopology` feature gate before running `init` — it is alpha in Cluster API and off by default: ```sh CLUSTER_TOPOLOGY=true clusterctl init --config clusterctl.yaml --infrastructure terraform ``` ## What gets installed Everything below lands in one namespace, `captf-system`; `clusterctl init` also installs `cert-manager` itself when it is missing, in its own namespace. - **CRDs** for the seven kinds: `TerraformCluster`, `TerraformClusterTemplate`, `TerraformMachine`, `TerraformMachineTemplate`, `TerraformMachinePool`, `TerraformMachinePoolTemplate` and `TerraformClusterIdentity` (the last is cluster-scoped; the rest are namespaced). See [The Kinds]() for what each one does, and [Custom Resources]() for every field. - **The manager**, a single-replica `Deployment` named `captf-controller-manager`, running as a non-root user, with a `ServiceAccount` of the same name. It serves its webhooks on `:9443`, its metrics on `:8443`, and its health and readiness probes on `:9440`. - **The admission webhook**: a `ValidatingWebhookConfiguration` named `captf-validating-webhook-configuration`, one rule per kind, backed by the `captf-webhook-service` `Service`. Its serving certificate is a `cert-manager` `Certificate` (`captf-serving-cert`, issued by the self-signed `Issuer` `captf-selfsigned-issuer`), mounted into the manager pod and kept current by `cert-manager`’s CA injection. - **RBAC** for the manager itself: the `ClusterRole` `captf-manager-role` and its `ClusterRoleBinding`, plus a leader-election `Role` and `RoleBinding` scoped to `captf-system`. The manager also creates a runner `ServiceAccount` and `RoleBinding` in each tenant namespace at first use, from the static `ClusterRole` `captf-runner`. See [RBAC]() for what each role grants and for the per-namespace runner setup. None of this installs a `TerraformClusterIdentity` or any Terraform\* object: those come from templates you apply afterward, covered in [Identities and Credentials]() and [Templates and ClusterClass](). ## The manager image and the runner image `clusterctl init` sets the manager container’s image to the release’s image, `ghcr.io/captf-io/cluster-api-provider-terraform:vX.Y.Z`. The same reference is also set as the `CAPTF_MANAGER_IMAGE` environment variable on that container: it is the default of `--runner-image`, the image the manager runs as the init container that injects the runner binary into every Job it creates. Point `--runner-image` at a different reference (for example a registry mirror the tenant namespaces’ Jobs can reach, or a build carrying a patched runner binary) when the manager’s own image is not the one you want Jobs to pull; see [`--runner-image`]() for the flag and [Job Environment]() for how the init container uses it. `spec.source.image` on a `Terraform*` object is unrelated: it names the role image the Job actually runs, never the runner. A custom deployment, or an overlay that replaces the manager’s `env`, must keep `POD_NAMESPACE` and `SERVICE_ACCOUNT_NAME`, set from `metadata.namespace` and `spec.serviceAccountName` through the downward API. See [Manager environment](). ## Optional components `config/prometheus` and `config/network-policy` are kustomize components: neither is part of `infrastructure-components.yaml`, so `clusterctl init` never installs them and `clusterctl upgrade` never touches them, and neither is a release asset — get them from a checkout of the tag you installed (see [Register the provider](<#register-the-provider>)), not from the release URL. A kustomize `Component` builds standalone: point a bare kustomization at one with no `resources:` of your own, and it emits only that component’s own objects, carrying the `captf-`/`captf-system` names `config/default` produces: ```yaml apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization components: - /config/prometheus ``` - **Prometheus** (`config/prometheus`): a metrics `Service`, a `ServiceMonitor`, the alerting rules, and the RBAC Prometheus needs to scrape the manager’s authenticated metrics endpoint. See [Enabling the Prometheus component]() for what it adds and how to wire it up, and how a later upgrade affects it. - **Network policy** (`config/network-policy`): a `NetworkPolicy` that restricts the manager pod to webhook, metrics and health-probe traffic inbound, and DNS, the API server and container registries outbound. It only has an effect on a CNI that enforces `NetworkPolicy`, and it covers only the manager pod, not runner Jobs; `config/network-policy/job-egress-sample.yaml` is a separate starting point to copy into each tenant namespace for its Job pods. See [Network exposure]() for why the manager’s and a Job’s exposure differ. Building both at once needs both under `components:` in the same kustomization. Building either on top of a from-source install of `config/default` (rather than the one `clusterctl init` already installed) also works, but is unnecessary for these two components: each one’s objects stand alone and never depend on `config/default`’s own resources being built alongside them. ## Verifying the install ```sh kubectl -n captf-system rollout status deployment/captf-controller-manager kubectl get crds -l cluster.x-k8s.io/provider=infrastructure-terraform kubectl get validatingwebhookconfigurations captf-validating-webhook-configuration kubectl -n captf-system get certificate captf-serving-cert ``` The rollout command returns once the manager pod is ready. The CRD list should show all seven kinds. The webhook configuration and the certificate must both exist and the certificate must report `Ready=True`, cert-manager was able to issue the webhook’s serving certificate: without it, webhook calls from the API server fail closed (`failurePolicy: Fail`) and every `Terraform*` create or update is rejected. `clusterctl init` itself prints the components it installed; `clusterctl describe cluster` (once you have created one) reports whether CAPTF’s objects are ready. > [!NOTE] > > **See also** > > - [Upgrades]() for upgrading an existing install. > - [Configuration]() for the manager flags you set on top of this default install. > - [RBAC]() for the manager’s and runner’s permissions in full. > - [Security Model]() for the trust boundary a `Terraform*` object’s Job operates inside. # Configuration The CAPTF manager takes every setting as a command-line flag; there is no config file. This page explains what each group of flags changes and when to change it. For the full flag list, types and defaults, see [Manager Flags](); this page does not repeat that table. > [!NOTE] > > **Before you begin** > > - Cluster-admin access to the management cluster, and `kubectl`. > - The shipped Deployment is named `captf-controller-manager` in the provider’s namespace (`captf-system` after `clusterctl init`; see [Installation]()). ## Changing a flag The manager’s arguments live on the `manager` container of the `captf-controller-manager` Deployment. > [!WARNING] > > **A strategic-merge patch that adds one flag drops the others** > > Because `args` is a plain list with no merge key, a strategic-merge patch that only adds one flag replaces the whole list and silently drops the others (including `--leader-elect` and the diagnostics flags the shipped manifest sets). Add a flag with a JSON patch instead, which appends to the existing list: ```sh kubectl patch deployment captf-controller-manager -n captf-system --type=json \ -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--="}]' ``` To manage the arguments as a whole (for example with a kustomize overlay over the released manifest), patch the full `args` list so no existing flag is lost: ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: captf-controller-manager namespace: captf-system spec: template: spec: containers: - name: manager args: - --leader-elect - --diagnostics-address=:8443 - --insecure-diagnostics=false - --webhook-port=9443 - --= ``` Confirm the change with: ```sh kubectl rollout status deployment/captf-controller-manager -n captf-system kubectl get deployment captf-controller-manager -n captf-system \ -o jsonpath='{.spec.template.spec.containers[0].args}' ``` The manager also logs its parsed flags once at start (`FLAG: --name="value"` lines), and refuses to start on an invalid combination, so a typo or an out-of-range value shows up in the Pod’s logs and status rather than running with a wrong value. > [!WARNING] > > **A hand-patched flag does not survive an upgrade** > > `clusterctl upgrade apply` reinstalls the provider’s manifests, so a hand-patched flag does not survive an upgrade unless the patch is reapplied; see [Upgrades](). ## Manager environment The manager Deployment must set `POD_NAMESPACE` (from `metadata.namespace`) and `SERVICE_ACCOUNT_NAME` (from `spec.serviceAccountName`) through the downward API. The shipped manifests do; check that a custom deployment or a kustomize overlay keeps them. Without them the manager logs a warning, and the webhook refuses `providerID` writes because it cannot tell which ServiceAccount is the manager, so machines never finish provisioning. The `--state-backups` flag and the encryption of Secrets at rest are covered in [Operator files and settings](). ## Namespace scoping and `--watch-filter` `--namespace` restricts the manager to one namespace: it only watches and reconciles namespaced objects (`TerraformCluster`, `TerraformMachine`, `TerraformMachinePool`, their `*Template` kinds, and the Secrets and ConfigMaps CAPTF reads) in that namespace. `TerraformClusterIdentity` is cluster-scoped and is always watched everywhere, regardless of `--namespace`. Unset (the default, and what `clusterctl init` installs), the manager watches every namespace. **A single manager instance therefore watches one namespace or all of them, never a chosen set**: running a second, differently-namespaced manager instance does not split load or ownership the way it might look like it should. The leader-election Lease name (`controller-leader-election-captf`) is a fixed constant, not namespaced or parameterized by `--namespace`, and the shipped install is a singleton — one `ValidatingWebhookConfiguration`, one webhook `Service`, one `Certificate`. Two manager Deployments with `--leader-elect` and different `--namespace` values both contend for the same Lease in the same manager namespace, so only one of them ever holds leadership and reconciles anything; the other sits idle regardless of which namespace it was pointed at. To share one management cluster across several provider instances or tenants, use `--watch-filter` instead: it limits reconciling to objects labeled `cluster.x-k8s.io/watch-filter: `; unset, the manager reconciles every object it can see. It is the convention CAPI and its other providers use for exactly this, and each instance still watches (and its webhook still admits) every namespace. The filter applies to the `Terraform*` objects themselves; a Job or Secret one of them owns is still acted on regardless of its own labels. > [!WARNING] > > **An object in an unwatched namespace is admitted but never reconciled** > > The `ValidatingWebhookConfiguration` has no `namespaceSelector`: it admits a `Terraform*` create or update in any namespace, whether or not any manager instance watches that namespace or the object’s `watch-filter` label matches one. A `Terraform*` object created in a namespace no running manager watches (outside `--namespace`’s one namespace, or not labeled for any instance’s `--watch-filter`) is admitted normally and then never reconciled: no Job ever starts for it, and its conditions never move past whatever they were set to (or left unset) at admission. Both flags apply to the [orphan sweep]() too: `--namespace` limits which namespaces it sweeps, and it always ignores `--watch-filter` when it decides whether a namespace still holds a `Terraform*` object, so another instance’s objects (including one this manager does not watch) keep a namespace from being swept. ## Per-kind concurrency `--terraformcluster-concurrency`, `--terraformmachine-concurrency`, `--terraformmachinepool-concurrency` and `--terraformmachinetemplate-concurrency` each cap how many objects of that kind the manager reconciles at once (default 10). Raise one when that kind’s objects queue behind each other under load — reconciles are lightweight (a handful of API reads and a status patch, unless a Job needs starting) and mostly wait on Job completion, so a higher number rarely costs much CPU. Lower one to reduce the manager’s burst of API calls against a small or rate-limited API server. ## Leader election `--leader-elect` (default `false`) turns on leader election so that only one of several manager replicas reconciles at a time; enable it whenever you run more than one replica, so a rolling update or a crash never leaves two managers reconciling the same objects together. The shipped Deployment runs one replica but sets `--leader-elect` anyway, ready for a scale-up. `--leader-elect-lease-duration`, `--leader-elect-renew-deadline` and `--leader-elect-retry-period` tune how fast a crashed leader is detected and replaced; the defaults (15s/10s/2s) match `kube-controller-manager` and rarely need changing. ## Sync period and the orphan sweep `--sync-period` (default 10 minutes) is the minimum interval at which the manager’s informers re-enqueue every cached object for reconciliation, on top of the normal event-driven reconciles; it reads only the local cache, never the API server, so it does not set the cadence of drift or health checks, which run on their own schedule (see [Drift and Health]()), and it does not recover a watch event that never reached the cache. Lowering it corrects a missed requeue sooner at the cost of more reconciles; raising it does the opposite. The same interval drives the orphan sweep, so lowering `--sync-period` also cleans up an orphaned namespace’s runner objects sooner. See [The orphan sweep]() for what it removes and how it decides. ## `--cluster-operation-gate` Keeps a `TerraformCluster`’s apply, destroy or restore from running at the same time as its machines’ and machine pools’ applies, destroys or restores, through a per-Cluster write Lease; the per-object run Lease that keeps two Jobs from starting for the same object is always on and cannot be turned off. Default `true`. Turn it off only if you accept a cluster’s own operation and its machines’ operations running concurrently — see [Run leases and the cluster operation gate]() for what the gate does and how the two sides wait for each other. ## `--runner-image` The image of the init container that copies the runner binary into every Job; it must be an image that contains a `/runner` binary, which in practice means a CAPTF manager or runner image, not a module image. Unset, it defaults to `$CAPTF_MANAGER_IMAGE`, the manager’s own image, which the shipped Deployment sets to whatever image it runs — so most installs never need to set the flag or the environment variable by hand. Set it to pin the init container to a specific published image independently of the manager’s own, for example while testing a new manager build against the current runner. The manager refuses to start when neither the flag nor the environment variable resolves to a valid image reference. See [Job Environment]() for where the copied binary ends up and [Security Model]() for why images are pinned by digest once a Job has run. ## `--runner-events` Default `true`. Have each Job’s runner post its own progress (`RunStarted`, `Step*`, `PlanSummary`, `ResourcesChanged`, `RunFinished`) as Events on the `Terraform*` object the Job is for, related to the Job itself; emission is best effort and never fails or slows the run. Turn it off to reduce Event volume on a cluster with many objects, or if the runner ClusterRole in your installation does not grant `events` `create` (see [RBAC]()). See [Events]() for the full list. ## `--state-backups` How many state backups to keep per object (default 5); every new state serial the manager observes is copied into a `captf-state-backup-*` Secret, and older copies beyond the count are pruned in the same pass. `--state-backups=0` takes no new backups but leaves existing ones in place and restorable. Lower it to reduce the Secret count and storage in a large installation; raise it for a longer recovery window. See [Terraform State]() for the backup naming and what is skipped, and [State Restore]() for the restore procedure. ## `--drift-default-interval` The drift check interval (default 30 minutes) an object falls back to when neither it nor, for a machine or pool, its `TerraformCluster`’s `spec.defaults.drift` sets one. For a `TerraformCluster` or `TerraformMachine`, it has no effect once `spec.drift.intervalSeconds` (own or inherited) is set, including to `0`, which disables drift. A `TerraformMachinePool`’s drift can never be disabled: a `0`, its own or inherited, falls back to this default instead. Change the default to shift the fleet-wide drift cadence without touching every object; see [Drift]() for setting an interval per object and [Drift and Health]() for how the schedule and its jitter work. ## Diagnostics address, authentication and TLS `--diagnostics-address` (default `:8443`) is where the manager serves Prometheus metrics, authenticated and authorized against the API server by default. `--insecure-diagnostics` (default `false`) turns that off and serves plain HTTP with no authentication instead; use it for local development only, never for a manager reachable from anything but your own workstation. `--tls-min-version`, `--tls-cipher-suites` and `--tls-curve-preferences` constrain the TLS the metrics and webhook servers negotiate, the same flags and defaults as the CAPI core providers. See [Observability]() for what the endpoint serves and how a scraper authenticates to it. `--webhook-port` (default `9443`) is where the manager serves admission webhooks; `--webhook-cert-dir`, `--webhook-cert-name` and `--webhook-key-name` say where it finds the serving certificate that cert-manager issues. The shipped manifests wire all of this together; see [Installation](). ## Logging `--logging-format` (default `text`) also accepts `json`. `-v` sets the log verbosity (default `2`); raise it while diagnosing a problem and lower it back afterward, since higher verbosities log more of each reconcile. `--vmodule` overrides the verbosity for individual source files, and only works with the text format. `--feature-gates` takes a comma-separated `key=value` list of the logging feature gates (`ContextualLogging`, `LoggingBetaOptions`, `LoggingAlphaOptions`); CAPTF itself registers no feature of its own. See [Manager Flags]() for the full flag and feature gate list. > [!NOTE] > > **See also** > > - [Manager Flags]() — every flag, its type and its default. > - [Installation]() — installing the provider and its webhook certificate. > - [Upgrades]() — what a provider upgrade changes. > - [RBAC]() — the orphan sweep and the runner’s permissions. > - [Observability]() — metrics, alerts and events. # Upgrades This page covers upgrading an installed CAPTF provider with `clusterctl`, how CAPTF versions its contract, what an upgrade can trigger on existing objects, where to read what changed, and how to roll back. See [Installation]() for installing CAPTF the first time. > [!NOTE] > > **Before you begin** > > - CAPTF already installed with `clusterctl init` (see [Installation]()), and a `clusterctl.yaml` naming the provider (`name: terraform`). > - `clusterctl`, at a version that supports upgrading the Cluster API release you run. ## Check what changed first Before upgrading, read what changed between your installed version and the target one: - The provider’s own release notes on its GitHub release, generated from the commits since the previous tag. - [The module contract changelog](), when the entries mention a contract or inputs-hash change (below); it records every contract-visible change, newest first, whether or not the contract’s own version number moved. ## Upgrade with clusterctl ```sh clusterctl upgrade plan ``` lists the provider versions clusterctl can upgrade each installed component to, grouped by the Cluster API contract they implement. Apply a plan, or name a version for CAPTF directly: ```sh clusterctl upgrade apply --contract v1beta2 clusterctl upgrade apply --infrastructure terraform:vX.Y.Z ``` > [!WARNING] > > **No release is published yet** > > CAPTF has not published a release yet, so there is nothing for `clusterctl upgrade` to find until one exists; until then, changing versions means reinstalling from a local repository, the same way as a first install (see [Installing from a local repository]()). ## Contract versioning CAPTF’s `metadata.yaml` lists its release series for `clusterctl`: each entry is a `(major, minor)` pair and the Cluster API contract it implements. The list is append-only — an entry, once published, is never removed or changed — so `clusterctl upgrade plan` can always resolve an older installed version’s contract. CAPTF has one series so far, `0.1`, implementing contract `v1beta2`; a later series only appears once a release under it exists, and only ever adds to the list. The contract version in `metadata.yaml` is separate from the module contract version (`v1alpha1`, [Module Contract]()): the first is what Cluster API’s own core and other providers require of CAPTF, the second is what CAPTF requires of a module image. A provider upgrade can change either, both, or neither. ## What an upgrade can trigger An upgrade replaces the installed manifests — CRDs, RBAC, the manager and its webhooks — but does not touch a `Terraform*` object’s spec. Things it can still trigger on existing objects, all driven by the manager’s own reconcile after it restarts on the new version: - **A one-time re-apply of every provisioned mutable object.** The inputs hash (`captf.io/inputs-hash`) covers a hash scheme identifier along with the contract version, role, image reference and rendered inputs (see [What is (and isn’t) in the inputs hash]()). A release that changes what the hash covers bumps the scheme, so every existing `TerraformCluster`’s and `TerraformMachinePool`’s inputs hash changes even though nothing in its spec did; the object’s next reconcile sees a hash that no longer matches its state and re-applies once (a `TerraformCluster` with `spec.applyPolicy: Manual` plans once instead and waits for approval; see [Plan Approval]()). A `TerraformMachine` is unaffected: its inputs hash is never recomputed once it is provisioned. - **A one-time cleanup of a `TerraformClusterIdentity`’s credentials Secret.** Earlier versions put an ownerRef on the Secret named in `spec.secretRef`, so deleting the identity garbage-collected it; current versions do not own that Secret. The identity’s reconcile removes a leftover ownerRef the first time it runs after the upgrade, independent of whether anything uses the identity. Let the upgraded manager reconcile every identity once before deleting one, so this cleanup runs first; see [Identities and Credentials](). - **Re-approval of plans waiting under `captf.io/approve-plan`.** Plan hashes now start with `p2:` and bind what each change does, plus output changes, imports and moves. A plan that was waiting for approval under an older `p1:` hash is not approved by the old annotation; approve the new `status.plan.planHash` instead (see [Plan Approval]()). A `status.plan` recorded under `p1:` is re-planned automatically. - **A one-time unknown baseline for pools that applied before exports were recorded.** A `TerraformMachinePool` that applied before CAPTF recorded its applied cluster exports has no baseline for the exports guard. On its first reconcile after the upgrade, CAPTF records the exports it can prove the pool applied (from its newest successful apply, or from its current inputs when no apply Job is retained and they match the state). Without that proof, every apply of the pool is guarded until one succeeds: a plan that deletes or replaces nothing applies and records the baseline, and a destructive plan waits for approval, with `ApplyJobSucceeded` saying the exports are unknown and giving the command. See [Machine pools](). The first two are examples, not the full list: any change the contract changelog records as covering the inputs hash or an identity’s Secret ownership behaves the same way on the next release that ships it. Read the changelog entries for the version you are moving to, since only they say whether either applies. ## Rolling back `clusterctl` has no dedicated rollback command: rolling back means installing the older version’s manifests the same way you install any version, `clusterctl upgrade apply --infrastructure terraform:vX.Y.Z` naming the earlier tag, or reinstalling from that tag’s local repository. > [!WARNING] > > **Rolling back re-applies every provisioned object again** > > Rolling back reverses an inputs-hash scheme change the same way upgrading applies one: the older manager computes the older scheme, so every provisioned mutable object re-applies once again on its first reconcile after the rollback. `v1alpha1` (the module contract, distinct from the `metadata.yaml` contract above) carries no compatibility guarantee between its own changes ([Versioning]()), so a rollback that crosses a contract-changing release can also mean the older manager and the module images or generated roots from the newer one disagree about a field name or an input; check the contract changelog for the versions in between before rolling back across one. # RBAC CAPTF ships a fixed set of ClusterRoles, one Role and the objects that bind them, and creates one more RoleBinding per tenant namespace as it goes. This page covers every one of them: what the manager itself can do, what the runner can do inside a tenant namespace, how a namespace gets its runner identity, and how CAPTF cleans that identity up again. For what a Terraform or OpenTofu module can do with the runner’s access once it has it, see [Security model](). > [!NOTE] > > **Before you begin** > > - Cluster-admin access to the management cluster, and `kubectl`. > - The manager’s own ClusterRole and ClusterRoleBinding are named `captf-manager-role` and `captf-manager-rolebinding` after `clusterctl init` (see [Installation]()); the runner ClusterRole is `captf-runner` regardless of which cluster or namespace it runs in. ## The manager’s ClusterRole The manager’s ServiceAccount, `captf-controller-manager` in the provider’s namespace, is bound to one ClusterRole scoped to exactly what reconciling `Terraform*` objects needs: | Resource | Verbs | Why | | --- | --- | --- | | `terraformclusters`, `terraformmachines`, `terraformmachinepools`, `terraformclusteridentities`, their `*Template` kinds | get, list, watch, create, update, patch, delete | Reconciles and owns every kind it serves | | The `/status` subresources of every kind above except `terraformmachinepooltemplates` (which has none) | get, update, patch | Writes status separately from the spec | | The `/finalizers` subresources of `terraformclusters`, `terraformmachines`, `terraformmachinepools` and `terraformclusteridentities` (never a `*Template` kind) | update | Adds and removes its own finalizers | | `clusters`, `clusters/status`, `machines/status`, `machinepools/status` | get, list, watch | Reads the Cluster API objects a `Terraform*` object belongs to | | `machines` | get, list, watch, patch | Reads Machines and patches the `cluster.x-k8s.io/remediate-machine` annotation it sets to request remediation | | `machinepools` | get, list, watch, patch | Reads MachinePools and writes back an autoscaled pool’s observed replicas | | `secrets` | get, list, watch, create, update, patch, delete | Reads and writes every Secret in [Secrets]() | | `namespaces` | get, list, watch | Evaluates a `TerraformClusterIdentity`’s `allowedNamespaces` selector | | `configmaps` | get, list, watch | Reads `spec.variablesFrom` ConfigMap sources labeled `captf.io/variables=true` | | `serviceaccounts` | get, list, watch, create, delete | Creates the default runner ServiceAccount, reads any override, and the [orphan sweep](<#the-orphan-sweep>) deletes unused ones | | `rolebindings` | get, list, watch, create, update, delete | Manages the per-namespace runner RoleBinding below | | `clusterroles`, resource name `captf-runner` | bind | Lets the manager bind that ClusterRole without holding its permissions itself | | `leases` | get, list, watch, create, update, delete | Its own run and cluster-operation-gate leases, and cleaning up the backend’s state lock lease on delete | | `jobs` | get, list, watch, create, patch, delete | Creates, watches and prunes the runner Jobs | | `jobs/finalizers` | update | Granted alongside `jobs`; the controller itself never sets a finalizer on a Job | | `pods` | get, list, watch | Reads a Job’s pod status for its outcome and the image digest it ran | | `pods/log` | get, list, watch | Granted alongside `pods`; the controller itself never reads a pod’s logs | | `events`, `events.k8s.io/events` | create, patch | Records reconcile events on `Terraform*` objects | | `authentication.k8s.io/tokenreviews` | create | Backs the authenticated diagnostics endpoint | | `authorization.k8s.io/subjectaccessreviews` | create | Backs the authenticated diagnostics endpoint and the identity webhook’s Secret-read check | The manager is never labeled to aggregate into a broader ClusterRole, and it holds no permission on any resource outside this table. ## The leader-election Role `captf-leader-election-role` is a namespaced Role, scoped to the provider’s namespace, bound to the manager’s ServiceAccount by `captf-leader-election-rolebinding`. It grants `leases` (get, list, watch, create, update, patch, delete) and `events` (create, patch): controller-runtime’s leader election uses a Lease, one per manager, and only matters when `--leader-elect` is on (see [Leader election]()). ## The runner ClusterRole `captf-runner` is the one ClusterRole every tenant namespace’s runner ServiceAccount is bound to. It is cluster-scoped and static: the manager never edits its rules, only binds it per namespace ([below](<#the-per-namespace-runner-serviceaccount-and-rolebinding>)). | Resource | Verbs | Why | | --- | --- | --- | | `secrets` | get, list, create, update, delete | The Terraform/OpenTofu Kubernetes state backend’s own requirements | | `leases` | get, create, update | Takes, releases and force-unlocks the state lock | | `events.k8s.io/events` | create | Reports run progress on the object that owns the Job | The Secret verbs are exactly what the state backend uses, no more: `get`, `create` and `update` read and write the state; `list` is needed because the backend lists its state chunks by label on every read and write, and lists workspaces during `init`; `delete` trims chunks when the state shrinks. None of those verbs can be scoped by name or label, so they apply to every Secret in the namespace — this is the basis of the trust boundary described in [Security model](). `leases` omits `delete`: the backend deletes its lock Lease only when a workspace is deleted, which the runner never does, and the manager deletes it itself once an object’s state is gone. `events.k8s.io/events` grants `create` only, never `patch` or `get`: the runner emits a series of one-shot events ([Reading events]()) and never reads or updates one; with `--runner-events=false` the Jobs pass no event target and nothing calls this permission. No Role is ever created for the runner. The manager binds this ClusterRole through the `bind` verb, restricted to its name, so the manager does not need to hold the ClusterRole’s own permissions itself. ## The per-namespace runner ServiceAccount and RoleBinding The first time a namespace needs a Job, the manager makes sure it has both: - a ServiceAccount, `captf-runner` by default, labeled `captf.io/managed=true`. An operator who pre-creates it themselves keeps whatever else they put on it. - a RoleBinding, always named `captf-runner`, whose `roleRef` is fixed to the `captf-runner` ClusterRole and whose subjects are recomputed on every reconcile: the union of every ServiceAccount some `Terraform*` object in the namespace runs its Job as. > [!NOTE] > > **A hand-written RoleBinding named captf-runner is never taken over** > > A RoleBinding named `captf-runner` that already exists without `captf.io/managed=true` is left alone and never modified: the manager reports an error rather than take over a hand-written binding of the same name. One it does manage is re-created, not patched, if something changes its `roleRef` — `roleRef` is immutable in Kubernetes — and otherwise only has its subject list updated. ## Custom ServiceAccounts and the runner opt-in `spec.jobs.serviceAccountName` (or the same field under a `TerraformCluster`’s `spec.defaults.jobs`) can name a ServiceAccount other than `captf-runner` for a Job to run as. Naming `captf-runner` itself is not an override and needs nothing further. Naming anything else needs that ServiceAccount to already exist **and** carry the label `captf.io/runner=true`; without both, the condition `RunnerRBACReady=False/ServiceAccountNotOptedIn` is set and no Job runs. Once it opts in, the manager adds it to the `captf-runner` RoleBinding’s subjects alongside every other ServiceAccount the namespace’s objects use; removing the label removes it from the binding again on the next reconcile. > [!WARNING] > > **The runner opt-in label is consent, not authorization** > > The label is consent, not authorization: whoever can label a ServiceAccount opts it in, whether or not they administer the namespace. It grants little by itself, because anyone who can already create Pods running as that ServiceAccount can read the namespace’s Secrets through a volume mount anyway; what it prevents is a ServiceAccount gaining the runner’s Secret *write* access merely by being named in `spec.jobs.serviceAccountName`. ## The orphan sweep At manager start, and then every `--sync-period` while this manager leads (see [Sync period and the orphan sweep]()), the manager deletes the `captf.io/managed=true` ServiceAccounts, RoleBindings and Leases of every namespace holding no `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` — what a [`clusterctl move`]() leaves behind in the source namespace, since `clusterctl` strips finalizers before deleting the source objects and the controller’s own delete-time cleanup never runs there. The decision is made through the manager’s uncached reader with no `--watch-filter` applied, so another manager instance’s objects, even ones this manager does not watch, still keep a namespace from being swept. Nothing is deleted for a single object here, so the sweep logs what it removed rather than recording an event. A namespace that still holds objects keeps its runner ServiceAccount and RoleBinding, but the RoleBinding’s subjects are pruned to the ServiceAccounts still in use: a `Terraform*` object’s own `spec.jobs.serviceAccountName`, else, for a machine or pool, its cluster’s `spec.defaults.jobs.serviceAccountName`, else `captf-runner`. This is how switching `spec.jobs.serviceAccountName` from one opted-in ServiceAccount to another eventually drops the old one’s Secret access, rather than leaving it bound until the whole namespace empties out. When a machine or pool’s cluster cannot be resolved yet, every cluster default in the namespace is kept rather than pruned, so a resolution delay never flaps the binding. ## The metrics-reader ClusterRole `captf-metrics-reader` lets a Prometheus instance read the manager’s authenticated `/metrics` endpoint; it ships only with the opt-in Prometheus component, not with the base install. See [Enabling the Prometheus component]() for what it binds to and how to point it at your own Prometheus. > [!NOTE] > > **See also** > > - [Security model]() > - [Secrets]() > - [Identities and Credentials]() > - [Configuration]() > - [`clusterctl move`]() # Secrets This page inventories every Secret CAPTF reads or writes: what it is named, which namespace it lives in, who creates and deletes it, what it contains, and whether it carries sensitive data. Read it before deciding who may read Secrets in a tenant namespace, or before writing a NetworkPolicy or admission policy that assumes only some Secrets matter. For what the runner’s own access to these Secrets means once a Job is running, see [Security model](). For how these Secrets are created, connected and deleted, with walkthroughs, see [Secret Management](). > [!NOTE] > > **Before you begin** > > - `get` access to Secrets in the namespaces you administer, to inspect any of these. > - Knowing which kind (`TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`), object name or identity a Secret belongs to; several of the name patterns below embed one of these. ## Every Secret CAPTF reads or writes | Name pattern | Namespace | Owner | Created / deleted | Contents | Sensitive | | --- | --- | --- | --- | --- | --- | | Identity’s source Secret (any name, `TerraformClusterIdentity.spec.secretRef`) | `secretRef.namespace`, any namespace | The operator | By the operator; CAPTF never creates, updates or deletes it, and removes any owner reference an older CAPTF version left on it | Cloud credentials (arbitrary keys and values, provider SDK conventions or file contents) | Yes | | `captf-creds-` (mirror) | Each namespace `allowedNamespaces` permits and that has resolved the identity | The `Terraform*` object(s) using the identity there, as non-blocking owner references | Created the first time an object in the namespace resolves the identity; rewritten whenever the source Secret’s data changes; deleted when the namespace stops being allowed, or when its last owning object is removed | A byte-for-byte copy of the source Secret’s data | Yes | | `captf-inputs-c-`, `captf-inputs-m-`, `captf-inputs-mp-` (durable inputs) | The object’s namespace | The `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` itself | Created before the object’s first Job; rewritten on every reconcile that re-renders inputs; deleted once the object’s state and infrastructure are gone | The rendered root module and tfvars, plus the pinned image, image digest and identity | Yes | | `captf-run-` (per-run) | The object’s namespace | The Job | Created just before the Job’s pod starts, from that reconcile’s rendered inputs; deleted once the controller has read the finished Job’s result | The same rendered root module and tfvars the Job runs with | Yes | | `tfstate-default-` and its `-part-N` chunks (state) | The object’s namespace | Unowned until the first successful apply, then the object, as a non-blocking owner reference | Created by the Terraform/OpenTofu Kubernetes state backend itself, not by CAPTF; deleted by CAPTF after a successful destroy, or immediately on deleting an object that never applied | The compressed Terraform state, plus the backend’s own workspace labels | Yes | | `captf-state-backup--` and its `-part-N` chunks (state backups) | The object’s namespace | The `Terraform*` object, as a non-blocking owner reference (never the state) | Created after a reconcile observes a state serial not backed up yet; pruned to a configured retention on the same pass; never deleted by the destroy cleanup above | A verbatim copy of the state Secrets they were taken from, at that serial | Yes | | Bootstrap data Secret (any name, `Machine.spec.bootstrap.dataSecretName` or `MachinePool.spec.template.spec.bootstrap.dataSecretName`) | The Machine’s or MachinePool’s namespace | The bootstrap provider (for example a `KubeadmConfig`), not CAPTF | By the bootstrap provider; CAPTF only reads it, uncached, and never labels, updates or deletes it | The bootstrap data (`value`) and its format (`format`) | Yes | | `spec.variablesFrom` source (any name, labeled `captf.io/variables=true`) | The object’s namespace | The operator | By the operator; CAPTF only reads it, and only while the label is present — an unlabeled Secret counts as missing | Arbitrary keys treated as module variable values | Yes | | Image pull secret (any name, named in `spec.jobs.imagePullSecrets`) | The object’s namespace | The operator | By the operator; for a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`, CAPTF only references its name on the Job’s pod, never reading or writing its contents. For a `TerraformMachineTemplate`, CAPTF also reads its own pull Secrets’ contents to authenticate the image inspection that resolves capacity, but never writes them | A `kubernetes.io/dockerconfigjson` registry credential | Yes | | `captf-webhook-service-cert` (webhook serving certificate) | The provider’s namespace | cert-manager, through its `Certificate` object | By cert-manager, on issuance and renewal; CAPTF never creates, reads or deletes it | A TLS key pair and CA bundle for the admission webhook | Yes | A name pattern that would exceed the 253-character Secret name limit — the mirror and the durable inputs Secret, both built from a user-chosen name — is shortened to its prefix plus a hash, still deterministic and unique. ## What the manager caches The manager’s main cache holds only Secrets labeled `captf.io/managed=true`: credential mirrors, durable and per-run inputs, state and state backups — the five Secrets above that carry that label, whether CAPTF or the state backend created them. A second, separate cache backs `spec.variablesFrom` watches; it holds Secrets (and ConfigMaps) labeled `captf.io/variables=true`, with their data stripped out before they are stored, so no variable value ever sits in memory there. Everything else — the identity’s source Secret, bootstrap data, image pull secrets, and a `spec.variablesFrom` Secret’s actual content when it is resolved — is read directly from the API server on demand and never watched. ## What never enters logs or status The durable and per-run inputs Secrets carry bootstrap data and module variable values in clear by design, and CAPTF never logs their data. State and its backups are read the same way: never logged. > [!WARNING] > > **Sensitive variable values are still written in clear into the inputs and state** > > A `spec.variablesFrom` value marked sensitive is redacted from the Job’s own plan and apply output and from the controller’s trace-level logs, but that redaction does not reach the Secrets in this table: like every other input, the value is still written in clear into the durable and per-run inputs Secrets and into the state. This is a separate guarantee from what CAPTF keeps out of `status` and events, which never carry Secret contents at all; see [what CAPTF keeps out of status, events and logs](). > [!NOTE] > > **See also** > > - [Security model]() > - [RBAC]() > - [Identities and Credentials]() > - [Job Inputs]() > - [Terraform State]() > - [Module Variables]() # Multi-Tenancy CAPTF can serve several tenants from one management cluster. The unit of isolation is the **namespace**: one tenant, one identity and one workload cluster per namespace. This page explains what that isolation is made of, where it is thin, and how to lay out a tenant. It builds on [Identities and Credentials](), [RBAC]() and the [Security Model](), and links to them instead of repeating them. ## The model A tenant is a namespace that holds that tenant’s Cluster API and `Terraform*` objects, their Secrets, and the Jobs CAPTF runs for them. What a tenant can do is decided by three things: | Layer | Controls | Where | | --- | --- | --- | | **Kubernetes RBAC** | Who may create or edit `Terraform*` objects, approve plans and read Secrets | Roles in the tenant namespace | | **The identity** | Which cloud credentials a namespace may use | `TerraformClusterIdentity.spec.allowedNamespaces` | | **The namespace itself** | What a Job can read, how much it can use, where it can connect | Pod Security labels, quotas, network policy | > [!WARNING] > > **A module runs with the namespace’s Secret access and the identity’s credentials** > > The point to hold on to is that **a module is code that runs with the namespace’s Secret access and the identity’s cloud credentials**. Whoever controls `spec.source.image` of an object, or can edit one, controls both. Everything below follows from that. ## Identities and `allowedNamespaces` A `TerraformClusterIdentity` is cluster-scoped. It names a credentials Secret (`spec.secretRef`, in any namespace) and says which namespaces may use it. | `allowedNamespaces` | Meaning | | --- | --- | | unset | No namespace may use it | | `list` | Exactly the named namespaces | | `selector` | Namespaces whose labels match | | `selector: {}` | **Every** namespace, including ones not yet created | | `list` and `selector` | The union | | `{}` | Rejected: it would read as “no one” or “everyone” | Protections the admission webhook adds: - **Creating** an identity, **changing** `spec.secretRef` and **widening** `allowedNamespaces` each run a `SubjectAccessReview`: the requester must be allowed to `get` the named Secret. A role that may manage identities but cannot read the Secret cannot point an identity at it, or grant more namespaces access to it. Narrowing is not re-checked. - **Deleting** an identity is refused while any object resolves to it or any namespace still holds its mirror. Practical rules: - Give each tenant its **own identity**, with credentials scoped to that tenant’s cloud account or role, and allow only that tenant’s namespace. Never use `selector: {}` for an identity with real credentials. - Keep the source Secret in a **platform namespace** tenants cannot read. The tenant’s namespace holds only the mirror. - Do not give tenants `create`, `update` or `patch` on `terraformclusteridentities`. The `SubjectAccessReview` limits what a requester can point an identity at, but identities are a platform object. ### The mirror per namespace For each identity a namespace uses, the manager keeps one Secret, `captf-creds-`, in that namespace: a copy of the source Secret. All the namespace’s objects that use the identity share it, and each is an owner of it. The Job mounts it into the module’s container as environment variables and read-only files. - **Revocation.** Drop a namespace from `allowedNamespaces` and the manager deletes the mirror there, whatever still references it, and starts no new Job for the namespace’s objects. A destroy waits, with `IdentityNotAllowed`, until access returns, or the object is abandoned. See [Revoke access](). - **The mirror is a Secret the tenant’s workloads can read.** Anyone with `get` on Secrets in the namespace can read the credentials. Scope the credentials, not just the access. - **Rotation is not instant.** A change to the source reaches the mirror on the manager’s next reconcile. See [Rotate credentials](). ## The runner reads every Secret in its namespace > [!CAUTION] > > **Any module image can read, replace or delete every Secret in its namespace** > > The Job runs as the `captf-runner` ServiceAccount, bound by a RoleBinding in each namespace to a ClusterRole with `get`, `list`, `create`, `update` and `delete` on Secrets, unscoped by name. The Kubernetes state backend needs that: it lists and writes its own chunks, and RBAC cannot name Secrets that do not exist yet. The ServiceAccount token is mounted in the Job pod, and the state backend uses it. As a result, **any module image can read, replace or delete every Secret in its namespace**: other objects’ state and inputs, the credential mirrors, and Cluster API Secrets such as a cluster’s kubeconfig and CA. See [What the runner can read, and why](). Consequences for a platform team: - **Do not put two trust domains in one namespace.** Two tenants, or a trusted platform cluster and an untrusted module author, in one namespace give each the other’s state and credentials. - **An untrusted or experimental module gets its own namespace**, with its own identity whose credentials can do no more than that experiment needs. - **The manager itself reads Secrets cluster-wide**, as any provider that uses the Kubernetes state backend does. Protect the manager’s namespace and its image supply chain accordingly. ### Override ServiceAccounts `spec.jobs.serviceAccountName` can name another ServiceAccount for a Job. It must exist and carry the label `captf.io/runner=true`, or the object reports `RunnerRBACReady=False`/`ServiceAccountNotOptedIn` and runs nothing. The label is consent: whoever may label a ServiceAccount opts it in. The manager then adds it to the namespace’s `captf-runner` RoleBinding, and prunes subjects no object uses. This is for a namespace administrator who needs a Job to run as a ServiceAccount with extra cloud-side permissions (workload identity), not a way to narrow the runner: the binding gives the Secret rights to every subject. See [Custom ServiceAccounts](). ## Who writes specs and who approves Approval is Kubernetes RBAC, by design: the approval annotations (`captf.io/approve-plan`, `captf.io/approve-destructive-plan`) and the manual-action annotations (`captf.io/restore-state`, `captf.io/abandon-infrastructure`) are authorized by who may `patch` the object, and the webhooks do not restrict them. RBAC cannot separate a patch that sets an annotation from one that edits the spec, so whoever holds `patch` on a `TerraformCluster` can do both. The pattern for a tenant, from [Operating the gates](): - **Approvers** hold `patch` on `terraformclusters` in the tenant namespace. - **Spec changes** come from a pipeline’s ServiceAccount, through a reviewed change (GitOps). People hold read-only access otherwise. - **Machines and pools are never gated**, so the split protects only the cluster module. See [Limits](). - An optional admission policy that guards the annotations is outside CAPTF; see that page. Which fields a spec writer can change also differs by kind. A `TerraformMachine`’s source, identity and variables are immutable; its `jobs`, `drift` and `remediation` policy is not. A `TerraformCluster` and a `TerraformMachinePool` are fully mutable by whoever has `update`. ## Quotas, limits and network policy - **Pod Security.** The Job pod meets `baseline` by default. Label each tenant namespace with the level you enforce. For `restricted`, the module image must run as non-root. See [Pod security](). - **ResourceQuota.** A tenant’s quota has to allow what CAPTF creates: Jobs and pods (finished Jobs are retained until pruned, 3 successful and 3 failed per op by default), Secrets (the state and its chunks, up to five backups of each, the durable inputs, the plan key, the mirror, and a per-run Secret per running Job) and Leases (a run lease per object, a write lease per Cluster and a state lock per object). The default Job pod requests 250m CPU and 512Mi, sets a 2Gi memory limit and **no CPU limit**. If the quota covers `limits.cpu`, add a `LimitRange` that defaults a CPU limit, or pods without one are refused (Kubernetes behavior). CAPTF has no quota-specific condition: a refusal appears as a reconcile error or a Job that never starts. See [Production Readiness](). - **Network policy.** The manager’s policy does not cover Job pods. Copy `job-egress-sample.yaml` into each tenant namespace and add the cloud and registry egress that tenant’s modules need. Egress is the control that limits where the identity’s credentials can be sent. See [Network exposure](). - **Several managers.** One manager watches one namespace or all of them, never a chosen set, and the leader-election lease name is fixed. To give groups of tenants their own manager instance, use `--watch-filter` and label the objects, and read [Namespace scoping and `--watch-filter`]() first: the webhook still admits every namespace, and an object no instance watches is never reconciled. ## An example tenant The platform team creates, per tenant `acme`: 1. A **namespace** `tenant-acme`, labeled for Pod Security (`pod-security.kubernetes.io/enforce: baseline`). 2. A **credentials Secret** `acme-cloud` in the platform namespace `captf-credentials`, holding credentials scoped to Acme’s cloud account. 3. An **identity**: ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: acme spec: secretRef: name: acme-cloud namespace: captf-credentials allowedNamespaces: list: - tenant-acme ``` 4. A **ResourceQuota** and **LimitRange** sized as above, and a **NetworkPolicy** for Job pods from the sample. 5. **Roles** in `tenant-acme`. The pipeline’s ServiceAccount authors clusters and machines; approvers are a smaller group with `patch` on `terraformclusters`. This is the pipeline’s Role: ```yaml apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: captf-spec-pipeline namespace: tenant-acme rules: - apiGroups: ["infrastructure.cluster.x-k8s.io"] resources: - terraformclusters - terraformmachines - terraformmachinepools - terraformclustertemplates - terraformmachinetemplates - terraformmachinepooltemplates verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] - apiGroups: ["cluster.x-k8s.io"] resources: ["clusters", "machinedeployments", "machinepools"] verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] - apiGroups: ["batch"] resources: ["jobs"] verbs: ["get", "list", "watch"] - apiGroups: [""] resources: ["pods", "pods/log", "events"] verbs: ["get", "list", "watch"] ``` Note what is **absent**: no `secrets` rule. The pipeline does not read the credential mirror or the state through the API, although the modules they run can (see above). The approvers’ Role is the one in [Operating the gates](), bound to a smaller group. Bind the Roles to groups with RoleBindings in the tenant namespace. Give the tenant no `ClusterRole` that reaches `terraformclusteridentities`. > [!NOTE] > > **See also** > > - [Identities and Credentials](). > - [RBAC]() and [Secrets](). > - [Security Model]() and [Security considerations](). > - [Approvals and Gates](). > - [Production Readiness](). # Production Readiness This page is a go-live checklist for the CAPTF manager and the namespaces it serves. Each item says what to check and why, and links to the page that has the detail. It also says plainly what CAPTF does **not** do for you, because several things a production cluster needs are the operator’s to provide. > [!WARNING] > > **Pre-alpha: treat this as a minimum, not a certification** > > CAPTF is pre-alpha: no release is published yet, and no end-to-end run against a live management cluster has happened (see [Project status]()). Treat the checklist as the minimum for a trial you intend to keep, not as a certification. ## The checklist | Area | Check | Detail | | --- | --- | --- | | Availability | Leader election is on and you know what failover costs | [Manager availability](<#manager-availability>) | | Sizing | Requests and limits fit your object count; concurrency is deliberate | [Sizing](<#sizing>) | | Webhooks | cert-manager is healthy and the serving certificate renews | [Webhook certificates](<#webhook-certificates>) | | Environment | `POD_NAMESPACE` and `SERVICE_ACCOUNT_NAME` are set | [Required environment](<#required-environment>) | | Secrets | etcd encryption is on; backups and DR are planned | [Secrets at rest](<#secrets-at-rest>) | | Monitoring | The eleven alerts are installed and routed | [Alerts and metrics](<#alerts-and-metrics>) | | Access | Who may write specs and who may approve is decided | [Approver and writer split](<#approver-and-writer-split>) | | Network | The manager and the Job pods have egress policies | [Network policy](<#network-policy>) | | Jobs | Pod Security level, Job policy defaults and quotas fit | [Runner Jobs](<#runner-jobs>) | | After go-live | You know what to watch in the first week | [First week](<#the-first-week>) | ## Manager availability - [ ] **Replicas.** The shipped Deployment runs **one** replica, with the leader election flag set. One replica is a valid production setting: while it restarts, running Jobs continue, and the manager picks them up again. No PodDisruptionBudget or anti-affinity is shipped; add them if you scale up. - [ ] **Leader election.** The binary’s `--leader-elect` defaults to `false`; the shipped manifest passes it. If you build your own manifest, enable it before you run more than one replica, or two managers reconcile the same objects. See [Configuration: leader election](). - [ ] **Failover time.** The manager does not release the election lease when it stops, so a replacement waits for the lease to expire: up to 15 seconds with the defaults (lease 15s, renew 10s, retry 2s). A failover loses nothing: Jobs are Kubernetes objects, and the names, the leases and the cache-lag checks make the resumed work idempotent. The webhooks run on every replica, so writes keep working while no manager leads. See [Leader election and failover](). - [ ] **The webhook is on the write path.** All webhooks fail closed. If every manager pod is down, creating or updating a `Terraform*` object fails, including Cluster API’s own writes. See [Webhook Unavailable](). ## Sizing - [ ] **Manager resources.** The shipped requests are 10m CPU and 64Mi of memory, with limits of 500m and 256Mi. They suit a small installation. The manager holds informers for Jobs and managed Secrets, so memory grows with the number of objects, Jobs and Secrets. Watch the manager’s working set and raise the limit before it reaches it. - [ ] **Concurrency.** `--terraformcluster-concurrency` and its machine, pool and template counterparts default to 10 each. They cap reconciles in flight per kind, not Jobs. A reconcile starts a Job and returns, so no flag caps the number of Jobs running at once; the bounds are one Job per object and your namespace quotas. Raise a concurrency flag when objects queue behind each other. See [Configuration: per-kind concurrency](). - [ ] **Resync.** `--sync-period` (10 minutes) is a safety net, not a schedule. Drift runs every `--drift-default-interval` (30 minutes) unless an object sets its own. A large fleet with the default drift interval starts a burst of Jobs every half hour. The jitter spreads it by up to a tenth of the interval; set longer intervals if that is too much. See [Requeue intervals and schedules](). - [ ] **Job pods.** The runner’s default request is 250m CPU and 512Mi, with a 2Gi memory limit and no CPU limit. A large plan needs more: set `spec.jobs.resources`. See [Tuning Jobs](). ## Webhook certificates The validating webhooks are served with a certificate cert-manager issues from a self-signed Issuer and injects into the webhook configuration. The certificate Secret is mounted into the manager pod, and the volume is not optional. - [ ] **cert-manager must be installed** (`cert-manager.io/v1`). Without it the certificate Secret is never created and the manager pod stays in `ContainerCreating`. - [ ] **Watch the certificate.** Check that the `Certificate` is `Ready` and its renewal works; an expired certificate stops all writes to `Terraform*` objects. - [ ] **Bring your own.** `--webhook-cert-dir`, `--webhook-cert-name` and `--webhook-key-name` point the server at another certificate source if you do not use cert-manager. You then own CA injection too. See [Installation]() and [Webhook Unavailable](). ## Required environment The manager needs `POD_NAMESPACE` and `SERVICE_ACCOUNT_NAME`, which the shipped Deployment sets from the downward API. From them it computes its own username. The `TerraformMachine` webhook lets only that user set `spec.providerID`. > [!WARNING] > > **A missing environment variable only logs a warning** > > If either is unset the manager only **logs a warning** and starts. Then the webhook rejects the manager’s own `providerID` write, and no machine finishes provisioning. If you template your own manifest, check the variables are present; look for the warning in the manager’s log at start. `CAPTF_MANAGER_IMAGE` supplies the runner image unless `--runner-image` is set. ## Secrets at rest > [!WARNING] > > **CAPTF adds no encryption of its own** > > Terraform state, state backups, rendered inputs and credential mirrors are Kubernetes Secrets. CAPTF adds **no encryption of its own**. - [ ] **Enable etcd encryption at rest** (an `EncryptionConfiguration`, or your platform’s equivalent) on the management cluster, and protect etcd snapshots as you would the state. See [No encryption at rest of its own](). - [ ] **State backups.** `--state-backups` (default 5) keeps that many copies per object in the same namespace. They protect against a bad apply or a deleted state Secret, **not** against losing the namespace or the cluster. Plan external backups; see [Disaster Recovery](). - [ ] **Size.** A state near 1 MiB per Secret, or rendered inputs near 1,000,000 bytes, need attention before they fail: see [Size Limits](). ## Alerts and metrics The Prometheus component is opt-in and **not** part of the release manifest. Build it from a checkout of the tag you installed, and apply it again after each upgrade. It adds a `ServiceMonitor` and a `PrometheusRule` named `captf-alerts` with eleven alerts. CAPTF ships rules, not routing: connect them to your Alertmanager. | Alert | Severity | Fires when | | --- | --- | --- | | [`CAPTFJobFailing`]() | warning | More than 2 failed or deadline-killed Jobs of a kind and op in 30 minutes | | [`CAPTFDestroyStuck`]() | critical | A destroy failed in the last 30 minutes, for 30 minutes | | [`CAPTFStateUnreadable`]() | critical | A state read error in 15 minutes | | [`CAPTFClusterDrift`]() | warning | A cluster reports drift for an hour | | [`CAPTFForceUnlocks`]() | warning | A stale lock was force-unlocked in the last hour | | [`CAPTFReconcileErrors`]() | warning | A controller errors at more than 0.1 per second for 10 minutes | | [`CAPTFJobSlow`]() | info | p90 Job duration above 30 minutes | | [`CAPTFJobQueueSlow`]() | warning | p90 time from Job creation to start above 5 minutes | | [`CAPTFStateNearSecretLimit`]() | warning | A state above 900 KiB | | [`CAPTFInputsNearLimit`]() | warning | Rendered inputs above 900,000 bytes | | [`CAPTFNoRecentSuccess`]() | warning | No successful refresh or drift for 6 hours | Route `CAPTFDestroyStuck` and `CAPTFStateUnreadable` to someone who can act: both mean infrastructure the controller can no longer manage safely. The ServiceMonitor scrapes over HTTPS with the Prometheus ServiceAccount `prometheus-k8s` in `monitoring`; edit the shipped RBAC if yours differs. See [Observability](). Three metrics have no alert and are worth a dashboard: `captf_jobs_active` (Jobs in flight), `captf_ready` (objects by readiness) and `captf_identity_denied_total` (denied identity use). Also watch `captf_lease_waits_total` for contention and `captf_state_backups_total` for backups being taken. See [Metrics](). ## Approver and writer split > [!WARNING] > > **Approval is Kubernetes RBAC, by design** > > **Approval is Kubernetes RBAC, by design.** Whoever may `patch` a `TerraformCluster` may set its approval annotations (`captf.io/approve-plan`, `captf.io/approve-destructive-plan`) and its manual actions (`captf.io/restore-state`, `captf.io/abandon-infrastructure`). The webhooks do not restrict them. RBAC cannot separate “sets the annotation” from “edits the spec” on one object, so whoever holds `patch` can do both. Decide who holds it: approvers get `patch`, and everyone else’s spec changes come through a reviewed path such as GitOps. See [Who can approve]() for example Roles and [Multi-Tenancy](). Gates also bind only the cluster: machines and pools are never gated, and a change of the cluster’s exports re-applies pools without approval. See [Limits](). ## Network policy - [ ] **The manager.** `config/network-policy` is an opt-in component. It allows inbound TCP 9443 (webhook), 8443 (metrics) and 9440 (probes) and outbound DNS, 6443 and 443; everything else is denied. It needs a CNI that enforces `NetworkPolicy`. Check that your API server’s address and port are covered by those rules, since the policy allows 6443 and 443 to any destination. - [ ] **Job pods.** The component does **not** cover them. A sample (`job-egress-sample.yaml`, not applied) is a starting point to copy into each tenant namespace. Its cloud-provider egress is deliberately left out: add what your modules need, which depends on the cloud and the registry. Jobs hold cloud credentials, so egress is the control that limits where they can go. See [Network exposure](). ## Runner Jobs - [ ] **Pod Security.** The Job pod defaults satisfy the `baseline` profile. `restricted` needs an image that runs as non-root, set through `jobs.podSecurityContext`, because `runAsNonRoot` is not defaulted. Label the tenant namespace with the level you enforce and test a Job in it. See [Pod security](). - [ ] **Job policy defaults.** Defaults are a one-hour deadline, a five-minute lock timeout, three retained successful and three failed Jobs per op, and the resources above. Set cluster-wide values in `spec.defaults.jobs` on the `TerraformCluster` so that machines and pools inherit them. See [Tuning Jobs]() and [Deadlines and lock timeouts](). - [ ] **ResourceQuota.** Quotas can block what CAPTF creates, and **no quota-specific condition, event or metric exists**. A rejected Job create is returned as a reconcile error and gives its leases back. A pod the quota refuses stays attached to a Job that never starts, which would show up as `CAPTFJobQueueSlow`. Per Job, the namespace holds a Job and a pod, a per-run Secret, the state, durable-inputs and backup Secrets, and a few Leases; each namespace also holds a `captf-runner` ServiceAccount and RoleBinding. Size object-count quotas for retained Jobs: finished Jobs stay until pruned. See [Reconcile Errors](). ## The first week | Watch | Why | Where | | --- | --- | --- | | Manager restarts and the start-up log | A missing environment variable or certificate shows here | `kubectl logs`; [Required environment](<#required-environment>) | | `Ready` on every object | The first sign that something is not provisioning | [Troubleshooting by Condition]() | | Failed and slow Jobs | A new module’s first applies fail and run long | [Failing Jobs](), [Slow Jobs]() | | `DestructivePlanBlocked` and `PlanAwaitingApproval` | Your approval path must work | [Approvals and Gates]() | | The first drift report | Drift settings, and whether the module’s first plan is clean | [Drift]() | | State size and backup counts | Growth, and that backups are taken | [Size Limits]() | | Webhook certificate and cert-manager | An expired certificate blocks writes | [Webhook Unavailable]() | | One restore and one delete rehearsal | You find the gaps before an incident | [State Restore](), [Disaster Recovery]() | > [!NOTE] > > **See also** > > - [Installation](), [Configuration]() and [Upgrades](). > - [Known Limitations]() and [Compatibility](). > - [Security Model](). # Observability This page is for operators running the manager: what its metrics endpoint serves and who may read it, how to wire up Prometheus, what each alert means and where to look first, and how to read conditions, events and logs once you are looking at a specific object or Job. ## The diagnostics endpoint The manager serves Prometheus metrics over HTTPS, authenticated and authorized against the API server by default: 1. The bearer token in the request is checked with a `TokenReview`. 2. The token’s access is checked with a `SubjectAccessReview` for `get` on the non-resource URL `/metrics`. > [!WARNING] > > **`--insecure-diagnostics` turns authentication off** > > `--insecure-diagnostics` turns both checks off and serves plain HTTP instead; use it only for local development. [Configuration]() covers that flag and the address, TLS version and cipher flags, and [Manager Flags]() lists them with their defaults. | Request | Result | | --- | --- | | `GET /metrics` without a token | `401 Unauthorized` | | `GET /metrics` with a token authorized for `get` on `/metrics` | `200`, serving `captf_build_info` and the rest of the series | There is no flag for the metrics server’s own certificate: mount a Secret with `tls.crt` and `tls.key` at `/tmp/k8s-metrics-server/serving-certs/` to serve your own, or leave it unmounted and the manager generates a self-signed certificate at startup. The same authenticated endpoint also serves a profiler (`GET /debug/pprof/*`) and an endpoint to change the log level at runtime (`PUT /debug/flags/v`; see [Logs and verbosity](<#logs-and-verbosity>)) once `--insecure-diagnostics` is off. Each request is authorized the same way, against its own non-resource URL and the HTTP method lowercased as the verb (`get` for the profiler, `put` for the log level). Nothing in the shipped manifests grants either, only `/metrics`, so bind your own ClusterRole to reach them. The series themselves are listed in [Metrics](). ## Enabling the Prometheus component `config/prometheus` is an opt-in kustomize component: a metrics `Service`, a `ServiceMonitor` and a `PrometheusRule` with the eleven alerts below, plus a `ClusterRole` that lets a Prometheus instance read the diagnostics endpoint. It is not part of `infrastructure-components.yaml` and is not a release asset, so `clusterctl init` never installs it and `clusterctl upgrade` never touches it: get `config/prometheus` from a checkout of the tag you installed, build it, and apply it yourself, and repeat that after every upgrade to pick up any change to the alert rules. See [RBAC]() for what the `ClusterRole` grants. > [!NOTE] > > **Before you begin** > > - the Prometheus Operator CRDs (`ServiceMonitor`, `PrometheusRule`) are installed in the cluster; > - you know the namespace and name of the ServiceAccount your Prometheus scrapes with, if it is not kube-prometheus’s default (`monitoring/prometheus-k8s`). 1. Build the component on its own — it needs no other resources, and its objects already carry the `captf-` names and the `captf-system` namespace that `config/default` produces: kustomization.yaml ```yaml apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization components: - /config/prometheus ``` ```sh kustomize build >captf-prometheus.yaml ``` 2. If your Prometheus does not run as `monitoring/prometheus-k8s`, patch the `captf-metrics-reader` `ClusterRoleBinding`’s subject to your ServiceAccount instead, before applying. 3. Apply the built manifest with `kubectl apply -f captf-prometheus.yaml`. If your provider install itself uses renamed objects (a non-default `namePrefix` or namespace), patch the `ServiceMonitor`’s and `PrometheusRule`’s selectors and the `Service` reference to match, since a component does not inherit an overlay’s name or namespace transformers. The `ServiceMonitor` scrapes over HTTPS with `insecureSkipVerify` set, since the metrics server’s certificate is self-signed by default; once you mount a CA-signed certificate, set `tlsConfig.ca` instead and drop `insecureSkipVerify`. To confirm it works without waiting on Prometheus, read the endpoint yourself with the same kind of token Prometheus uses: ```sh kubectl port-forward -n captf-system svc/captf-controller-manager-metrics 8443:8443 & curl -sk -H "Authorization: Bearer $(kubectl create token -n --duration=5m)" \ https://localhost:8443/metrics | grep captf_build_info ``` `` and `` name a ServiceAccount already bound to `captf-metrics-reader` (or your own equivalent grant). A `captf_build_info` line back confirms both the endpoint and the token’s authorization; `401` or `403` means the token or the binding, not the manager. `make promtool-check` and `make promtool-test` validate the alert rules before you ship a change to them; see [Make Targets](). ## Useful queries ```promql # p95 Job duration by op, over the last day, for Jobs that ran their course histogram_quantile(0.95, sum by (le, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1d]))) # Slowest runner steps (p90) by kind, op and step histogram_quantile(0.9, sum by (le, kind, op, step) (rate(captf_job_step_duration_seconds_bucket[6h]))) # Failures by error kind and step sum by (op, error_kind, step) (increase(captf_job_errors_total[1d])) # p90 queue time (creation to the source container's start) by op histogram_quantile(0.9, sum by (le, op) (rate(captf_job_queue_seconds_bucket[1h]))) # The ten largest states, and how close each is to one Secret's 1 MiB topk(10, captf_state_bytes) captf_state_bytes / (1024 * 1024) # Objects whose drift check or health refresh is overdue time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 3600 ``` No Grafana dashboard ships with the component; the queries above are the panels one would build. The full series list, with every label, is in [Metrics](). ## Alerts `config/prometheus` ships eleven alerts on the series above. Two things help in reading them: - The failure counters count transitions, not reconciles: they go up once when an object or Job newly enters the bad state, so `increase(...) > 0` means something newly broke, not that it is still broken. - `CAPTFClusterDrift`, `CAPTFStateNearSecretLimit`, `CAPTFInputsNearLimit` and `CAPTFNoRecentSuccess` name the object directly, with `namespace` and `name` labels. The rest aggregate by `kind` (with `op` or `reason`), except `CAPTFReconcileErrors`, which aggregates by `controller` alone; find the affected object through its conditions, as each section below says. Every rule’s expression, `for` and severity are in [Alerts](); each heading below links there. ### CAPTFJobFailing Jobs of a kind and op failed or hit their deadline more than twice in 30 minutes. Find the objects with `ApplyJobSucceeded=False` or `DriftJobSucceeded=False` and read `status.lastRun` and the Job’s logs: ```sh kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \ | jq -r '.items[] | select(.status.conditions[]? | .type == ("ApplyJobSucceeded", "DriftJobSucceeded") and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"' ``` See [Failing Jobs]() and [the rule](). ### CAPTFDestroyStuck An object’s destroy keeps failing while it deletes; its finalizer and state stay in place, so nothing is orphaned. See [Stuck Destroy]() and [the rule](). ### CAPTFClusterDrift A `TerraformCluster`’s last drift check found changes and it has stayed that way for an hour. With `drift.action: Report` an operator decides next; with `Remediate` the remediation apply is failing, check `ApplyJobSucceeded`. See [Drift]() and [the rule](). ### CAPTFStateUnreadable A state Secret turned unreadable. Find the object with `StateReadable=False` and read that condition’s reason and message: ```sh kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \ | jq -r '.items[] | select(.status.conditions[]? | .type == "StateReadable" and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"' ``` See [Unreadable State]() and [the rule](). ### CAPTFForceUnlocks A Job force-unlocked a state lock whose holder pod no longer existed. Find out why the previous Job died before it happens again. See [Stale State Lock]() and [the rule](). ### CAPTFReconcileErrors A controller keeps returning errors from its reconcile loop. Read the manager’s logs for the failing kind. See [Reconcile Errors]() and [the rule](). ### CAPTFJobSlow Jobs of a kind and op are taking longer than expected at the 90th percentile. Compare it with the Jobs’ `activeDeadlineSeconds` and find the slow step with `captf_job_step_duration_seconds`. See [Slow Jobs]() and [the rule](). ### CAPTFJobQueueSlow Jobs are waiting too long between creation and the source container starting: scheduling, image pulls or the runner’s own init copy. Look for Pending runner pods and their events. See [Slow Jobs]() and [the rule](). ### CAPTFStateNearSecretLimit An object’s compressed state is approaching the 1 MiB a Secret can hold. `captf_state_resources` shows how many resources it manages. See [Size Limits]() and [the rule](). ### CAPTFInputsNearLimit An object’s rendered inputs are approaching the size no Job will start past. See [Size Limits]() and [the rule](). ### CAPTFNoRecentSuccess An object’s scheduled drift check or health refresh has not succeeded in six hours. Read `DriftJobSucceeded` and `status.lastRun`, the same as for a failing Job. See [Failing Jobs]() and [the rule](). ## Reading conditions `kubectl describe` on any `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` or `TerraformClusterIdentity` shows its conditions: a type, a status of `True`, `False` or `Unknown`, a reason and a message. Most conditions are normal polarity (`True` is healthy); a few are inverted, such as `Deleting`. `Unknown` most often means the object is waiting on something else to finish, not that anything failed: it does not fail `Ready`. See [retry backoff]() for how a wait like this is treated. Every condition type CAPTF sets, its polarity, and every reason and message it can carry are in [Conditions](). ## Reading events Every stage of an object’s life emits a Kubernetes Event on it: the manager once per transition, and a Job’s runner in real time while it runs. Read them in order with: ```sh kubectl events --for terraformcluster/ -n kubectl events --for terraformmachine/ -n kubectl events --for terraformmachinepool/ -n kubectl events --for terraformclusteridentity/ -n default ``` `TerraformClusterIdentity` is cluster-scoped, but its events still land in the `default` namespace. Add `-o wide` for a `SOURCE` column that distinguishes the manager’s events from a runner’s. Notes never carry credentials, tfvars, output, plan values or raw stderr: a step failure’s note is the runner’s curated summary, not its log output. Every reason, its type and what it means are in [Events](). ### Runner events A Job’s runner posts its own progress (`RunStarted`, `StepStarted`, `StepSucceeded`, `StepFailed`, `PlanSummary`, `ResourcesChanged`, `RunFinished`) as Events on the object the Job is for, related to the Job itself. Emission is best effort: each request has its own short timeout, and after a few consecutive failures the runner stops emitting for the rest of that run without failing or slowing it. Turning `--runner-events` off (see [Configuration]()) skips them entirely, which is worth doing on a large fleet since a single scheduled drift check alone produces several of them. The runner’s own `events` `create` grant is in [RBAC](). ## Logs and verbosity The manager logs at a default verbosity where the usual reconcile flow is visible; raising it shows more detail down to per-request tracing, and lowering it keeps only errors and irreversible actions such as force unlocks. Credentials, bootstrap data, tfvars content and output values are never logged, whatever the level. Change the level without restarting the manager through the same authenticated diagnostics endpoint: ```sh curl -sk -X PUT -H "Authorization: Bearer " --data '' \ https://
/debug/flags/v ``` The token needs its own authorization for `put` on `/debug/flags/v`; the shipped manifests do not grant it. `--v`, `--vmodule` and `--logging-format` set the level, per-file overrides and the output format at startup instead; see [Configuration]() and [Manager Flags](). Read the manager’s own logs, and a Job’s, with: ```sh kubectl logs -n captf-system deploy/captf-controller-manager -c manager kubectl logs -n job/ -c source ``` > [!NOTE] > > **See also** > > - [Metrics]() — every series, its type, labels and meaning. > - [Alerts]() — every rule’s expression, `for` and severity. > - [Conditions]() — every condition, reason and message. > - [Events]() — every event reason, type and meaning. > - [Configuration]() — the flags behind the diagnostics endpoint, runner events and logging. > - [RBAC]() — the manager’s and Prometheus’s RBAC. > - [Runbooks]() — the recovery procedures the alerts above link to. > - [Troubleshooting by Condition]() — from a condition and reason to its cause and action. # Operating the Gates This page is the reference for running the gates day to day: what the conditions and events mean, how retries and upgrades behave, and who can approve. The commands are in [Plan Approval](). ## What you see Everything below is on the `TerraformCluster`. Read it with: ```sh kubectl get terraformcluster -n \ -o jsonpath='{range .status.conditions[?(@.type=="ApplyJobSucceeded")]}{.status}/{.reason}: {.message}{"\n"}{end}' ``` ### `ApplyJobSucceeded` reasons | Status / reason | Meaning | What to do | | --- | --- | --- | | `Unknown`/`PlanAwaitingApproval` | `Manual`: a plan is ready and nothing applies until it is approved. The message has the counts and the exact approve command | Review `status.plan` and the plan Job’s log, then approve | | `Unknown`/`PlanChanged` | The approved apply re-planned, the plan no longer matched, and nothing was applied. The message has the new plan | Review the new plan, then approve the new hash | | `False`/`DestructivePlanBlocked` | `Automatic`: the plan deletes or replaces something and the inputs hash is not approved. The message names the resources and the inputs hash | Review the plan, then approve the inputs hash | | `True`/`ApplySucceeded` | The last apply succeeded | None | | `False`/`ApplyFailed` | The apply failed; the message names the failing step | See [Failing Jobs]() | | `False`/`JobPolicyInvalid` | The merged Job policy is inconsistent (`lockTimeoutSeconds` is not below `activeDeadlineSeconds`), so no Job starts | Fix the timeouts; see [Tuning Jobs]() | | `Unknown`/`WaitingForRunLease` | Another Job of this object holds the run lease | Wait; it retries every 30 seconds | | `Unknown`/`WaitingForMachineOperations` | The cluster’s apply or destroy waits for machine and pool Jobs to finish | Wait | | `Unknown`/`WaitingForClusterOperation` | A machine’s or pool’s apply or destroy waits for the cluster’s Job (seen on those kinds) | Wait | | `False`/`IdentityNotAllowed` | A destroy cannot start because the identity no longer allows the namespace | See [Credentials]() | `PlanAwaitingApproval` and `PlanChanged` are `Unknown` on purpose: a plan waiting for you is not a failure and does not make `Ready` false. The full list of reasons is in [Conditions](). ### Events | Event | Type | When | | --- | --- | --- | | `PlanReady` | Normal | A plan Job planned a change; once per plan Job. For an empty plan its message says the apply runs without an approval | | `PlanApproved` | Normal | The apply of an approved plan started (not for an empty plan) | | `PlanApplied` | Normal | The approved apply succeeded and `status.plan` was cleared | | `PlanChanged` | Warning | The plan changed since it was approved | | `DestructivePlanBlocked` | Warning | A guarded apply stopped on a delete or replace; once per blocked Job | | `DestructivePlanApprovalConsumed` | Normal | The controller removed a destructive-plan approval after an apply | An `InputsChanged` event also marks the start of a plan. The condition transitions for `PlanAwaitingApproval` and `PlanChanged` emit no event of their own. See [Events]() for the full list. `captf_plan_approvals_total{kind,result}` counts applied approvals (`approved`) and changed plans (`changed`), and `captf_destructive_plan_approvals_consumed_total{kind}` counts consumed destructive-plan approvals; see [Metrics](). ## Timing A waiting plan, and a blocked apply, are re-checked at least every ten minutes, and a changed annotation or new inputs re-trigger the reconcile at once. The controller does not validate an approval’s format: it compares the string with the plan hash or the inputs hash, and ignores any other value. ## Upgrades A `status.plan` recorded by an older release has a hash that does not start with `p2:`. After the upgrade the controller plans again by itself, and the new hash needs a fresh approval. An old approval annotation does not match. See [Upgrades](). ## Who can approve **Approval authority is Kubernetes RBAC, by design.** An approval is an annotation on the object, and whoever may `patch` the `TerraformCluster` may set `captf.io/approve-plan`; for the destructive-plan guard, whoever may patch the object may set `captf.io/approve-destructive-plan`. CAPTF’s admission webhook does not restrict these annotations, and will not: the Kubernetes authorization you already run is the single place that decides who approves. There is no second list of approvers to keep in step. > [!WARNING] > > **Whoever holds patch can both change the module and approve the change** > > The consequence to plan for: RBAC cannot look inside a patch, so it cannot separate “sets the annotation” from “edits the spec” on the same object. Whoever holds `patch` can both change what the module does and approve the change. The way to separate the two is to separate **who holds `patch`**, not to restrict an annotation. The recommended pattern: - **Approvers** are the principals with `patch` on `terraformclusters` in the namespace, with read access to Jobs, pods and logs so they can read the plan. - **Spec changes** from everyone else come through a reviewed path, such as a GitOps pipeline whose ServiceAccount holds the write Role. People propose a change in review; an approver reads `status.plan` and the plan log, and annotates. - **Do not give approvers’ rights to automation that edits specs**, and do not sync the approval annotations from Git: a controller that does would re-add a consumed approval. Example Roles, one for the approvers and one for the pipeline: ```yaml apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: captf-plan-approver namespace: rules: - apiGroups: ["infrastructure.cluster.x-k8s.io"] resources: ["terraformclusters"] verbs: ["get", "list", "watch", "patch"] - apiGroups: [""] resources: ["pods"] verbs: ["get", "list"] - apiGroups: [""] resources: ["pods/log"] verbs: ["get"] - apiGroups: ["batch"] resources: ["jobs"] verbs: ["get", "list"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: captf-spec-pipeline namespace: rules: - apiGroups: ["infrastructure.cluster.x-k8s.io"] resources: ["terraformclusters", "terraformmachines", "terraformmachinepools"] verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] ``` Bind `captf-plan-approver` to the approvers’ group and `captf-spec-pipeline` to the pipeline’s ServiceAccount, and give everyone else read-only access. An approver can still change the spec, so give the Role only to people you would trust to do so. Machines and pools are not gated at all, so this split protects the cluster module only; see [Limits](). If you also want an admission policy that rejects the annotations unless the requester is in a group, that is **optional and outside CAPTF**: a Kubernetes `ValidatingAdmissionPolicy`, Kyverno or Gatekeeper rule can do it, and must still let the manager’s own ServiceAccount remove a consumed approval. CAPTF ships none and does not need one. > [!NOTE] > > **See also** > > - [Plan Approval](). > - [Manual plan approval]() and [The destructive-plan guard](). > - [RBAC](). # Other Manual Actions Not every wait is an approval. Some conditions need a person to do something that is not “approve the plan”. This page lists them, with what CAPTF does and where the procedure is. ## Restore a state `captf.io/restore-state=` asks for a backup to be pushed back as the state. A value that is not a serial, or that names no complete backup, sets `RestoreJobSucceeded=False`/`RestoreBackupNotFound` and starts nothing. A restore takes precedence over apply, drift and refresh, waits on the run lease and, for a cluster, the cluster write lease, and is never gated by an approval. See [Backups and restore]() and the [state restore runbook](). A failed restore is not retried for the same serial, and does not count toward backoff. To retry it: remove the annotation, wait until `status.lastRestoredSerial` clears, then set the annotation again. Deleting the failed Job does not retry. ## Abandon an object `captf.io/abandon-infrastructure=` releases a deleting object whose destroy cannot run: held on missing or unreadable state, a failed last destroy, a missing durable inputs Secret, or an identity or credentials the destroy cannot start with. The value must equal the UID; anything else is ignored. > [!CAUTION] > > **Abandoning leaves the infrastructure running, untracked** > > The infrastructure keeps running, untracked. An object whose destroy can run is destroyed as usual, even with the annotation set. See the [stuck destroy runbook](). ## Fix an inconsistent Job policy `ApplyJobSucceeded=False`/`JobPolicyInvalid` means the merged Job policy has a `lockTimeoutSeconds` that is not below `activeDeadlineSeconds`, counting inherited and built-in defaults (300 and 3600 seconds). No Job other than a destroy starts until you change one of the two values. See [Tuning Jobs](). ## Release a foreign state lock `StateReadable=False`/`StateLocked` means the state lock is held by something that is not one of the object’s own runner pods, for example a workstation running `terraform state rm`. The controller does not unlock it: every Job for the object waits `lockTimeoutSeconds` for it and fails. A lock held by the object’s own pod that is now gone is unlocked by the next Job on its own. See the [stale lock runbook](). ## Opt a ServiceAccount in A Job runs as the ServiceAccount `captf-runner`. If you set another ServiceAccount in `spec.jobs.serviceAccountName`, it must exist and carry the label `captf.io/runner=true`. Otherwise `RunnerRBACReady` is `False`/`ServiceAccountNotOptedIn`, the ServiceAccount is not bound to the runner role, no Job starts and the controller retries every 30 seconds. Add the label, or use the default. See [RBAC](). > [!NOTE] > > **See also** > > - [Runbooks](). > - [Operating the gates](). > - [Deletion and Teardown]() for the whole delete path. # Disaster Recovery This page covers recovering from the loss of Terraform state, of a namespace, or of the whole management cluster. It starts with what CAPTF does and does not give you, then lists what to back up outside the cluster, how to rebuild, and a drill to rehearse it. > [!WARNING] > > **Most of this page is untested guidance** > > Most of this page is **guidance derived from how the controller behaves**, not a tested procedure. CAPTF is pre-alpha, and no recovery has been run against a live management cluster. Where a step depends on a tool outside CAPTF (Velero, `kubectl`, `jq`) it is marked as guidance: check it in a scratch cluster before you rely on it. The drill at the end is how. ## What the in-cluster backups are not CAPTF keeps up to `--state-backups` (default 5) copies of each object’s state as Secrets named `captf-state-backup--`. They are useful for a state Secret that was deleted or damaged. They are **not** disaster recovery: - They are **in the same namespace**, so deleting or losing the namespace loses them with the state. - They are **owned by the object**, so Kubernetes garbage-collects them when it goes, including after a finalizer is removed by hand. - They are **taken after each write** by the controller, when it sees a new state serial. There is no copy taken before an apply, and a run of bad applies pushes the last good copy out of the window. - **Anyone who can delete Secrets in the namespace can delete them**, and so can a module running in a Job. See [Backups are not disaster recovery]() and [Backups and restore](). For what the controller does with them when the state is lost, see [State Restore](). ## What to back up outside the cluster Back up the things that cannot be recreated from the rest. The table says what each CAPTF object is, how to select it, and what happens if it is lost. | Object | Select by | If lost | | --- | --- | --- | | **State Secrets** `tfstate-default-` and `-part-N` | Label `tfstate=true` | The state is gone: `StateReadable=False`/`StateLost` for an object that applied, no Job runs, a deletion is held | | **State backups** `captf-state-backup-*` | Label `captf.io/state-backup=true` | No in-cluster restore source | | **Durable inputs** `captf-inputs--` | Name prefix `captf-inputs-` (label `captf.io/managed=true` with the owner-kind label) | The `captf.io/applied` marker, the pinned digest and the pinned identity go; a `TerraformMachine` can no longer render a destroy (`DestroyFailed`) | | **`Terraform*` objects** | The kinds | The spec is the module’s inputs; without it nothing can be rebuilt | | **`TerraformClusterIdentity`** (cluster-scoped) | The kind | Objects report `IdentityNotFound`; no Job, not even a destroy | | **Identity source Secrets** | The Secret named by each identity’s `spec.secretRef` | `SecretNotFound` on the identity and every object using it | | **`variablesFrom` ConfigMaps and Secrets** | The labeled sources | `VariablesSourceNotFound`; mutable kinds stop | | **Cluster API objects** (`Cluster`, `Machine`, `MachineDeployment`, `MachinePool`, control plane and bootstrap objects) | The kinds | The owners of the `Terraform*` objects | | **Module images**, by digest | Your registry | A destroy needs the pinned image to exist | Not worth backing up, because the controller re-creates them: the credential mirror `captf-creds-*`, the plan key `captf-plankey-*`, the run and cluster Leases, the per-run Secrets, the Jobs, and the runner ServiceAccount and RoleBinding. Losing the plan key invalidates the hash of every plan that waits for approval; the next reconcile makes a new plan to approve. Notes on the selectors: - `captf.io/managed=true` matches all of these Secrets and also Leases, the runner ServiceAccount and RoleBinding and the per-run Secrets. If you use it, exclude Leases and the `captf-run-*` Secrets. The per-run Secrets hold rendered variables and are transient. - The state Secrets and the backups are plain Kubernetes Secrets: protect the backup as you protect the Secrets. See [No encryption at rest of its own](). - The suffix in a Secret’s name is derived from the namespace, the kind and the object’s name, not from its UID, so it is the same after a restore. ### Two ways to take the copy **An etcd snapshot** captures everything, including owner-reference UIDs, so a restore returns the cluster to a point in time with nothing to repair. The cost is that it is the whole cluster at that moment. **A per-namespace backup** (Velero or `kubectl`) is finer but has one trap: owner references. Every CAPTF Secret above, except the identity’s, is owned by its `Terraform*` object, by UID. A restored Secret that still names the **old** UID has an owner that does not exist, and Kubernetes garbage collection deletes it. Restore the Secrets without their `ownerReferences`: CAPTF re-owns them. Each object’s first unpaused reconcile with no active Job gives its state Secrets, backups, durable inputs and plan key one owner reference to the object’s new UID, and replaces its old entry in the credential mirror, emitting an `OwnerReferencesRepaired` event. A Secret that still names the old UID would be repaired too if it lasted that long, but garbage collection normally deletes it first, and nothing is re-owned while the object is paused. Take the copy so that you can do this. For example, with `kubectl` and `jq` (guidance, not a tested procedure): ```sh ns= for sel in tfstate=true captf.io/state-backup=true; do kubectl get secret -n "$ns" -l "$sel" -o json \ | jq 'del(.items[].metadata.ownerReferences, .items[].metadata.uid, .items[].metadata.resourceVersion, .items[].metadata.creationTimestamp, .items[].metadata.managedFields)' > "backup-$sel.json" done ``` Do the same for each `captf-inputs-*` Secret. If you use Velero, check how your version restores `ownerReferences` before you rely on it, and test with one namespace. How often to take the copy decides how much you lose. The state changes on every apply, so a copy older than the last apply restores an out-of-date state: see [After a restore](<#after-a-restore>). ## Rebuild a lost management cluster Infrastructure that Terraform created keeps running when the management cluster is lost; only the record of it is gone. The order below matters because of how the controller treats a **new** object whose state it cannot see. > [!CAUTION] > > **A new object with no visible state creates a second set of resources** > > A new `Terraform*` object whose state Secret is missing and that shows no sign of an earlier apply reads as “never applied”, and its first apply creates a **second** set of resources next to the live ones. The signs of an earlier apply are the `captf.io/applied` marker or a pinned digest on the durable inputs Secret, or any state backup. With any of them the missing state reads as lost and nothing runs. So restore the Secrets **before** the objects can reconcile, or restore the objects paused. 1. **Install the platform.** Create a management cluster, install cert-manager and Cluster API, and install CAPTF at the same version, so the CRDs match the backed-up objects. See [Installation](). 2. **Restore the identities and credentials.** The `TerraformClusterIdentity` objects and their source Secrets. `variablesFrom` sources go back too. 3. **Restore the Cluster API objects.** The `Cluster` and the machine owners come first: a `Terraform*` object whose owner is missing waits at `DependenciesReady`, and one whose owner reference names a UID that no longer matches reports `OwnerMismatch` and runs nothing. Restore them with the Cluster paused (`spec.paused: true`), so nothing starts yet. 4. **Restore the Secrets**: the state, the backups and the durable inputs, **without their `ownerReferences`**. CAPTF adds the owner references back on each object’s first unpaused reconcile. 5. **Restore the `Terraform*` objects, paused.** Add the `cluster.x-k8s.io/paused` annotation to each, or keep the Cluster paused. A paused object runs bookkeeping only. 6. **Check before you unpause.** For each object: `kubectl get` shows the state Secret and the durable inputs; the suffix in `status.stateSecretSuffix` matches the Secret names. 7. **Unpause one object, then the rest.** Start with a machine or a small cluster. Afterwards look for an `OwnerReferencesRepaired` event, or for owner references to the new UID on its Secrets. On the first reconcile the controller reads the state and compares the inputs hash the state recorded with the hash of the current inputs: - They match when the spec, the module image and CAPTF’s rendering are unchanged since the last apply. **No apply runs.** Drift checks follow on their schedule. - They differ when anything changed. A mutable kind (a cluster or pool) applies; under `applyPolicy: Manual` it plans and waits for you. The repair skips an object while it is paused, and skips its state Secrets while a Job holds the run lease, because the backend’s own write would conflict; it retries on the next reconcile. A repair that fails is logged and tried again, and never fails the reconcile. A normal destroy removes the state, the durable inputs and the plan key by label, whoever owns them. ### After a restore - **A state older than the infrastructure.** If the state predates changes Terraform made after it, the live resources it does not list are untracked. Run a drift check (`drift.action: Report`) before you let anything apply, and compare the report with the cloud’s own inventory. - **`captf.io/abandon-infrastructure`** takes the object’s UID. A restored object has a new UID, so a value copied from the old one does nothing. ## Move instead of rebuild If the old management cluster is still running, `clusterctl move` is the supported way to bring objects to a new one. It carries the state Secrets, the backups and the durable inputs, and it rewrites owner references for you. A file-level restore leaves that to CAPTF’s own repair, above. You must copy the identity source Secrets yourself. See [clusterctl move]() and [clusterctl move]() for what moves and what does not. ## Recover from total state loss State is gone, with no external copy and no usable in-cluster backup. Choose by what you have: | You have | Do | | --- | --- | | An in-cluster backup | Annotate `captf.io/restore-state=`. See [State Restore]() | | An external copy of the state Secrets | Restore them without `ownerReferences`, as above; the next reconcile reads them | | A state file from elsewhere | `terraform state push` it to the object’s backend, then reconcile; see [Manual recovery]() | | Nothing | Re-import, or abandon: see [Total State Loss and Import]() | **Re-import** and the other routes, with commands and the module side, are in the [Total State Loss and Import runbook](). In short: for an object that applied before, every Job is held while its state is gone, so a module’s `import` blocks cannot run; rebuild the state from a workstation (route 1), or abandon the object and recreate it with `import` blocks (route 2). An object that never applied can import on its first apply. That flow has not been exercised end to end, and whichever route you use, run a drift check before anything applies. > [!CAUTION] > > **Abandon leaves the infrastructure running and untracked** > > **Abandon** applies to a deleting object. If the object must go and its state is lost or unreadable, the abandon annotation releases the finalizer without a destroy and leaves the infrastructure running and untracked. Clean it up through the cloud. See [Held deletions](). ## What a deleted namespace loses Every Secret CAPTF keeps is in the object’s namespace, so a deleted namespace removes the state, **the backups**, the durable inputs, the plan key, the mirror, the Leases, the Jobs and the runner ServiceAccount, while the cloud resources stay. An object whose status marks it provisioned and whose state is gone is held (`StateLost`) and can be abandoned. A *moved* object looks never-applied once its Secrets are gone; see [Terminating namespaces](). Only the cluster-scoped identity and a source Secret in another namespace survive. ## A recovery drill Rehearse on a scratch management cluster with the `noop` module, which creates nothing in a cloud, so the drill costs nothing and cannot damage anything. Repeat it after each CAPTF upgrade and when you change the backup method. 1. **Set up.** Install CAPTF in a scratch cluster. Create an identity, a `TerraformCluster` and a `TerraformMachine` on the `noop` module and wait for `Ready=True`. 2. **Back up** with the method you will use in production: the etcd snapshot, or the per-namespace export. 3. **Destroy the cluster**, or delete the namespace with the objects in it. 4. **Restore** following [Rebuild a lost management cluster](<#rebuild-a-lost-management-cluster>) on a fresh scratch cluster. 5. **Verify.** Each object reaches `StateReadable=True` and `Ready=True`; **no `apply` Job started** (`kubectl get jobs`, and the `JobCreated` events); a drift check reports `NoDrift`. 6. **Break it on purpose.** Restore once with the Secrets’ original `ownerReferences` and watch them get collected; restore once with objects unpaused and no state, and watch what happens when the durable inputs are present and when they are not. You learn the failure modes in a place where they are harmless. 7. **Record** the time it took and what you changed in this page’s steps for your environment. > [!NOTE] > > **See also** > > - [Production Readiness](). > - [Secrets]() and [Operator files and settings](). > - [Stuck Destroy]() and [Unreadable State](). > - [Deletion and Teardown](). # State Restore The `kubernetes` backend keeps only the latest state for each object; there is no history to roll back to inside the backend itself. CAPTF keeps **versioned backups** of every object’s state instead, and can push one back into the backend when you ask it to. This applies the same way to a `TerraformCluster`, a `TerraformMachine` and a `TerraformMachinePool`: all three keep backups and restore the same way. > [!NOTE] > > **Before you begin** > > - `get` access to the object and its Secrets, and `patch` access to the object (`kubectl annotate` patches it), in its namespace, on the management cluster. > - Replace ``, `` and `` below with the object’s namespace, name and Kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`; the lowercase singular works with `kubectl`). `` is `status.stateSecretSuffix`. Run every command against the management cluster. ## When to use it - `StateReadable=False`/`StateLost`: the state Secret of an object that applied before is gone. - `StateReadable=False`/`StateCorrupt` or `StateInconsistent` that does not clear on its own (a chunk deleted or overwritten by something other than the runner). - A state overwritten with the wrong content — someone ran `terraform state push` or `state rm` from a workstation against the wrong object. > [!CAUTION] > > **A restore makes Terraform forget every resource created after the backup** > > Restore is for disasters. It rewrites the object’s state, so every resource created **after** the backup’s serial is no longer in the state: Terraform forgets it, and it keeps running unmanaged in the cloud. The next plan shows the difference (for a machine, the next drift check). Do not use a restore to undo an ordinary change: change the inputs back instead. For how backups are taken, how many are kept, and what is never backed up, see [Terraform State](). ## List the backups `status.stateBackups` lists up to 16 complete backups, newest first (serial, time taken, compressed bytes — see [Common Fields]() for its fields). It is refreshed when a backup is taken or pruned and when a restore is requested. ```sh kubectl get -n -o jsonpath='{.status.stateBackups}' kubectl get secrets -n -l captf.io/state-backup=true,captf.io/state-backup-suffix= \ -o custom-columns=NAME:.metadata.name,SERIAL:.metadata.annotations.captf\.io/state-backup-serial,TAKEN:.metadata.annotations.captf\.io/state-backup-taken-at ``` `` is `terraformcluster`, `terraformmachine` or `terraformmachinepool`; `` is `status.stateSecretSuffix`. ## Restore Pick the serial and annotate the object: ```sh kubectl annotate -n captf.io/restore-state= --overwrite ``` What happens: 1. The controller finds the newest complete backup of that serial. None, or a value that is not a serial: `RestoreJobSucceeded=False`/`RestoreBackupNotFound`, no Job. 2. A restore takes precedence over apply, drift and refresh. A deleting object normally never restores, with one exception: when deletion is held because the state is lost or unreadable (see [deleting while state is unreadable]()), the restore runs first and the normal destroy follows. A restore waits for a running Job to finish, then takes the object’s run lease (and, for a `TerraformCluster` under the cluster operation gate, the cluster write lease; a machine’s or pool’s restore waits for its `TerraformCluster`’s apply or destroy) like an apply — see [run leases and the cluster operation gate](). A wait shows as `RestoreJobSucceeded=Unknown`/`WaitingForRunLease` (or the cluster-gate reasons). 3. Before the Job starts, the controller re-reads the annotation live, so a stale cached object cannot start a second restore. A restore Job starts. Its root module declares only the `kubernetes` backend, so it needs neither the durable inputs nor any provider. It runs `init` (backend config as usual), `state push -force` (holding the state lock; `-force` because the backup’s serial is older, or the lineage differs, or there is no state at all), and `state list`. The Job fails if the listed state has no managed resource although the backup had some. 4. On success: `RestoreJobSucceeded=True`/`StateRestored`, one `StateRestored` event, the annotation is removed, `status.lastRun` names the restore Job, and the next reconcile reads the restored state as usual, adopting the inputs hash the backup carried (`StateAdopted` — see [Events]()). The push writes back exactly the backup’s serial rather than advancing it, and that serial is already backed up with this same content, so nothing new is written to `status.stateBackups`. 5. On failure: `RestoreJobSucceeded=False`/`RestoreFailed`, one `StateRestoreFailed` warning event. The restore is **not retried** for the same serial. When a restore Job finishes, succeeded or failed, the controller records the serial it consumed in `status.lastRestoredSerial` and refuses to run `captf.io/restore-state` for that serial again. Deleting the failed restore Job does not trigger a retry. To try the same serial again, remove the annotation, wait until `status.lastRestoredSerial` is cleared (the controller clears it once the annotation is gone), then set the annotation again; or set another serial. ```sh kubectl annotate -n captf.io/restore-state- kubectl get -n -o jsonpath='{.status.lastRestoredSerial}' ``` ## Confirm > [!TIP] > > ```sh > kubectl get -n \ > -o jsonpath='{range .status.conditions[*]}{.type}={.status}/{.reason}{"\n"}{end}' > kubectl logs job/ -n -c source # state list output > kubectl events --for / -n > ``` > > Expect `StateReadable=True`/`StateRead`, `status.observedStateSerial` at exactly the restored serial (the push writes it back unchanged, it does not advance it), and `RestoreJobSucceeded=True`/`StateRestored`. Then look at the difference between the restored state and reality **before** anything applies it: - A `TerraformCluster` or `TerraformMachinePool` (both mutable) whose current inputs differ from the restored state’s inputs hash applies them right after the restore (`InputsChanged`). Only the `TerraformCluster`’s apply is guarded: a plan that deletes or replaces anything stops for approval (see [The destructive-plan guard]()), so read the blocked Job’s plan before approving it; a `TerraformMachinePool` applies unguarded. Do not pause the object to prevent this: a paused object starts no Job, the restore included. - Run a drift check to list what the restored state no longer matches; with drift action `Report` nothing is changed. See [Drift](). - Resources created after the backup are not in the state. Import them (`terraform import` from a workstation against the object’s backend) or delete them in the cloud. ## Manual recovery with no backup If `StateLost`, `StateCorrupt` or `StateInconsistent` never clears because no backup was taken before the loss (see the [state-unreadable runbook]()), there is no CAPTF-side restore: reconstruct the state from a workstation against the object’s own backend, then let the next reconcile read it back. The object’s backend is the `kubernetes` backend, configured exactly: ```hcl terraform { backend "kubernetes" { secret_suffix = "" # status.stateSecretSuffix namespace = "" # the object's own namespace in_cluster_config = false # true only from inside the cluster config_path = "~/.kube/config" # a kubeconfig for the management cluster labels = { "captf.infrastructure.cluster.x-k8s.io/owner-kind" = "" "captf.infrastructure.cluster.x-k8s.io/owner-name" = "" "cluster.x-k8s.io/cluster-name" = "" "captf.io/managed" = "true" "clusterctl.cluster.x-k8s.io/move" = "" } } } ``` - `` is exactly `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`. - `` and `` are the object’s own name and its Cluster’s name, each verbatim if 63 characters or fewer, or the first 16 hex characters of its sha256 otherwise (the Kubernetes label value limit) — the same rule `status.stateSecretSuffix`’s own derivation follows; see [Terraform State](). - Getting the `labels` map wrong does not fail loudly: the backend uses it, together with the suffix and workspace, to find the object’s existing state Secrets, so a mismatched value makes it list none and behave as if the object had no state at all, rather than erroring. Match it exactly, including the empty string on the last key. Run `terraform init` with this backend block, then either `terraform import` each resource the cloud provider’s own inventory shows is missing, or `terraform state push ` a state file reconstructed another way. Once it is pushed, the next reconcile reads it back the same as a restored backup (`StateReadable` turns `True`/`StateRead`): run a drift check before anything applies again, since the object’s own record of its last-applied inputs was lost along with the state, and the first reconcile after the push may see the current inputs as changed. ## Caveats - Restoring an older state makes Terraform forget every resource created after that serial: those resources become unmanaged and are neither updated nor destroyed with the object. - A backup is taken of what the backend stored. A state that was already wrong when it was written (a bad apply) is backed up just as faithfully. - A backup older than the object’s current inputs makes a `TerraformCluster` or `TerraformMachinePool` re-apply its current inputs after the restore. - A paused object (or Cluster) starts no restore until it is resumed. - Backups share the namespace’s Secret quota with everything else. > [!NOTE] > > **See also** > > - [Terraform State]() — the backend, Secret naming and how backups are taken and pruned. > - [Stuck destroy runbook]() — recovering when the object cannot be destroyed at all. > - [Conditions reference]() — every `RestoreJobSucceeded` reason. # Total State Loss and Import An object’s Terraform state is gone, and there is no backup to restore. The infrastructure still exists in the cloud. This runbook says what works today, in the order to try it, and how `import` blocks in a module fit in. > [!WARNING] > > **This flow has not been exercised against a live cluster** > > **None of this flow has been exercised end to end against a live management cluster.** Each step is derived from the controller’s code and marked where it depends on behavior outside CAPTF (Terraform, Cluster API, `kubectl`). Rehearse it on a scratch cluster first; the [recovery drill]() is the place. > [!NOTE] > > **Before you begin** > > - Replace ``, `` and `` with the object’s namespace, Kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`) and name. > - You need `kubectl` access to the management cluster, and Terraform or OpenTofu with the module and its providers, usually through the module’s own image (see [step 1 of route 1](<#route-1-rebuild-the-state-from-a-workstation>)). ## Which case is it ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A[State is gone] --> B{An in-cluster or
external copy?} B -- yes --> R[Restore it: State Restore] B -- no --> C{Did the object ever apply?} C -- never --> N[Import blocks work on the first apply] C -- yes --> H[StateLost: every Job is held] H --> R1[Route 1: rebuild the state
from a workstation] H --> R2[Route 2: abandon and recreate
with import blocks] ``` “Ever applied” means **any** of: `status.initialization.provisioned`, the `captf.io/applied` marker on the durable inputs Secret, a pinned image digest on it, or any state backup Secret of the object. With any of them, a missing state reads as `StateReadable=False`/`StateLost` and **no Job runs**: not an apply, not a refresh, not a drift check. So an `import` block in the module cannot run, because nothing runs the module. No annotation clears the marker. ## A copy exists Restore it. With an in-cluster backup, annotate `captf.io/restore-state=`; see [State Restore](). With an external copy of the state Secrets, restore them without their `ownerReferences`; see [Disaster Recovery](). ## The object never applied A new object, or one that never reached a successful apply and has no durable inputs, backup or pinned digest, has no state and reads `StateNotFound`. Its first apply runs the module, so `import` blocks run on it. Two properties of that first apply: - The first apply of an object with no state is **not gated** by `applyPolicy: Manual`. It runs at once, so a wrong import id adopts the wrong resource without a prompt. Under `Automatic` the destructive-plan guard runs but imports are not deletions, so it does not stop them. - Read the plan before you create the object: run the module locally with the same variables and look for `(import)` lines. ## Route 1: rebuild the state from a workstation This keeps the object, its name and its Cluster API links as they are. You re-create the state with Terraform, push it to the backend CAPTF reads, and CAPTF carries on. The object has no Job running while `StateLost` holds, so nothing contends for the state. If it is a pool or cluster, pause its Cluster (`spec.paused: true`) as an extra guard: a paused object starts no Job. 1. **Get what you need.** ```sh kubectl get -n -o jsonpath='{.status.stateSecretSuffix}{"\n"}' kubectl get -n -o jsonpath='{.metadata.labels.cluster\.x-k8s\.io/cluster-name}{"\n"}' mkdir root for f in main.tf.json terraform.tfvars.json; do kubectl get secret -n captf-inputs-- \ -o go-template='{{index .data "'$f'" | base64decode}}' > root/$f done ``` `` is `c`, `m` or `mp`. The first two commands give the suffix and the cluster name. The durable inputs Secret holds the rendered root module (`main.tf.json`) and variables (`terraform.tfvars.json`) of the last successful apply, if it still exists. If it is gone too, you must reconstruct the root and variables yourself; see [Runtime Environment](). The rendered root declares an empty `kubernetes` backend, which `init` completes from `-backend-config` flags; you do not write a backend block. 2. **Build the backend flags**, exactly as CAPTF’s Jobs pass them. A label map that differs does not fail: the backend lists the object’s chunks by it and finds none. The labels are: ```text {"captf.infrastructure.cluster.x-k8s.io/owner-kind"="", "captf.infrastructure.cluster.x-k8s.io/owner-name"="", "cluster.x-k8s.io/cluster-name"="", "captf.io/managed"="true", "clusterctl.cluster.x-k8s.io/move"=""} ``` The owner and cluster names are used verbatim up to 63 characters, and otherwise as the first 16 hex characters of their SHA-256. See [Terraform State](). Do not change the workspace: it is `default`. 3. **Initialize in the module’s image**, so the runtime, providers and module are the ones CAPTF uses. The image’s own binary is at `/captf/runtime`, and the module at `/captf/module` (see the [Image Contract]()). Mount the root files and a kubeconfig for the management cluster, and point the backend at it with `in_cluster_config=false`: ```sh run() { docker run --rm -it --entrypoint /captf/runtime \ -v "$PWD/root:/captf/work/root" -v "$KUBECONFIG:/kube/config:ro" \ -w /captf/work/root @ "$@" } run init -input=false \ -backend-config=secret_suffix= -backend-config=namespace= \ -backend-config=in_cluster_config=false -backend-config=config_path=/kube/config \ -backend-config='labels={"captf.infrastructure.cluster.x-k8s.io/owner-kind"="",...}' ``` The `@` is `captf.io/image` and `captf.io/image-digest` from the durable inputs Secret, or `status.source` on the object. This follows [Stuck Destroy](), which runs the same binary this way; only the backend arguments differ (a kubeconfig instead of the pod’s own token). 4. **Import each resource** the cloud inventory shows, in the same container: ```sh run import -var-file=terraform.tfvars.json '
' '' run state list ``` Repeat per resource. Imports write the state to the backend, which creates the `tfstate-default-` Secret with its own labels. 5. **Plan it locally**, with `run plan -var-file=terraform.tfvars.json`, until it shows no changes or only changes you intend. This is the safest moment to find what you missed: CAPTF will apply next. 6. **Let CAPTF read it.** On the next reconcile the state reads, and `StateReadable` turns `True`/`StateRead`. The state has no `captf.io/inputs-hash` annotation yet, so what follows depends on the kind: | Kind | What CAPTF does | | --- | --- | | `TerraformCluster`, `Automatic` | Reason `StateWithoutInputsHash`: an apply of the current inputs starts, guarded by the destructive-plan guard | | `TerraformCluster`, `Manual` | A plan Job runs and the apply waits for approval, unless the plan has no changes | | `TerraformMachinePool` | An apply starts (pools are not gated) | | `TerraformMachine`, provisioned | **StateLost stays**: “the state carries no inputs hash although the object is provisioned”. See below | The apply writes the inputs hash, and CAPTF then owns the state Secrets. If your local plan was clean, that apply changes nothing. For a provisioned **`TerraformMachine`**, which never re-applies, the state must carry a hash to be accepted. > **Setting the inputs hash by hand is guidance, not a supported procedure** > > Setting any non-empty value on the base Secret works in the code, but is guidance, not a supported procedure: > > ```sh > kubectl annotate secret -n tfstate-default- captf.io/inputs-hash=restored > ``` Replacing the machine instead (delete its Machine; the MachineDeployment or control-plane provider creates a new one) is the supported route, and leaves the old instance to clean up by hand. Prefer it for machines. 7. **Unpause** the Cluster, if you paused it, and confirm: see [Confirm it worked](<#confirm-it-worked>). ## Route 2: abandon and recreate Use this when you cannot rebuild the state, or the object cannot continue. It releases the object without a destroy, recreates it, and has the new object’s **first apply adopt the existing infrastructure through `import` blocks**. It needs a module written for it; see [Write import blocks that are safe to leave in](<#write-import-blocks-that-are-safe-to-leave-in>). > [!CAUTION] > > **Abandoning deletes the state and leaves the infrastructure running** > > The abandon deletes the state, the durable inputs and the plan key, and the infrastructure keeps running with nothing managing it until the new object adopts it. A wrong import id adopts the wrong resource. 1. **Unpause the Cluster.** A paused object never runs a delete, so an abandon does nothing while the Cluster is paused. 2. **Delete the dependents first** where there are any. A `TerraformCluster` with machines or pools of its own waits at `DeletionBlocked` and is not released until they are gone, even with the abandon annotation. See [Order and finalizers](). 3. **Delete the object and abandon it.** The annotation is honored for a deletion held on `StateLost`: ```sh kubectl delete -n --wait=false uid=$(kubectl get -n -o jsonpath='{.metadata.uid}') kubectl annotate -n captf.io/abandon-infrastructure="$uid" ``` The controller removes the finalizer, runs its cleanup and emits an `InfrastructureAbandoned` event. The infrastructure keeps running. See [Held deletions](). 4. **Wait for the leftovers to go**, before you reuse the name. The abandon deletes the state, the durable inputs and the plan key. The state backups are removed by garbage collection once the old object is gone: ```sh kubectl get secret -n -l captf.io/state-backup=true ``` Wait until none of the old object’s backups remain. This matters because the state suffix depends on the namespace, kind and name, **not on the UID**, so a new object with the same name has the same suffix and finds the old object’s backups. Any backup still present makes the new object read as “ever applied” with no state, so it is held on `StateLost` again. Delete any such backup by hand. 5. **Recreate the object** with the same spec plus the variables that drive the imports (below). Reusing the name keeps the Cluster API links; a new name needs the links changed. See [Cluster API objects](<#cluster-api-objects>). 6. **The first apply imports.** With no state and no marker, the object reads `StateNotFound` and the apply of the new object runs at once (not gated under `Manual`, see above). The plan lists resources as `(import)`. A wrong id adopts the wrong resource, so check the variables before you create. ### Cluster API objects The right action depends on the kind. Everything in this table is guidance, based on how the objects reference each other, and has not been run. | Kind | Owned by | What to do | | --- | --- | --- | | `TerraformCluster` | The `Cluster`, through `spec.infrastructureRef` | Delete the `TerraformCluster` and create one with the **same name**. The Cluster’s reference is by name, so it points at the new object, and Cluster API sets itself as owner again. The new object waits at `DependenciesReady`/`WaitingForOwner` until it does | | `TerraformMachine` | A `Machine`, through its `infrastructureRef` | The webhook refuses to delete it while a live Machine references it. **Replace the machine**: delete the `Machine`, and let the `MachineDeployment` or control-plane provider create a new one with a new name and a new instance. The old instance must be cleaned up by hand. Importing is not practical: a machine’s variables are fixed by its template, so they cannot name one instance’s id | | `TerraformMachinePool` | A `MachinePool`, through its `infrastructureRef` | Delete it and create one with the same name. The new object waits at `WaitingForOwnerMachinePool` until Cluster API sets the owner | ## Write import blocks that are safe to leave in Route 2 and the never-applied case depend on the module. Terraform’s `import` block is declarative: it runs on every apply, so write it so that leaving it in is harmless. - **Drive the ids from a variable.** Declare a map, default empty, set by `spec.variables` only when adopting: ```hcl variable "adopt_ids" { type = map(string) default = {} } import { for_each = var.adopt_ids to = aws_lb.api id = each.value } ``` `for_each` on an `import` block needs Terraform 1.7 or OpenTofu 1.7 or later; the reference modules only require 1.5. The example’s address and resource type are placeholders. With the default empty map the block does nothing, so it can stay in the module. - **After the adoption has applied, empty the map.** An `import` block for a resource already in the state at that id is a no-op, but an old id that no longer matches fails the plan. Clear the variable once the adoption is done. - **One id per resource.** Name each resource’s id in the map, not a list, so a reordering cannot swap them. - **Pair with `prevent_destroy`** on what an adoption must never remove. - **Imports are not deletes.** The destructive-plan guard does not stop them, and on a `TerraformCluster` under `Manual` an import-only plan after the first apply waits for approval. The plan shows `(import)`. See [Module Design Patterns](). ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}{"\n"}{end}' > kubectl get -n -o jsonpath='{.status.stateSecretSuffix}{"\n"}' > kubectl events -n --for / > ``` > > Expect `StateReadable` `True`/`StateRead`, no unexpected `JobFailed`, and for a cluster or pool an apply that changed nothing. Run a drift check before you let anything change the infrastructure. > [!NOTE] > > **See also** > > - [Disaster Recovery](). > - [State Restore]() and [Unreadable State](). > - [Held deletions](). > - [Module Design Patterns](). # clusterctl move > [!WARNING] > > **Do not leave the source Cluster paused when you run clusterctl move** > > `clusterctl move` refuses to run against a Cluster that is already paused — it stops with an error naming the paused Clusters — and it pauses the Cluster itself for the duration of the move and unpauses it afterwards. So **do not** leave the source Cluster paused when you run it. This page walks through moving a Cluster’s `Terraform*` objects safely, what does and does not come along, and cleaning up what a move leaves behind. > [!NOTE] > > **Before you begin** > > - `clusterctl` configured with both the source and target management cluster’s kubeconfigs. > - The target cluster running the same or a newer CAPTF version. > - Replace `` and `` with the Cluster’s namespace and name below. ## Procedure The documented sequence pauses first only to drain in-flight Jobs, then unpauses before invoking `clusterctl move` (which pauses and unpauses it again on its own): 1. **Pause the Cluster** so the controller starts no new Job on any `Terraform*` object of it: ```sh kubectl patch cluster -n --type=merge -p '{"spec":{"paused":true}}' ``` A Job already running is left to finish, unless it never started at all — its per-run inputs Secret is missing and every pod it created is still `Pending` a minute after the Job was created — in which case the paused reconcile deletes it so the next reconcile starts it again. The paused reconcile still does Job bookkeeping and clears `clusterctl.cluster.x-k8s.io/block-move` once no Job is active, even while paused ([the clusterctl move block]()). 2. **Wait until no `Terraform*` object of the Cluster carries `block-move`:** ```sh kubectl get terraformclusters,terraformmachines,terraformmachinepools -n \ -l cluster.x-k8s.io/cluster-name= \ -o custom-columns='KIND:.kind,NAME:.metadata.name,BLOCK-MOVE:.metadata.annotations.clusterctl\.cluster\.x-k8s\.io/block-move' ``` Repeat until every row’s `BLOCK-MOVE` column is ``. The annotation is set before a Job is created and removed once no Job is active, so a cleared annotation means the object has no Job in flight. 3. **Unpause the Cluster** — required, or step 4 fails immediately: ```sh kubectl patch cluster -n --type=merge -p '{"spec":{"paused":null}}' ``` 4. **Run `clusterctl move`** to the target management cluster: ```sh clusterctl move --namespace --to-kubeconfig ``` `clusterctl move`’s own backoff does not wait for long Jobs (about two minutes, while `jobs.activeDeadlineSeconds` defaults to one hour), which is why step 2 waits explicitly before the move itself starts. ## Moving every namespace `clusterctl move` always moves one namespace: an unset `--namespace` falls back to your kubeconfig context’s current namespace, not to every namespace on the cluster, and there is no flag that moves all of them in one invocation. Moving every namespace means repeating the procedure above once per namespace. A `TerraformClusterIdentity` is cluster-scoped and several namespaces can reference the same one. This is safe to move namespace by namespace: `clusterctl move` never deletes a cluster-scoped object from the source, only namespaced ones, so the first namespace’s move copies the identity to the target and leaves it in place on the source; each later namespace’s move finds it already on the target and skips re-creating it rather than erroring. You do not need to move the namespaces that share an identity in any particular order, and you still copy its credentials Secret to the target only once (see [below](<#the-identitys-credentials-secret-does-not-move>)), however many namespaces that use it you move. ## What moves and what doesn’t `clusterctl move` discovers CRD-kind objects, ConfigMaps and Secrets in a Cluster’s owner reference chain, copies them to the target, then strips finalizers and deletes them from the source. The controller’s own delete path therefore never runs on the source. | Moves | Doesn’t move | | --- | --- | | State Secrets (`tfstate-default-` and its chunks) — owned by the object | The state lock Lease (`lock-tfstate-default-`) — not a discovered kind. **Recreated by the backend at the target’s next `init`.** | | State backups (`captf-state-backup--` and their chunks) — owned by the object | The runner ServiceAccount (`captf-runner`) and its RoleBinding — not discovered kinds. **Recreated by the controller before the target’s first Job.** | | The durable inputs Secret (`captf-inputs--`) — by owner reference only | The run lease and, for a TerraformCluster, the cluster write lease — not discovered kinds. A lease whose Job was active during the move is left on the source; see [run leases](). | | The plan key (`captf-plankey--`) and the mirrored credentials Secret (`captf-creds-`) — by owner reference only | The identity’s own credentials Secret — deliberately not owned and not labeled for move. **Copy it to the target yourself; see below.** | | | Every `variablesFrom` ConfigMap or Secret — deliberately not owned. **Recreate it on the target, or label it for move yourself** (see [Module Variables]()). | Only the state Secrets and the state backups carry the `clusterctl.cluster.x-k8s.io/move` label (the backend labels include it). The durable inputs Secret, the plan key and the credential mirror carry no move label: they move because the owner-reference chain `clusterctl` follows leads to them. The identity’s source Secret has neither label nor owner reference and does not move. The per-run Secret (`captf-run-`) is owned by its Job, is deleted when the Job finishes and is not meant to move; a move waits for running Jobs through `block-move`. See [Annotations and labels]() for the label and every other one this page’s objects carry. `status` is never restored by move: it is rebuilt from state on the target’s first reconcile. A `TerraformCluster` or `TerraformMachinePool` (both mutable) re-reads its `variablesFrom` sources on every reconcile, so it waits at `DependenciesReady=False`/`VariablesSourceNotFound` on the target until the source exists there, whether or not it was already provisioned. A `TerraformMachine` (immutable) stops needing its sources once provisioned, so only one moved before its first apply waits. ### The identity’s credentials Secret does not move The `TerraformClusterIdentity` object itself is cluster-scoped and carries a move-hierarchy label, so `clusterctl` moves it as part of the global hierarchy every namespace’s move includes. Its credentials Secret is different: it is namespace-scoped and used by every namespace the identity allows, so CAPTF deliberately removes any owner reference it might have added and never labels it for `clusterctl move`. Labeling it for move yourself would make `clusterctl` delete it from the source once the move of the namespace it lives in finished — deleting credentials that other allowed namespaces on the source still use, and `clusterctl move` never moves more than one namespace per invocation regardless (see [Moving every namespace](<#moving-every-namespace>) above), so there is no invocation that “finishes last” to safely hang the deletion off. Copy it yourself every time, on every namespace’s move; it is a duplicate, not a move. Because the identity object moves but its Secret does not, the identity reports `Ready=False`/`SecretNotFound` on the target until you copy the Secret there yourself, keeping its name and the keys the identity names; every `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` that uses it reports `IdentityAllowed=False`/`SecretNotFound` for the same reason. `kubectl get -o yaml | kubectl apply -f -` carries the source Secret’s `resourceVersion` and `uid` along, which `apply` rejects against a target that has no existing object with that `resourceVersion`; strip the identifying metadata first: ```sh kubectl --kubeconfig get secret -n -o json \ | jq 'del(.metadata.resourceVersion, .metadata.uid, .metadata.creationTimestamp, .metadata.managedFields, .metadata.ownerReferences)' \ | kubectl --kubeconfig apply -f - ``` See [Identities and Credentials]() for creating, rotating and revoking that Secret. ### TerraformMachine deletion during a move `clusterctl` annotates each object with `clusterctl.cluster.x-k8s.io/delete-for-move` before deleting it on the source. The `TerraformMachine` delete webhook honors that annotation only while the machine’s Cluster (its `cluster.x-k8s.io/cluster-name` label) has `spec.paused: true`, which `clusterctl move` sets before it deletes anything. On an unpaused Cluster the annotation changes nothing: deleting a `TerraformMachine` whose Machine is live is refused, because it would skip drain and the lifecycle hooks. This check exists only on `TerraformMachine`: a `TerraformMachinePool` is never delete-guarded, moved or not. ## Confirm it worked > [!TIP] > > ```sh > kubectl --kubeconfig get terraformclusters,terraformmachines,terraformmachinepools -n > kubectl --kubeconfig get cluster -n -o jsonpath='{.spec.paused}' > ``` > > The first command lists the `Terraform*` objects on the target with the same names they had on the source; the second prints nothing (or `false`), confirming `clusterctl move` unpaused the Cluster once it finished. Expect `Ready=False` until you complete the [manual steps above](<#the-identitys-credentials-secret-does-not-move>) and any `variablesFrom` source the objects need, then a Job starts on the target and the object reaches the same `Ready` state it had on the source. See the [other runbooks]() for any condition that does not clear on its own. ## Orphaned objects left on the source After a successful move, the source namespace still holds the runner ServiceAccount and RoleBinding (`captf-runner`), and any run or cluster write lease whose Job was active during the move or had not been bookkept yet (`captf-run-`, `captf-cluster-`) — all labeled `captf.io/managed=true`. Jobs and their per-run Secrets are not left behind: they have owner references to the deleted source objects, so the garbage collector removes them. The manager sweeps these automatically: on start, and then every `--sync-period` (default 10 minutes), it removes the `captf.io/managed=true` ServiceAccounts, RoleBindings and Leases of every namespace that holds no `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` — exactly what a completed move leaves behind. See [RBAC]() for the full mechanism, including how it treats a namespace an administrator still manages by hand. > [!WARNING] > > **Run it only once the namespace holds no Terraform\* object** > > Run it only once the namespace holds no `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`: while any remain, their runner and lock are still in use. To clean up without waiting for the next sync, once the namespace holds no `Terraform*` object: ```sh kubectl delete lease,serviceaccount,rolebinding -l captf.io/managed=true -n ``` > [!NOTE] > > **See also** > > - [The Reconcile Lifecycle]() — the `block-move` annotation and run leases in full. > - [Identities and Credentials]() — creating and rotating the credentials Secret this page tells you to copy. > - [RBAC]() — the runner ServiceAccount and RoleBinding the sweep manages. > - [Module Variables]() — `variablesFrom` sources and why they don’t move either. # clusterctl move `clusterctl move` is a delete that skips the controller’s delete path. It copies a Cluster’s objects to another management cluster, then deletes them from the source with their finalizers stripped. No destroy runs and no cleanup runs on the source; the infrastructure is not touched, because the target takes it over. The procedure is the [clusterctl move runbook](). This page explains how the pieces behave. ## How the source is deleted `clusterctl` pauses the Cluster, annotates each object `clusterctl.cluster.x-k8s.io/delete-for-move`, and deletes it. - The `TerraformMachine` delete webhook honors the annotation only while the owning Cluster, found through the machine’s `cluster.x-k8s.io/cluster-name` label, has `spec.paused=true`. On an unpaused, missing or unlabeled Cluster the ordinary check applies and the delete is refused while a live Machine references the object. See [Order and finalizers](). - A paused object runs only the paused branch, so the controller does not start a destroy for the objects being deleted. - Before any of this, `clusterctl` waits for the `block-move` annotation to clear. The controller sets it before it creates a Job, persists it before the Job exists, and clears it once no Job is active, also while paused. A Job the Job cache has not shown yet keeps the annotation: the controller checks the API server and the live run lease before it clears. See [The clusterctl move block](). ## What moves `clusterctl` follows a Cluster’s owner-reference chain and any object carrying the move label. | Object | Moves by | Notes | | --- | --- | --- | | The `Terraform*` objects | Owner chain | Spec and metadata only | | State Secrets, every chunk | `clusterctl.cluster.x-k8s.io/move` label, and the owner reference once the object owns it | The label is set by the backend config on every state Secret | | State backups | The same label and owner reference | | | The durable inputs Secret | Owner reference | No move label; it carries the `captf.io/applied` marker and the pinned digest | | The plan key Secret | Owner reference | | | The credential mirror | Owner reference | Recreated if missing | | The `TerraformClusterIdentity` | Its CRD’s `move-hierarchy` label | Cluster-scoped; the first namespace’s move copies it, later ones skip it | ## What does not move - **Status.** The target rebuilds it from the state on its first reconcile: outputs, health, `provisioned`. What status carried that nothing else does is the “ever applied” record, which is why the `captf.io/applied` marker lives on a Secret that does move. See [ever applied](). - **The Leases.** The state lock Lease is not a kind `clusterctl` moves; the backend recreates it at the target’s next `init`. The run and cluster write leases are taken afresh. A lease whose Job was active during the move stays on the source. - **The runner ServiceAccount and RoleBinding.** The target controller creates them before its first Job. - **The credential source Secret.** The Secret a `TerraformClusterIdentity` names is deliberately neither owned nor labeled for move: labeling it would make `clusterctl` delete it from the source, where other namespaces may still use it. **You must recreate it on the target.** Until then the identity reports `Ready=False`/`SecretNotFound` and every object using it `IdentityAllowed=False`/`SecretNotFound`. See the [runbook step]() for the command. - **`variablesFrom` sources.** Not owned, so not moved. A `TerraformCluster` or `TerraformMachinePool` waits at `DependenciesReady=False`/`VariablesSourceNotFound` on the target until the source exists there. The `captf.io/managed` Leases, ServiceAccounts and RoleBindings left on the source are removed by the [namespace RBAC sweep]() once the namespace holds no `Terraform*` object. ## After the move On the target, the first reconcile reads the moved state. A state that did not arrive, while the durable Secret’s marker or a backup says the object applied, reads as `StateLost`, not as a new object, so it is never applied a second time next to the live resources; restore a backup or investigate. Periodic checks have no history to count from, so the drift and health schedules start from the object’s creation time with a per-object jitter rather than all at once. Deleting the source namespace after a move has a limit; see [the known limit](). > [!NOTE] > > **See also** > > - [clusterctl move runbook](). > - [Lifecycle walkthroughs: `clusterctl move`](). # Troubleshooting # Troubleshooting by Condition Every CAPTF object tells you what it is waiting for in its status conditions. This chapter is the lookup: given a condition and a reason, what it means, what usually causes it and what to do. It is organized as one page per kind of lookup, not per symptom, so use it when you already see a reason. When you only have a symptom, start from the [runbooks](). ## Choose a starting point - **Runbooks by Symptom** --- You only have a symptom or an alert. Each runbook goes from diagnosis to fix. - **Nothing Is Happening** --- An object has no Job and is waiting. - **Every Condition** --- Every condition type and reason, with status, meaning, likely cause, action and a link. - **Every Event** --- The `Warning` events, with what to do, and the informational ones. - **Jobs** --- Jobs that keep failing. - **State** --- State that is lost, corrupt or locked. - **Deletion** --- A deletion that does not finish. - **Access** --- Identities, credentials and RBAC that block an object. This page covers where to start and how Ready is built. ## Start here Work from the top level down. Each step narrows the search. 1. **Read `Ready`.** `kubectl get -n ` shows it. If it is `True` and nothing seems wrong, you are done. If it is `False` or `Unknown`, go on. 2. **Read its inputs.** `Ready` summarizes other conditions, and its message names the ones that decide it. Which ones count depends on the kind and on whether the object has been provisioned (see [below](<#what-feeds-ready>)). 3. **Find the `False` or `Unknown` condition.** List them: ```sh kubectl get -n -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\t"}{.message}{"\n"}{end}' ``` Three conditions have **negative polarity**, where `True` is the problem: `Deleting`, `DriftDetected` and `DeletionBlocked`. Every other condition is healthy when `True`. 4. **Read its reason and message.** The reason is what the table keys on. The message adds the specifics: a Job name, a key, an annotation to set. 5. **Look it up** in the [conditions table](), under the condition’s type. The row gives the cause, the action and the page that goes deeper. 6. **Check the events.** `kubectl events -n --for /` shows what the controller did and when. The [events page]() explains each `Warning`. ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A[Ready] --> B{Status?} B -- True --> C[Healthy: check DriftDetected
and ApplyJobSucceeded if unsure] B -- False or Unknown --> D[Read the Ready message
and list the conditions] D --> E{Which input is not True?} E -- DependenciesReady --> F[An owner, the cluster or bootstrap data] E -- IdentityAllowed, CredentialsMirrored, RunnerRBACReady --> G[Credentials and RBAC] E -- ApplyJobSucceeded --> H[A Job, a lease, an approval or a gate] E -- StateReadable --> I[The state] E -- OutputsValid, InfrastructureHealthy --> J[The module or the instance] F --> K[Conditions table] G --> K H --> K I --> K J --> K ``` ## What feeds Ready | Phase | `TerraformCluster` and `TerraformMachine` | `TerraformMachinePool` | | --- | --- | --- | | Before provisioned | `DependenciesReady`, `IdentityAllowed`, `CredentialsMirrored`, `RunnerRBACReady`, `ApplyJobSucceeded`, `StateReadable`, `OutputsValid`, `InfrastructureHealthy`, `Deleting` | The same nine | | After provisioned | `InfrastructureHealthy`, `Deleting` | `InfrastructureHealthy`, `ApplyJobSucceeded`, `Deleting` | Two consequences follow: - **After provisioning, a failing apply does not turn a cluster’s or machine’s `Ready` false.** A failed re-apply of a cluster must not flip the Cluster’s `InfrastructureReady`, which would suspend every MachineHealthCheck. Check `ApplyJobSucceeded`, `StateReadable` and `DriftDetected` yourself when the object looks healthy but is not changing. - **These conditions never feed `Ready`**, so check them by symptom: `Paused`, `RestoreJobSucceeded`, `DriftJobSucceeded`, `DriftDetected`, `DeletionBlocked`, `EndpointAvailable` and `AutoscalingActive`. `CapacityResolved` belongs to the template kinds and `Ready` to the identity, which sets it from its Secret. ## By symptom | You see | Look at | | --- | --- | | A new object that never provisions | `DependenciesReady`, then the credential conditions, then `ApplyJobSucceeded` | | Nothing is happening and there is no Job | `ApplyJobSucceeded` wait reasons; [Nothing is happening]() | | Jobs keep failing | `ApplyJobSucceeded`, `DriftJobSucceeded`; [Failing Jobs]() | | A deletion that does not finish | `Deleting`, `DeletionBlocked`, `StateReadable`; [My object will not delete]() | | The state is lost, corrupt or locked | `StateReadable`; [Unreadable State]() | | A change waits for a person | `ApplyJobSucceeded` with `PlanAwaitingApproval` or `DestructivePlanBlocked`; [Approvals and Gates]() | | An instance is unhealthy | `InfrastructureHealthy`; [Machine Remediation]() | | The module’s outputs are rejected | `OutputsValid`; [Module Contract]() | > [!NOTE] > > **See also** > > - [Conditions reference]() and [Events reference](): the generated lists these pages build on. > - [Observability]() for metrics and alerts. > - [Runbooks](). # Runbooks If you already have a condition and a reason, [Troubleshooting by Condition]() looks it up directly. Each runbook below covers one symptom: what it looks like, why it happens, how to check, how to fix it, and how to confirm the fix worked. Alerts and condition reasons are cross-references, not a substitute for reading the page: start from whichever alert fired or condition you see, but read the whole runbook before you act. | Runbook | Covers | Alerts and conditions | | --- | --- | --- | | [Failing Jobs]() | A `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` whose Jobs keep failing. | [`CAPTFJobFailing`](), [`CAPTFNoRecentSuccess`](); `ApplyJobSucceeded=False` and `DriftJobSucceeded=False` | | [Stuck Destroy]() | A `destroy` Job that cannot succeed, and how to remove the object’s finalizer safely. | [`CAPTFDestroyStuck`](); `ApplyJobSucceeded=False`/`DestroyFailed` | | [Unreadable State]() | A state Secret CAPTF cannot parse. | [`CAPTFStateUnreadable`](); `StateReadable=False`/`StateCorrupt`, `StateInconsistent` or `StateEncrypted` | | [State Restore]() | Restoring a Terraform or OpenTofu state from a CAPTF-managed backup. | `StateReadable=False`/`StateLost`; `RestoreJobSucceeded` | | [Stale State Lock]() | Clearing a state lock left behind by a killed or evicted runner. | [`CAPTFForceUnlocks`](); `StateReadable=False`/`StateLocked` | | [Size Limits]() | A state or rendered inputs approaching the Secret size limit. | [`CAPTFStateNearSecretLimit`](), [`CAPTFInputsNearLimit`](); `ApplyJobSucceeded=False`/`InputsTooLarge` | | [Slow Jobs]() | Jobs that take a long time to run, or a long time to start. | [`CAPTFJobSlow`](), [`CAPTFJobQueueSlow`]() | | [Reconcile Errors]() | The controller itself failing to reconcile. | [`CAPTFReconcileErrors`]() | | [Identities and Credentials]() | An object that cannot resolve or mirror its `TerraformClusterIdentity`. | `IdentityAllowed=False`, `CredentialsMirrored=False` | | [Webhook Unavailable]() | Writes to a `Terraform*` object failing because the admission webhook cannot be reached. | None | | [Total State Loss and Import]() | The state is gone with no backup: rebuild it from a workstation, or abandon and recreate the object with `import` blocks. | `StateReadable=False`/`StateLost` | | [clusterctl move]() | Moving a Cluster’s `Terraform*` objects with `clusterctl move`: the procedure, what does and does not come along, and cleaning up what a move leaves behind. | None | For how deletion works and a flowchart from a stuck object to the right runbook, see [Deletion and Teardown](). For why an object has no Job and is waiting, see [Nothing Is Happening](). Every runbook above applies to `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` alike unless it says otherwise. > [!NOTE] > > **See also** > > - [Observability]() — the metrics and alerts these runbooks are reached from. > - [Conditions reference]() — every condition type and reason named above. # Nothing Is Happening An object that is not changing and has no Job is waiting on something. The controller always records what: a condition, an event, or a requeue reason in its log. Find it in this order. ```sh # The conditions that carry a wait, and Ready's own message kubectl get -n \ -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\t"}{.message}{"\n"}{end}' # The Jobs of the object, newest last kubectl get jobs -n -l captf.infrastructure.cluster.x-k8s.io/owner-name= --sort-by=.metadata.creationTimestamp # The events kubectl events -n --for / ``` Then match what you see. ## By wait reason | You see | Meaning | What to do | | --- | --- | --- | | `ApplyJobSucceeded=Unknown`/`WaitingForRunLease`, message names a Job | A live Job, or another manager, holds the object’s run lease or the Cluster’s write lease | Wait; it clears when the holder finishes. If the holder Job does not exist, the lease frees after a minute. Never delete a Lease by hand. See [Leases]() | | `…/WaitingForClusterOperation` on a machine or pool | Its `TerraformCluster` applies, destroys or restores | Wait. The cluster Job is named in the message; look at that Job if it is slow | | `…/WaitingForMachineOperations` on a cluster | Machine and pool applies, destroys or restores of the Cluster are in flight; new ones wait behind the cluster | Wait; the message names up to five of them | | The same reasons on `DriftJobSucceeded` or `RestoreJobSucceeded` | The op that waits is a refresh, a drift check or a restore | The same | | `ApplyJobSucceeded=Unknown`/`PlanAwaitingApproval` | Manual approval of a plan | [Plan Approval]() | | `ApplyJobSucceeded=False`/`DestructivePlanBlocked` | The destructive-plan guard stopped an apply | [Destructive-plan guard]() | | `ApplyJobSucceeded=False`/`JobPolicyInvalid` | The merged `lockTimeoutSeconds` is not below `activeDeadlineSeconds` | Fix one of them; see [Deadlines]() | | `ApplyJobSucceeded=False`/`ApplyFailed`, message `Job : disappeared while it ran…` | An apply Job was deleted while it ran; an apply of the current inputs is due and guarded | Wait for it, or approve its plan if it waits. See [Retries]() | | `ApplyJobSucceeded=False`/`ApplyFailed`, `JobDeadlineExceeded`, `ImagePullFailed`, `ImageInvalid`, `InputsTooLarge` | The last Job failed and the op is in backoff, or nothing can start until the cause is fixed | [Failing Jobs](), [Size Limits](). The delay is in [Retries]() | | `StateReadable=False`/`StateLost`, `StateCorrupt`, `StateEncrypted`, `StateInconsistent` | No Job runs until the state reads | [Unreadable State](), [State Restore]() | | `StateReadable=False`/`StateLocked` | A foreign holder has the state lock; every Job waits `lockTimeoutSeconds` and fails | [Stale State Lock](); see [The state lock]() | | `DependenciesReady=False` or `Unknown` | An owner, bootstrap data, cluster exports or a variables source is not there yet | [Reconcile Errors]() | | `IdentityAllowed`, `CredentialsMirrored` or `RunnerRBACReady` not `True` | The credentials for the Job are not ready; retried every 30 seconds | [Identities and Credentials]() | | `Paused=True` | The object or its Cluster is paused: bookkeeping only, no Job starts | Unpause | | `RestoreJobSucceeded=False`/`RestoreBackupNotFound` | The restore annotation names no complete backup | [State Restore]() | | `DeletionBlocked=True` | A cluster waits for its machines and pools | [Deletion order]() | ## No condition explains it | Situation | What it is | What to do | | --- | --- | --- | | A Job exists and runs for a long time | The Job is the work; the object waits on it | Look at its pod logs; see [Slow Jobs](). It ends at `activeDeadlineSeconds` | | A Job exists, its pod is `Pending` | The image, the volume or the schedule | A Job older than a minute whose per-run Secret is missing and whose pods are all `Pending` is deleted and started again by the controller | | Idle, healthy, `Ready=True` and no Job | Nothing is due: the next drift, health or membership deadline | Wait, or see [Schedules]() for when the next check is due | | The object looks right but is stale | The Job cache lags, or the manager is not the leader | [Cache lag](); check the manager pods and [Leader election]() | | No condition changes and no events | The manager is not reconciling this object | [Reconcile Errors](); check `--watch-filter` and `--namespace` | | Writes to the object fail | The webhook is unavailable | [Webhook Unavailable]() | ## What not to do > [!CAUTION] > > **Do not delete a Lease or a failed Job to unstick a wait** > > - **Do not delete a Lease** to unstick a wait. Leases are taken over by the controller once the holder finishes, does not exist after a minute, or is past its deadline plus the backstop. > - **Do not delete a failed Job** to skip a backoff. It is the record the retry and the remediation cap count from. > - **Do not edit the object’s annotations** to clear `block-move`; the controller clears it when no Job is active. > [!NOTE] > > **See also** > > - [Runbooks](). > - [Conditions reference](). # Conditions Every condition CAPTF sets, with what each reason means and what to do about it. Find the object’s condition type below, then its status and reason. The tables list **every** status and reason a type can take, so a reason that is healthy is here too. [Start here]() explains how to get from `Ready` to the right row. Jump to: [Ready](<#ready>) [Paused](<#paused>) [Deleting](<#deleting>) [DependenciesReady](<#dependenciesready>) [IdentityAllowed](<#identityallowed>) [CredentialsMirrored](<#credentialsmirrored>) [RunnerRBACReady](<#runnerrbacready>) [ApplyJobSucceeded](<#applyjobsucceeded>) [StateReadable](<#statereadable>) [RestoreJobSucceeded](<#restorejobsucceeded>) [OutputsValid](<#outputsvalid>) [InfrastructureHealthy](<#infrastructurehealthy>) [DriftJobSucceeded](<#driftjobsucceeded>) [DriftDetected](<#driftdetected>) [DeletionBlocked](<#deletionblocked>) [EndpointAvailable](<#endpointavailable>) [AutoscalingActive](<#autoscalingactive>) [CapacityResolved](<#capacityresolved>). How to read the tables: - **Status** is the condition’s status: `True`, `False` or `Unknown`. For the three negative-polarity types (`Deleting`, `DriftDetected` and `DeletionBlocked`) `True` is the problem. - **Meaning** is what the reason says about the object. - **Likely cause** is what usually puts the object there. A transient wait is said to be transient. - **What to do** is the next action. “Wait” means the controller will move on by itself, and a link leads to how long that takes. - The generated [Conditions reference]() lists the same reasons with the code’s own wording. ## Ready Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool, TerraformClusterIdentity. Polarity: normal (`True` is healthy). The only condition Cluster API reads. It summarizes other conditions; the message names the inputs that decide it. A `TerraformClusterIdentity` sets it from its credentials Secret instead. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `Ready` | Every input of Ready is `True`. | A healthy object. | Nothing. | [Ready summarization]() | | `True` | `SecretFound` | Identity only: the credentials Secret named by `spec.secretRef` exists. | A healthy identity. | Nothing. | [Identities]() | | `False` | `NotReady` | At least one input of Ready is `False`. The message names it. | Any failing input: dependencies, credentials, the apply, the state, outputs or health. | Read the message, then find the named condition below. Before the object is provisioned the inputs are `DependenciesReady`, `IdentityAllowed`, `CredentialsMirrored`, `RunnerRBACReady`, `ApplyJobSucceeded`, `StateReadable`, `OutputsValid`, `InfrastructureHealthy` and `Deleting`. | [Start here]() | | `False` | `SecretNotFound` | Identity only: the credentials Secret does not exist. | Never created, `spec.secretRef` names the wrong Secret or namespace, or the Secret was not copied after a `clusterctl move`. | Create or copy the Secret. | [Identity runbook]() | | `Unknown` | `ReadyUnknown` | No input is `False`, but at least one is `Unknown`. | The object is waiting: for its owner, the cluster infrastructure, bootstrap data, a first apply or a lease. | Find the `Unknown` input in the status. This is usually transient. | [Start here]() | ## Paused Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). Whether reconciliation is paused. A paused object runs bookkeeping only and starts no Job. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `Paused` | The object, or its Cluster, is paused. | `spec.paused` on the Cluster, or the `cluster.x-k8s.io/paused` annotation on the object. `clusterctl move` pauses on purpose. | Unpause when the pause is no longer wanted. A deletion under a paused object waits, and `Deleting` says so. | [Order and finalizers]() | | `False` | `NotPaused` | The object is not paused. | The normal state. | Nothing. | [Lifecycle]() | ## Deleting Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: negative (`True` is the problem). Whether the object is being deleted. A deleting object’s message can say what the destroy waits for. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `Deleting` | The object has a deletion timestamp. | A delete, usually through Cluster API. The message may read `The destroy Job waits for its credentials`, or `Deletion waits until the object is unpaused` when the object or its Cluster is paused: a paused object keeps its finalizer, starts no Job and never runs a destroy, so that `clusterctl move` can delete source objects safely. | Wait for the destroy. If the message says unpaused, unpause. Otherwise follow the flowchart. | [My object will not delete]() | | `False` | `NotDeleting` | The object has no deletion timestamp. | The normal state. | Nothing. | [Deletion and Teardown]() | ## DependenciesReady Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). The gates before a Job can render inputs: the owner, the Cluster, the cluster infrastructure and exports, bootstrap data and variables. While a gate is closed nothing is rendered or run. While the object is deleting, owner gates do not apply: the destroy still runs. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `DependenciesReady` | Every gate is open. | The normal state. | Nothing. | [Lifecycle]() | | `False` | `ClusterNotTerraform` | The owning Cluster’s `infrastructureRef` is not a `TerraformCluster`. | A machine or pool was created for a Cluster that uses another infrastructure provider. | Point the Cluster at a `TerraformCluster`, or move the machine to the right Cluster. | [Kinds]() | | `False` | `OwnerMismatch` | An owner reference of the expected kind resolves to an object that does not reference this one back. The reference is treated as forged or stale. | A wrong or missing `infrastructureRef`, a UID mismatch, or a `cluster-name` label that disagrees with the owner. | Fix the owner’s `infrastructureRef` or the label. The message says which check failed. | [Kinds]() | | `False` | `OwnerNotFound` | The owner Machine, MachinePool or Cluster is gone. | The owner was deleted before this object, or a stale owner reference remains. | Delete this object if it is orphaned. A deleting object still destroys from its durable inputs. | [Deletion order]() | | `False` | `VariablesInvalid` | A `variablesFrom` source has a key that is not a Terraform identifier or is reserved, or a value that is not UTF-8 or valid JSON. The message names the key, never the value. | A malformed ConfigMap or Secret. | Correct the source. No Job starts until you do. | [Module variables]() | | `False` | `VariablesSourceNotFound` | A required `variablesFrom` ConfigMap or Secret is missing or lacks the `captf.io/variables=true` label. | Not created, wrong name, a missing label, or not recreated after a `clusterctl move`. | Create the source and label it, or mark the reference optional. | [Module variables]() | | `False` | `WaitingForOwnerMachine` | A fresh `TerraformMachine` has only a control-plane owner reference so far. | The control-plane provider created it before Cluster API set the Machine owner reference. | Wait. If it persists, check the Machine’s `infrastructureRef` and the object’s owner references. | [Control planes]() | | `False` | `WaitingForOwnerMachinePool` | A `TerraformMachinePool` has owner references but none to its MachinePool yet. | Cluster API has not set the MachinePool owner reference. | Wait. If it persists, check the MachinePool’s `infrastructureRef`. | [Machine pools]() | | `Unknown` | `WaitingForOwner` | The owner is not set yet. For a machine or pool: the Cluster named by the `cluster.x-k8s.io/cluster-name` label is not found. | The object was just created, or the label is missing or wrong. | Wait. Check the label and the Cluster’s name. | [Lifecycle]() | | `Unknown` | `WaitingForClusterInfrastructure` | The `TerraformCluster` is not provisioned yet. | The cluster’s first apply has not finished or failed. | Read the `TerraformCluster`’s `ApplyJobSucceeded` and `Ready`. | [Failing Jobs]() | | `Unknown` | `WaitingForClusterExports` | The cluster module’s `exports` output is not readable. | The cluster’s state is not readable, or its module declares no exports output. | Check the `TerraformCluster`’s `StateReadable` and `OutputsValid`. | [Unreadable state]() | | `Unknown` | `WaitingForBootstrapData` | The Machine’s or MachinePool’s bootstrap data Secret is not set or not found. | The bootstrap provider has not produced it yet, or the control plane is waiting. | Check the bootstrap config object and the control plane. Nothing in CAPTF creates this Secret. | [Control planes]() | ## IdentityAllowed Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). Whether the object’s `TerraformClusterIdentity` exists, allows the namespace and has its credentials Secret. While it is not `True` no Job starts, except that a deletion that needs no Job still finishes. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `IdentityAllowed` | The identity exists, allows the namespace and its Secret exists. | The normal state. | Nothing. | [Identities]() | | `False` | `IdentityNotFound` | No identity resolved. | No `identityRef` on the object or the cluster defaults, or the named identity does not exist. | Set `identityRef.name` to an existing identity, or create it. | [Identity runbook]() | | `False` | `NamespaceNotAllowed` | The identity’s `allowedNamespaces` excludes this namespace. The mirror is revoked. | The namespace was never added, or was removed. | Add the namespace to the identity. | [Identity runbook]() | | `False` | `SecretNotFound` | The identity allows the namespace but its credentials Secret does not exist. | Not created, deleted, or not copied after a `clusterctl move`. | Create the Secret at the name and namespace in `spec.secretRef`. | [Identity runbook]() | | `Unknown` | `IdentityCheckFailed` | The check could not be completed. | A read of the identity, the namespace labels or the Secret failed for a reason other than not found, such as an API error. | Usually clears on its own. If it persists, read the manager’s log for the error. | [Identity runbook]() | ## CredentialsMirrored Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). Whether the identity’s credentials were copied into the object’s namespace as the mirror Secret the Job mounts. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `Mirrored` | The mirror exists and is current. | The normal state. | Nothing. | [Identities]() | | `False` | `MirrorFailed` | Creating or updating the mirror failed. | An API error, or a Secret with the mirror’s name that is not a mirror of this identity. | Remove or rename the conflicting Secret. Another message names an API error: fix it and retry. | [Identity runbook]() | | `Unknown` | `MirrorPending` | No mirror yet. | `IdentityAllowed` is not `True`, or this is before the first mirror. | Fix `IdentityAllowed`; this follows it. | [Identity runbook]() | ## RunnerRBACReady Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). Whether the runner ServiceAccount exists and is bound to the runner role in the namespace. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `RBACReady` | The runner ServiceAccount and RoleBinding are in place. | The normal state. | Nothing. | [RBAC]() | | `False` | `RBACFailed` | Creating the ServiceAccount or the `captf-runner` RoleBinding failed. | An API error, or a RoleBinding named `captf-runner` that CAPTF does not own. | Remove or rename the conflicting RoleBinding. Otherwise fix the error in the message. | [Identity runbook]() | | `False` | `ServiceAccountNotOptedIn` | An override ServiceAccount lacks the `captf.io/runner=true` label. | `spec.jobs.serviceAccountName` names a ServiceAccount that does not exist or is not labeled. | Label it, create it, or drop the override. The controller retries every 30 seconds. | [Identity runbook]() | ## ApplyJobSucceeded Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). The outcome of the newest apply, plan or destroy, and the reason an apply or destroy waits or cannot start. A plan Job stands for the apply it plans. The `Unknown` wait reasons here are also used by `DriftJobSucceeded` and `RestoreJobSucceeded`. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `ApplySucceeded` | The newest apply succeeded. | The normal state. | Nothing. | [Jobs]() | | `True` | `DestroySucceeded` | The destroy succeeded; cleanup follows. | A deletion that completed its destroy. | Nothing. | [Cleanup]() | | `False` | `ApplyFailed` | An apply Job failed. The message names the Job and the failed step. A message `Job : disappeared while it ran and may have applied part of its change; an apply of the current inputs is due` means the Job was deleted while it ran (cluster or pool): an apply of the current inputs stays due and is guarded, with a suffix saying whether it is guarded. This message takes precedence over an older blocked or plan-changed apply, and stays until an apply started after the vanished Job (carrying `captf.io/after-interrupted-apply`) reports its own outcome. | A module error, a provider or cloud error, a lock timeout, or a runner error. For the `disappeared` message, a `kubectl delete job` of a running apply. | Read `status.lastRun` and the Job’s pod logs. The controller retries with backoff. For `disappeared`, let the due apply run, or approve its plan if it waits; it clears when an apply started afterwards succeeds. | [Failing Jobs]() | | `False` | `DestroyFailed` | A destroy Job failed, or a destroy cannot be rendered because the durable inputs Secret is missing. | A provider or cloud error; a dependency still in use; a deleted inputs Secret. | Fix the cause. It retries forever. If it cannot succeed, abandon the object or clean up by hand. | [Stuck destroy]() | | `False` | `JobDeadlineExceeded` | The Job hit `activeDeadlineSeconds`. It counts toward backoff. | A slow module or provider, or a hung call. | Raise `spec.jobs.activeDeadlineSeconds`, or find what hangs. | [Deadlines]() | | `False` | `ImagePullFailed` | The pod stayed in `ErrImagePull` or `ImagePullBackOff` past the deadline. | A wrong image or tag, a missing pull secret, or a registry outage. | Fix the image reference or `imagePullSecrets`. | [Failing Jobs]() | | `False` | `ImageInvalid` | The runner reported an image-layout error: no `/captf/module` or a non-executable command. | The image does not follow the image contract. | Rebuild the image to the contract; check it with `tfcapi-lint`. | [Image contract]() | | `False` | `InputsTooLarge` | The rendered root module and variables exceed what a Secret can carry. No Job starts. | Large variables, bootstrap data or exports. | Shrink the inputs. | [Size limits]() | | `False` | `JobPolicyInvalid` | The merged Job policy has a `lockTimeoutSeconds` that is not below `activeDeadlineSeconds`. No Job starts except a destroy. | A cluster default merged with a machine’s policy, or one of the two left at its default. | Lower the lock timeout or raise the deadline. | [Deadlines]() | | `False` | `DestructivePlanBlocked` | A `TerraformCluster` apply, or a `TerraformMachinePool` apply of a change of the cluster’s exports, stopped before a plan that deletes or replaces resources. It counts toward no backoff. For a pool the change is held: the pool keeps applying the exports of its last successful apply, and its `Ready` is `False` until the change is approved. | The guard: no approval names this hash (a cluster’s inputs hash, a pool’s approval hash: the inputs hash without `bootstrap_data`). For a pool, the exports changed and the plan deletes or replaces something; after a failed guarded apply every pool apply is guarded; and a pool that applied before exports were recorded, without proof of them, has every apply guarded (the message says the exports of the last successful apply are unknown). | Review the plan in the message and run the command it gives (`captf.io/approve-destructive-plan=`), or change the inputs. A pool’s approval is removed if the exports return to the applied ones. | [Destructive-plan guard]() | | `False` | `IdentityNotAllowed` | A pending destroy cannot start because the identity no longer allows the namespace, or is gone. | The identity’s `allowedNamespaces` changed, or the identity was deleted. | Allow the namespace again, or abandon the object. | [Held deletions]() | | `Unknown` | `NoApplyYet` | No apply has completed. | A new object, or one waiting on a gate or credentials. | Check `DependenciesReady`, `IdentityAllowed` and the Jobs. | [Jobs]() | | `Unknown` | `PlanAwaitingApproval` | Under `applyPolicy: Manual`, the plan in `status.plan` waits for approval. | A change of inputs under Manual policy. | Review the plan and annotate `captf.io/approve-plan` with its hash. | [Manual plan approval]() | | `Unknown` | `PlanChanged` | An approved apply planned other changes and stopped before applying them. | The inputs or the infrastructure changed after approval. | Review the new plan in `status.plan` and approve it. | [Manual plan approval]() | | `Unknown` | `WaitingForClusterOperation` | A machine’s or pool’s apply, destroy or restore waits for its `TerraformCluster`’s. | The cluster is applying or destroying. No Job exists yet and nothing counts as a failure. | Wait. The message names the cluster’s Job. | [Leases]() | | `Unknown` | `WaitingForMachineOperations` | A `TerraformCluster`’s apply, destroy or restore waits for its machines’ and pools’ operations in flight. | Machine or pool Jobs are running. New ones wait behind the cluster. | Wait. The message names up to five Jobs. | [Leases]() | | `Unknown` | `WaitingForRunLease` | Another live Job, perhaps another manager’s, holds the run lease or the Cluster’s write lease. | A Job of the same object or a second cluster operation is running. | Wait. Never delete a Lease by hand. | [Slow Jobs]() | ## StateReadable Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). Whether the object’s Terraform state can be read. While it is `False`, no Job runs, and a deleting object’s deletion is held, except for `StateLocked`. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `StateRead` | The state was read. | The normal state. | Nothing. | [Unreadable state]() | | `False` | `StateLost` | The state Secret of an object that applied before is missing, or an immutable object’s state carries no inputs hash. | A state Secret deleted by hand or with its namespace. For a provisioned machine: a state without an inputs hash. | Restore a backup with `captf.io/restore-state`. A deleting object can be abandoned instead. | [Unreadable state]() | | `False` | `StateCorrupt` | The state cannot be decoded, exceeds the reader’s caps, or has an unsupported version. | A hand edit, a partial write, or a state too large. | Restore a backup taken before the corruption. | [Unreadable state]() | | `False` | `StateEncrypted` | The state carries OpenTofu’s client-side encryption. CAPTF cannot read it. | State encryption is enabled in the module. | Disable it and push a decrypted state; no backup exists. | [Unreadable state]() | | `False` | `StateInconsistent` | The state chunks do not form one complete state. | A missing or duplicated chunk, or a stray Secret in the set. | Repair the set or restore a backup. | [Unreadable state]() | | `False` | `StateLocked` | Something other than the object’s own runner holds the state lock. Every Job waits `lockTimeoutSeconds` for it and fails. | A workstation running Terraform against the same backend. | Release it with `force-unlock` once the holder is gone. | [Stale lock]() | | `Unknown` | `StateNotFound` | No state exists yet. | The object has never applied. | Nothing; the first apply creates it. | [Unreadable state]() | ## RestoreJobSucceeded Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). The outcome of a state restore requested with `captf.io/restore-state`. Never an input of Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `StateRestored` | The restore Job pushed the backup into the backend. | A completed restore. | Nothing; the annotation is removed. | [State restore]() | | `False` | `RestoreBackupNotFound` | The annotation names no complete backup, or is not a serial. No Job starts. | A typo, a pruned backup or an incomplete one. | Pick a serial from `status.stateBackups`. | [State restore]() | | `False` | `RestoreFailed` | The restore Job failed. It is not retried for the same serial. | A lock, a bad backup or a runner error. | Read the Job’s logs, remove the annotation, wait for `status.lastRestoredSerial` to clear, and set it again. | [Other manual actions]() | | `Unknown` | `WaitingForRunLease` | Another live Job holds the run lease. | A Job of the object is running. | Wait. | [Leases]() | | `Unknown` | `WaitingForClusterOperation` | A machine’s or pool’s restore waits for its cluster’s operation. | The cluster is applying, destroying or restoring. | Wait. | [Leases]() | | `Unknown` | `WaitingForMachineOperations` | A cluster’s restore waits for its machines’ and pools’ operations. | Machine or pool Jobs are in flight. | Wait. | [Leases]() | ## OutputsValid Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). Whether the module’s outputs, read from state, match the role’s contract. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `OutputsValid` | The outputs are valid. | The normal state. | Nothing. | [Module contract]() | | `True` | `InstancesTruncated` | A pool’s `instances` output had more than 1000 entries; `status.instances` keeps the first 1000. The pool still provisions. | A very large pool. | Usually nothing. Count against the contract’s cap. | [MachinePool role]() | | `False` | `OutputsMissing` | A required output is not declared. | The module lacks an output its role requires. | Add the output. Check the module with `tfcapi-lint`. | [tfcapi-lint]() | | `False` | `OutputsInvalid` | An output breaks the contract or a Cluster API marker. | A wrong type, shape or value. The message names the output. | Fix the module’s output. | [Module contract]() | | `False` | `FailureDomainMismatch` | A machine’s `failure_domain` output differs from the one its Machine requested. | The module placed the instance elsewhere. | Honor the requested failure domain in the module. | [Machine role]() | | `False` | `ProviderIDChanged` | `provider_id` changed after it was first written. It is immutable. | The module replaced the instance outside Cluster API’s replacement, or the output is unstable. | Do not edit `spec.providerID`. Delete the Machine to replace the instance. | [Machine role]() | | `Unknown` | `OutputsPending` | Required outputs are null, or there is no state yet. | The first apply has not finished. | Wait for the apply. | [Jobs]() | ## InfrastructureHealthy Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). The health the module reports, read from state at each refresh. After provisioning it is the input that moves Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `Healthy` | The instance is running and healthy. | The normal state. | Nothing. | [Drift and health]() | | `False` | `Provisioning` | The first apply is running. | The first apply started and has not finished. | Wait for the Job. | [Jobs]() | | `False` | `InstancePending` | The module reports `pending`. | The instance is still starting. | Wait; refresh runs on a doubling schedule up to five minutes. | [Schedules]() | | `False` | `InstanceUnhealthy` | The module reports running but not healthy. | A failing health check on the instance. | Look at the instance. Remediation can replace it. | [Machine remediation]() | | `False` | `InstanceDegraded` | The module reports `degraded`. | A partial failure of the instance. | As above. | [Machine remediation]() | | `False` | `InstanceStopped` | The module reports `stopped`. | The instance was stopped outside CAPTF. | Start it, or let remediation replace it. | [Machine remediation]() | | `False` | `InstanceTerminated` | The module reports `terminated`, or the instance vanished. | The instance was deleted outside CAPTF. | Let remediation replace it, or delete the Machine. | [Machine remediation]() | | `Unknown` | `WaitingForProvisioning` | Before the first apply. | A new object. | Nothing; see `ApplyJobSucceeded`. | [Jobs]() | | `Unknown` | `HealthUnknown` | The module reports `unknown`, or no health at all. | The module has no health signal yet. | Check the module’s `health` output. | [Drift and health]() | | `Unknown` | `ProviderIDMissing` | `provider_id` turned null after provisioning. It is reported terminated only if the next sample is null too. | A transient read, or the instance vanished. | Wait one sample. | [Drift and health]() | ## DriftJobSucceeded Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: normal (`True` is healthy). The outcome of the newest refresh or drift Job. Never an input of Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `DriftChecked` | The newest refresh or drift Job succeeded. | The normal state. | Nothing. | [Drift]() | | `False` | `DriftJobFailed` | A refresh or drift Job failed. | A provider error, a lock timeout or a module error. | Read the Job’s logs. The check retries with backoff. | [Failing Jobs]() | | `False` | `DriftJobDeadlineExceeded` | The Job hit `activeDeadlineSeconds`. | A slow provider read. | Raise the deadline. | [Deadlines]() | | `Unknown` | `DriftNotChecked` | No drift check has run. | A new object, or drift is disabled. | Nothing, or enable drift checks. | [Drift]() | | `Unknown` | `DriftJobRunning` | A drift Job runs. | A drift check is in progress. | Wait. | [Drift and health]() | | `Unknown` | `WaitingForRunLease` | A refresh or drift waits for the run lease. | Another live Job of the object holds it. | Wait. | [Leases]() | ## DriftDetected Carried by: TerraformCluster, TerraformMachine, TerraformMachinePool. Polarity: negative (`True` is the problem). Whether the last drift check found differences. `True` is the problem, though it never feeds Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `DriftReported` | Drift was found and the action is `Report`. | Infrastructure changed outside CAPTF. | Decide whether to accept it, or set the action to `Remediate`. | [Drift]() | | `True` | `DriftPending` | Drift was found with action `Remediate`, and the remediation apply has not succeeded yet. The message names a failed, blocked or changed remediation Job. | The apply waits for backoff, a lease or an approval, or failed. | Wait, or read `ApplyJobSucceeded`. After the failed limit it waits for the next drift check. | [Retries]() | | `True` | `DriftRemediating` | A remediation apply runs. | Action `Remediate` and drift found. | Wait. | [Drift]() | | `False` | `NoDrift` | The last check found none. | The normal state. | Nothing. | [Drift]() | | `Unknown` | `DriftNotChecked` | No check has run. | A new object, or drift is disabled. | Nothing. | [Drift]() | ## DeletionBlocked Carried by: TerraformCluster. Polarity: negative (`True` is the problem). Whether a deleting `TerraformCluster` waits for dependents. Never an input of Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `DependentsExist` | Machines or pools of the cluster still exist. No destroy runs. | Cluster API has not finished deleting them, or one is stuck. | List them with `-l cluster.x-k8s.io/cluster-name=` and work on the stuck one. | [Order and finalizers]() | | `False` | `NotBlocked` | Nothing blocks the deletion. | The normal state. | Nothing. | [Deletion and Teardown]() | ## EndpointAvailable Carried by: TerraformCluster. Polarity: normal (`True` is healthy). Whether the cluster has a control-plane endpoint. Never an input of Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `EndpointAvailable` | A valid endpoint exists in the module output, the spec or the Cluster. | The normal state. | Nothing. | [Cluster role]() | | `False` | `WaitingForEndpoint` | The cluster is provisioned and neither the module’s `control_plane_endpoint` output nor `Cluster.spec.controlPlaneEndpoint` is set. | The module does not output an endpoint and none was supplied. | Output one from the module or set it on the Cluster. | [Cluster role]() | ## AutoscalingActive Carried by: TerraformMachinePool. Polarity: normal (`True` is healthy). Whether a pool’s replicas follow the autoscaler. Informational; never an input of Ready. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `ReplicasManagedByModule` | Both autoscaler annotations are valid and the controller writes the observed replicas back to the MachinePool. | An autoscaled pool. | Nothing. | [Machine pools]() | | `False` | `AutoscalingDisabled` | Neither annotation is set; `spec.replicas` is the only source of capacity. | A fixed-replica pool. | Nothing, unless you want autoscaling. | [Machine pools]() | | `False` | `AutoscalingAnnotationsInvalid` | An annotation is present but the pair is incomplete, unparsable or has min above max. The pool applies without autoscaling. | A typo or one annotation missing. | Fix the annotations; the message names the problem. | [Machine pools]() | | `False` | `ReplicasManagedExternally` | The annotations are valid but another controller owns `spec.replicas` through `cluster.x-k8s.io/replicas-managed-by`. Observed replicas are not written back. | Another autoscaler manages the MachinePool. | Choose one owner. A `Warning` event is emitted on the transition. | [Machine pools]() | ## CapacityResolved Carried by: TerraformMachineTemplate. Polarity: normal (`True` is healthy). Whether a template’s capacity and node info were read from its image’s labels. | Status | Reason | Meaning | Likely cause | What to do | See | | --- | --- | --- | --- | --- | --- | | `True` | `CapacityResolved` | Both labels parsed. | The normal state. | Nothing. | [Templates]() | | `True` | `CapacityNotDeclared` | The image carries neither label. | Not every image declares capacity. | Nothing, unless a ClusterClass autoscaler needs it. | [Templates]() | | `False` | `ImageInspectFailed` | The registry fetch or authentication failed. | A wrong reference, a registry outage or missing credentials. | Fix the reference or credentials. A `Warning` event is emitted. | [Templates]() | | `False` | `CapacityLabelInvalid` | A label is present but invalid. | A malformed capacity or node-info label on the image. | Fix the label in the image. | [Image contract]() | > [!NOTE] > > **See also** > > - [Events](). > - [Runbooks](). > - [Conditions reference](). # Events The controller records an event when something happens to an object, once per transition or occurrence rather than on every reconcile. Read them with: ```sh kubectl events -n --for / ``` Most `Warning` events repeat what a condition already says, at the moment it changed, and name the Job involved. Use them for the timeline: what happened first. The [Events reference]() is the generated list; this page adds what to do. > [!NOTE] > > **Events expire; conditions do not** > > Kubernetes keeps events for a limited time (one hour by default), so the condition is the record to rely on. Messages never carry credentials, variable values, output values or raw stderr, and are cut at 512 bytes. ## Warning events | Reason | Emitted on | When | What to do | See | | --- | --- | --- | --- | --- | | `ConditionChanged` | TerraformCluster, TerraformMachine, TerraformMachinePool | Any other owned condition changed. It is a `Warning` when the condition enters its bad state, `Normal` otherwise. The note reads `: /: `. | Find the type and reason in the [conditions table](). | [Conditions]() | | `DestructivePlanBlocked` | TerraformCluster, TerraformMachinePool | An apply stopped before a plan that deletes or replaces resources: a cluster apply, or a pool apply of a change of the cluster’s exports. Once per blocked Job, in place of `JobFailed`. | Review the plan summary in the condition and approve the hash it names, or change the inputs. A blocked pool change is held while the pool keeps applying its last exports. | [Destructive-plan guard]() | | `DigestUnknown` | TerraformCluster, TerraformMachine, TerraformMachinePool | No image digest is pinned, so an operation runs the spec’s image reference, or a Job succeeded without a readable digest. | Usually clears after the next successful apply. Pin the image by digest in the spec if you need it fixed. | [Job inputs]() | | `DriftDetected` | TerraformCluster, TerraformMachine, TerraformMachinePool | A drift check found a difference. Once per finding. | Decide whether to accept it or remediate. | [Drift]() | | `ForceUnlocked` | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job was started with a stale state lock to force-unlock: its holder pod no longer exists. | Nothing, unless it repeats. Frequent unlocks mean runners are being killed: check evictions and memory limits. | [The state lock]() | | `IdentityNotAllowed` | TerraformCluster, TerraformMachine, TerraformMachinePool | `IdentityAllowed` entered `False`, whatever the reason: `IdentityNotFound`, `NamespaceNotAllowed` or `SecretNotFound`. Again when the reason or message changes. | Read the condition’s reason and fix it. | [Identity runbook]() | | `IdentitySecretNotFound` | TerraformClusterIdentity | The identity’s credentials Secret went missing. | Recreate the Secret at `spec.secretRef`. | [Identity runbook]() | | `ImageInspectFailed` | TerraformMachineTemplate | The registry could not be read for the template’s capacity labels. | Fix the image reference or the registry credentials. | [Templates]() | | `InfrastructureAbandoned` | TerraformCluster, TerraformMachine, TerraformMachinePool | A deletion was released by `captf.io/abandon-infrastructure` without a destroy. The note names the cause. | Clean up the cloud resources: they keep running and are untracked. | [Held deletions]() | | `InstanceUnhealthy` | TerraformCluster, TerraformMachine, TerraformMachinePool | `InfrastructureHealthy` became `False` for an unhealthy, degraded, stopped or terminated instance. | Look at the instance. Remediation can replace it. | [Machine remediation]() | | `JobDeadlineExceeded` | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job hit `activeDeadlineSeconds`. Once per Job. | Raise the deadline or find what hangs. | [Deadlines]() | | `JobFailed` | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job failed, or an apply or destroy could not start (`ApplyJobSucceeded` is `False` without a Job). Once per Job. | Read the condition and the Job’s logs. | [Failing Jobs]() | | `JobInterrupted` | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job was stopped from outside: a drain, an eviction or a deletion. It retries without backoff. | Nothing, unless it repeats: find what keeps stopping the pod. | [Retries]() | | `OutputsInvalid` | TerraformCluster, TerraformMachine, TerraformMachinePool | `OutputsValid` entered `False`: `OutputsMissing`, `OutputsInvalid`, `FailureDomainMismatch` or `ProviderIDChanged`. Again when the reason or message changes. | Fix the output named in the condition. | [Module contract]() | | `PlanChanged` | TerraformCluster | An approved apply planned other changes and stopped before applying them. Once per such Job. | Review the new plan and approve it. | [Manual plan approval]() | | `RemediationRequested` | TerraformMachine | The owner Machine was annotated with `cluster.x-k8s.io/remediate-machine`. | Expected with a MachineHealthCheck. Check why the instance is unhealthy. | [Machine remediation]() | | `ReplicasManagedExternally` | TerraformMachinePool | Valid autoscaler annotations, but another controller owns `spec.replicas`. | Pick one owner for the replicas. | [Machine pools]() | | `StateLocked` | TerraformCluster, TerraformMachine, TerraformMachinePool | The state lock is held by something other than the object’s runner. | Release it with `force-unlock` once the holder is gone. | [Stale lock]() | | `StateLost` | TerraformCluster, TerraformMachine, TerraformMachinePool | A provisioned object’s state is gone or carries no inputs hash. | Restore a backup. | [Unreadable state]() | | `StateRestoreFailed` | TerraformCluster, TerraformMachine, TerraformMachinePool | A restore Job failed. It is not retried for the same serial. | Read the Job’s logs, then retry with a fresh annotation. | [Other manual actions]() | | `StateUnreadable` | TerraformCluster, TerraformMachine, TerraformMachinePool | The state could not be read: corrupt, encrypted or inconsistent. | See the `StateReadable` reason. | [Unreadable state]() | | `StuckJobDeleted` | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job that could never start, because its per-run Secret was missing and no pod started, was deleted. It starts again. | Nothing, unless it repeats: check quota and API errors in the manager’s log. | [Naming and adoption]() | | `StepFailed` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | A runtime step failed. The note carries the runner’s curated summary, never raw stderr. Needs `--runner-events`. | Read the Job’s logs for the full output. | [Failing Jobs]() | | `RunFinished` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | The run ended. A `Warning` when it did not succeed: failed, interrupted, blocked before a destructive plan, or stopped because the approved plan changed. Needs `--runner-events`. | As for `StepFailed`. | [Failing Jobs]() | ## Informational events These are `Normal` and need no action. They are the record of what the controller did. | Reason | Emitted on | Meaning | See | | --- | --- | --- | --- | | `JobCreated` | TerraformCluster, TerraformMachine, TerraformMachinePool | A Job started: op, attempt, image and why. | [Jobs]() | | `JobSucceeded` | TerraformCluster, TerraformMachine, TerraformMachinePool | An apply, destroy, refresh or drift Job succeeded. | [Jobs]() | | `WaitingForRunLease` | TerraformCluster, TerraformMachine, TerraformMachinePool | An operation waits for the run lease. Once per wait. | [Leases]() | | `WaitingForClusterOperation` | TerraformMachine, TerraformMachinePool | An operation waits for its cluster’s. | [Leases]() | | `WaitingForMachineOperations` | TerraformCluster | An operation waits for its machines’ and pools’. | [Leases]() | | `DestructivePlanApprovalConsumed` | TerraformCluster | The approved destructive apply succeeded and the annotation was removed. | [Destructive-plan guard]() | | `PlanReady` | TerraformCluster | A plan Job planned a change under `Manual`: counts, hash and the approve command. | [Manual plan approval]() | | `PlanApproved` | TerraformCluster | The apply of an approved plan started. | [Manual plan approval]() | | `PlanApplied` | TerraformCluster | The approved plan was applied and the annotation removed. | [Manual plan approval]() | | `DeletionStarted` | TerraformCluster, TerraformMachine, TerraformMachinePool | The first reconcile with a deletion timestamp. | [Deletion and Teardown]() | | `Destroyed` | TerraformCluster, TerraformMachine, TerraformMachinePool | The destroy succeeded and cleanup ran. | [Cleanup]() | | `FinalizerRemoved` | TerraformCluster, TerraformMachine, TerraformMachinePool | The finalizer was removed: after a destroy, with no state, or without an owner. | [Order and finalizers]() | | `Paused` | TerraformCluster, TerraformMachine, TerraformMachinePool | The `Paused` condition became `True`. | [Order and finalizers]() | | `Resumed` | TerraformCluster, TerraformMachine, TerraformMachinePool | The `Paused` condition became `False` again. | [Lifecycle]() | | `Provisioned` | TerraformCluster, TerraformMachine, TerraformMachinePool | `status.initialization.provisioned` latched true. | [Lifecycle]() | | `ProviderIDSet` | TerraformMachine, TerraformMachinePool | `spec.providerID` was written. | [Machine role]() | | `ControlPlaneEndpointSet` | TerraformCluster | `spec.controlPlaneEndpoint` was written from the module output. | [Cluster role]() | | `FailureDomainsChanged` | TerraformCluster | `status.failureDomains` changed. | [Cluster role]() | | `InputsChanged` | TerraformCluster, TerraformMachine, TerraformMachinePool | The inputs hash differs from the state’s and an apply starts. | [Choosing the operation]() | | `DigestPinned` | TerraformCluster, TerraformMachine, TerraformMachinePool | An image digest was recorded, or re-pinned after an apply. | [Job inputs]() | | `StateAdopted` | TerraformCluster, TerraformMachine, TerraformMachinePool | The state a successful apply wrote was adopted with its inputs hash. | [Terraform State]() | | `StateBackedUp` | TerraformCluster, TerraformMachine, TerraformMachinePool | A new state serial was copied into a backup. | [Backups]() | | `StateRestored` | TerraformCluster, TerraformMachine, TerraformMachinePool | A restore Job pushed a backup and the annotation was removed. | [State restore]() | | `DriftResolved` | TerraformCluster, TerraformMachine, TerraformMachinePool | `DriftDetected` went from `True` to `False`. | [Drift]() | | `DriftRemediationStarted` | TerraformCluster, TerraformMachine, TerraformMachinePool | An apply remediating drift started. | [Drift]() | | `InstanceHealthy` | TerraformCluster, TerraformMachine, TerraformMachinePool | `InfrastructureHealthy` became `True`. | [Drift and health]() | | `RemediationWithdrawn` | TerraformMachine | The instance read healthy again and the remediation annotation was removed. | [Machine remediation]() | | `ReplicasWrittenBack` | TerraformMachinePool | An autoscaled pool’s observed replicas were written to `MachinePool.spec.replicas`. | [Machine pools]() | | `IdentitySecretFound` | TerraformClusterIdentity | The credentials Secret appeared. | [Identities]() | | `MirrorCreated` | TerraformCluster, TerraformMachine, TerraformMachinePool | The namespace’s credential mirror was created for this object. | [Identities]() | | `OwnerReferencesRepaired` | TerraformCluster, TerraformMachine, TerraformMachinePool | Secrets of this object (state, backups, durable inputs, plan key, its mirror entry) were owned by it again after a restore, or because a chunk had no owner reference. | [State adoption]() | | `MirrorRemoved` | TerraformCluster, TerraformMachine, TerraformMachinePool | The mirror was deleted: its last user went, or the identity no longer allows the namespace. | [Cleanup]() | | `CapacityResolved` | TerraformMachineTemplate | A template’s capacity or node info changed from its image labels. | [Templates]() | | `RunStarted` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | The runtime is ready and the first step is about to run. | [Job Environment]() | | `StepStarted` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | A runtime step started. | [Job Environment]() | | `StepSucceeded` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | A runtime step finished. | [Job Environment]() | | `PlanSummary` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | A plan the runner parsed: counts only. | [Job Environment]() | | `ResourcesChanged` | TerraformCluster, TerraformMachine, TerraformMachinePool (runner) | What an apply or destroy step changed: counts only. | [Job Environment]() | > [!NOTE] > > **See also** > > - [Conditions](). > - [Observability](). # Failing Jobs This page helps you diagnose a Job that failed, or an object whose apply, destroy, drift check or health refresh keeps failing. It covers the `CAPTFJobFailing` and `CAPTFNoRecentSuccess` alerts and reading a failure directly from an object’s status, for a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`. > [!NOTE] > > **Before you begin** > > - `kubectl` access to read the failing object and its Job’s pod logs, in its namespace, on the management cluster. > - The object’s kind, namespace and name. `CAPTFJobFailing` carries only `kind` and `op`, so find the object through its conditions (step 1); `CAPTFNoRecentSuccess` already carries `namespace` and `name`. ## 1\. Find the failing objects ```sh kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \ | jq -r '.items[] | select(any(.status.conditions[]?; (.type=="ApplyJobSucceeded" or .type=="DriftJobSucceeded") and .status=="False")) | "\(.kind) \(.metadata.namespace)/\(.metadata.name)"' ``` `ApplyJobSucceeded` covers apply and destroy; `DriftJobSucceeded` covers drift checks and health refreshes. `CAPTFNoRecentSuccess` fires when a scheduled drift check or refresh has not succeeded for six hours, which usually means one of these Jobs is failing repeatedly rather than missing once. ## 2\. Read the condition’s reason The reason on `ApplyJobSucceeded` or `DriftJobSucceeded` says where to look next; see [Conditions]() for the full list of reasons each condition can carry. The ones that matter here: | Reason | What it means | Go to | | --- | --- | --- | | `ApplyFailed`, `DestroyFailed` or `DriftJobFailed` | A Job ran and its runtime failed partway through. `status.lastRun` carries the detail. | Step 3 | | `ImagePullFailed` | A container stayed in `ErrImagePull` or `ImagePullBackOff` until the Job’s deadline. No step ever ran, so `status.lastRun` carries nothing useful. | [Image pull failures](<#image-pull-failures>) | | `ImageInvalid` | The image does not follow the module contract. | [Image layout errors](<#image-layout-errors>) | | `JobDeadlineExceeded` | The Job’s pod ran past `activeDeadlineSeconds` without finishing. | [Deadline exceeded](<#deadline-exceeded>) | | `DestructivePlanBlocked` or `PlanChanged` | Not a failure: an apply stopped on purpose, waiting for an approval, and does not count toward `CAPTFJobFailing`. | [Plan Approval]() | ## 3\. Read status.lastRun for a completed run When the runner actually started and produced a result — every reason above except `ImagePullFailed` and a `JobDeadlineExceeded` that hit before any step ran — `status.lastRun.error.kind` classifies what happened. See the full field list in [Last run](). - `step`: a runtime command failed. `error.step` names it — usually `init`, `validate`, `plan`, `show-json`, `apply`, `apply-refresh-only` or `destroy`; occasionally `force-unlock` (retrying past a stale lock left by a previous run; see the [stale state lock runbook]()) or `prepare` (the runner’s own environment setup, before any runtime command ran) — and `error.summary` gives the runner’s curated one-line reason, at most 512 bytes — the module’s or the cloud provider’s own error, never raw output. Read the rest from the Job’s pod logs (step 4). - `image-layout`: the image does not follow the module contract; the same cause as `ImageInvalid` above. See [Image layout errors](<#image-layout-errors>). - `interrupted`: the runner was sent `SIGTERM` before it could finish — a node drain, an eviction, the Job being deleted, or the deadline. It is retried without counting toward backoff; see [Retry backoff](<#retry-backoff>). - `blocked` or `plan-changed`: see [Plan Approval](). `status.lastRun.steps` lists every step the runner completed, with its exit code and duration — useful for finding a slow step even when the run as a whole succeeded. ## 4\. Read the Job’s pod logs ```sh kubectl get job -n kubectl logs -n job/ -c source ``` `` is `status.lastRun.job` for a finished run, or `status.activeJob.name` for one still running. The `source` container runs the pinned image and the module; a separate init container only copies the runner binary into it and rarely fails on its own (that failure surfaces as an image pull failure on the manager’s `--runner-image` instead, a cluster-wide problem rather than one specific to this object). `status.lastRun.error.summary` is a curated, size-limited copy of the failing step’s output, not the raw log, because the object’s status is readable by anyone who can `get` it — the full log is only in `kubectl logs`. By default the newest three failed Jobs of each operation are kept (`spec.jobs.failedJobsHistoryLimit`; see [Tuning Jobs]()), so the pod and its logs are usually still there for a fresh alert. `kubectl events --for / -n --types=Warning` also carries a curated summary of each failure as a `JobFailed` or `StepFailed` event, even after the Job itself is pruned. ## Image pull failures `ImagePullFailed` means a container was stuck pulling its image until the Job’s deadline — never that the module ran and failed. Two images can be at fault: - the `source` container’s image, `spec.source.image` — check for a typo, a missing pull secret, or a tag that was deleted from the registry; - the runner init container’s image, the manager’s `--runner-image` (defaults to the manager’s own image) — a cluster-wide problem, not specific to this object. ```sh kubectl get pods -n -l batch.kubernetes.io/job-name= kubectl describe pod -n ``` The `Events` section names the image and the pull error (`ErrImagePull`/`ImagePullBackOff`). Fix the image reference or the pull credentials (`spec.jobs.imagePullSecrets`; see [Tuning Jobs]()), then let the next reconcile retry — no manual retrigger is needed. ## Image layout errors `ImageInvalid` (`status.lastRun.error.kind: image-layout`) means the runner started but the image itself does not follow the module contract: it is missing the module directory the contract requires, or its entry point is not executable. See [Image Contract]() for the layout an image must follow. Republish the image and point `spec.source.image` at the new tag or digest to try again. > [!WARNING] > > **No retry fixes an image layout error** > > This is a problem with the image, not with the object’s spec or a transient failure. ## Deadline exceeded `JobDeadlineExceeded` means the Job’s pod ran past `spec.jobs.activeDeadlineSeconds` (one hour unless set) without finishing. Kubernetes sends the runner `SIGTERM`; it gets most of the Job’s termination grace period to finish an in-flight provider call and write state before being killed, so infrastructure created before the deadline is not lost even though the Job is marked failed. Common causes: - a module that is simply slow relative to its deadline — raise `spec.jobs.activeDeadlineSeconds`; see [Tuning Jobs](); - a state lock wait: the first step that locks the state (`plan`, `apply`, `apply-refresh-only` or `destroy` — `init` never takes the lock) waits at most `spec.jobs.lockTimeoutSeconds` (five minutes unless set) for it before failing on its own — well inside the default deadline, so a lock wait alone should rarely be the cause; see [Slow Jobs]() and the [stale state lock runbook]() if it is; - a provider call that never returns — the module’s or the cloud API’s problem; the log up to where it stopped is the only record when the runner is killed before it can write a result. ## Retry backoff A Job that failed for `step`, `image-layout` or a pull failure counts toward that operation’s retry backoff; a `blocked` or `plan-changed` Job never counts, since it waits for an approval instead of a timer. See [The Reconcile Lifecycle]() for the delay itself. A deadline kill always counts, even when the runner caught its `SIGTERM` in time and reported itself `interrupted` (`status.lastRun.error.kind: interrupted`): a step that always hangs must reach the retry cap. Only an interruption that is not a deadline kill, such as a drain or an eviction, retries on the next reconcile with no wait. See [Retries and backoff](). > [!WARNING] > > **There is no way to skip a real backoff by hand** > > Retrying immediately against an unresolved cause (a bad image, a still-held lock) only fails again the same way. ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n -o json \ > | jq '.status.conditions[] | select(.type=="ApplyJobSucceeded" or .type=="DriftJobSucceeded")' > ``` > > The condition returns to `status: "True"`, and the next `status.lastRun.error` is empty. > [!NOTE] > > **See also** > > - [Approvals and Gates]() for the conditions and events of a waiting plan or a blocked apply. > - [Slow Jobs]() — a Job that runs but takes far longer than expected. > - [Stale State Lock]() — a lock held long enough to fail a Job on its own. > - [Plan Approval]() — approving a blocked or changed plan. > - [Tuning Jobs]() — deadlines, lock timeouts, history limits and pull secrets. # Slow Jobs This page helps you find out why a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`’s Jobs are taking a long time — either to run, or to start. It covers the `CAPTFJobSlow` and `CAPTFJobQueueSlow` alerts. The two measure different things and point at different causes, so start by working out which one fired. > [!NOTE] > > **Before you begin** > > - `kubectl` access to the object’s namespace and to Leases and pods there, on the management cluster. > - Both alerts carry `kind` and `op`, not an object name; find the specific object through `status.activeJob` or the Jobs of that kind and op. ## 1\. Tell the two alerts apart - `CAPTFJobSlow` measures the wall time of a **completed** Job, start to finish. Go to [A Job that runs long](<#a-job-that-runs-long>). - `CAPTFJobQueueSlow` measures the time from a Job’s **creation** to its module container actually starting. Go to [A Job slow to start](<#a-job-slow-to-start>). Neither counts time spent waiting *before* a Job is even created — that wait shows up as a condition stuck in an `Unknown` state instead. If `status.activeJob` names no Job yet the object is not idle either, go to [Waiting for a lease](<#waiting-for-a-lease>) first. ## Waiting for a lease Before creating a Job, the controller takes the object’s own run lease (`captf-run-`), and, for an apply, destroy or restore that names a cluster, the Cluster’s write lease (`captf-cluster-`) as well. See [Run leases and the cluster operation gate]() for how the two work together, and [Configuration]() for the `--cluster-operation-gate` flag. While either is held by someone else, `ApplyJobSucceeded` (or `DriftJobSucceeded` for a refresh or drift, `RestoreJobSucceeded` for a restore — see [Restoring State]()) stays `Unknown`, with one of these reasons, and no Job exists yet: | Reason | What it means | | --- | --- | | `WaitingForRunLease` | Another live Job, usually another manager instance’s, already holds this object’s run lease. | | `WaitingForClusterOperation` | A `TerraformMachine`’s or `TerraformMachinePool`’s operation is waiting for its `TerraformCluster`’s to finish first. | | `WaitingForMachineOperations` | A `TerraformCluster`’s operation is waiting for its machines’ and machine pools’ applies and destroys, already in flight, to finish first; any new one of theirs waits for the cluster in turn. | The condition’s message names the Lease and the Job (or manager) holding it, and the controller checks again every 30 seconds — this is normal and usually resolves on its own once the holder finishes. ```sh kubectl get lease -n -l captf.io/lease=run,captf.infrastructure.cluster.x-k8s.io/owner-name= kubectl get lease -n -l captf.io/lease=cluster ``` `spec.holderIdentity` is the holder’s Job name; the `captf.io/lease-op` and `captf.io/lease-acquired-at` annotations say what it is running and since when. A lease is only ever taken over by the controller itself, once its holder Job has finished, does not exist after a short grace period, or is older than its own deadline plus a backstop. > [!CAUTION] > > **Never delete or edit a Lease by hand** > > If its holder Job is genuinely stuck, that is a [failing Job]() or a [slow one](<#a-job-that-runs-long>) to chase down instead, not a lease problem. A wait that outlasts the holder’s `activeDeadlineSeconds` means the holder cannot finish; look at that Job directly. ## A Job that runs long `CAPTFJobSlow` fires on the Job’s own duration, so the module (or the cloud it calls) is doing the work, not the controller withholding anything. Narrow it down: ```text # p90 duration by step, for this kind and op histogram_quantile(0.9, sum by (le, step) (rate(captf_job_step_duration_seconds_bucket{kind="",op=""}[6h]))) ``` `status.lastRun.steps` on the object itself shows the same breakdown for its most recent run, without a Prometheus query. Once you know which step is slow: - **`plan`, `apply`, `apply-refresh-only` or `destroy` is slow**: these are the steps that lock the state (`init` never does), so a slow but *successful* one may have been contending for the lock — every locking step waits at most `spec.jobs.lockTimeoutSeconds` (five minutes unless set) for it before failing on its own. Check `StateReadable` for reason `StateLocked` and see the [stale state lock runbook](). Otherwise it is a large or slow module, or a cloud provider that is itself slow — nothing in the controller to tune. Compare against `spec.jobs.activeDeadlineSeconds` ([Tuning Jobs]()) before it starts hitting the deadline and turning into a [failing Job]() instead of a slow one. - **every step is proportionally slow**: check the Job pod’s own resource use against `spec.jobs.resources` ([Tuning Jobs]()) — a Terraform or OpenTofu process with several providers is memory-hungry, and throttling or swapping shows up as slowness everywhere, not one step. ## A Job slow to start `CAPTFJobQueueSlow` covers everything between the Job existing and its module container running: pod scheduling, both image pulls (the runner init container’s and the module’s own), and the runner binary copy. ```sh kubectl get pods -n -l captf.infrastructure.cluster.x-k8s.io/op= --field-selector=status.phase=Pending kubectl describe pod -n ``` Read the pod’s `Events`: - `FailedScheduling`: exhausted node capacity, a `ResourceQuota`, or a scheduling constraint (affinity, taints) the runner’s pod cannot satisfy. Check `spec.jobs.resources`, `spec.jobs.podSecurityContext` overrides and the namespace’s quota; see [Tuning Jobs](). - `Pulling` / `ErrImagePull` / `ImagePullBackOff`: a slow, rate-limited or unreachable registry. A pull that never succeeds eventually hits the Job’s deadline and becomes a [failing Job]() instead; a pull that succeeds but is merely slow is this alert’s whole story. A controller’s own concurrency limit (`--terraformcluster-concurrency`, `--terraformmachine-concurrency`, `--terraformmachinepool-concurrency`, ten each unless set; see [Configuration]()) does not appear here: it bounds how many objects of a kind reconcile at once, not how fast a Job someone already created gets scheduled and pulled. A concurrency limit too low for the fleet shows up as objects going longer between reconciles overall — check the object’s own `status.observedGeneration` and event timestamps for that, not this alert. ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n -o json | jq '.status.lastRun.steps' > ``` > > The next run’s step durations (or `captf_job_duration_seconds` and `captf_job_queue_seconds` for the kind and op) drop back under the alerts’ thresholds — 30 minutes p90 duration, 5 minutes p90 queue time. > [!NOTE] > > **See also** > > - [Failing Jobs]() — a Job that stops instead of merely running long. > - [Stale State Lock]() — clearing a lock a Job is waiting on. > - [The Reconcile Lifecycle]() — run leases and the cluster operation gate in full. > - [Tuning Jobs]() — deadlines, lock timeouts and resources. # Reconcile Errors This page helps you diagnose the `CAPTFReconcileErrors` alert: one of the manager’s own controllers is repeatedly failing to reconcile, as opposed to an object’s Job failing (see [Failing Jobs]()). A reconcile error means the controller could not even finish deciding what to do — it is not a statement about any one object’s infrastructure. > [!NOTE] > > **Before you begin** > > - `kubectl` access to the manager’s Deployment and its logs, in the `captf-system` namespace (or wherever it is installed), on the management cluster. ## 1\. Read the manager’s logs ```sh kubectl logs -n captf-system deploy/captf-controller-manager --tail=200 ``` Each reconcile that returns an error logs a line carrying `controller` (one of `terraformcluster`, `terraformmachine`, `terraformmachinepool`, `terraformmachinetemplate` or `terraformclusteridentity` — the alert’s `controller=~"terraform.*"` matches all five), the object’s `namespace` and `name`, and the error itself. Reconciles for a `TerraformMachine`, `TerraformCluster` or `TerraformMachinePool` also carry that object’s own key (`TerraformMachine`, `TerraformCluster`, `TerraformMachinePool`) plus its owning `Cluster`, `Machine` or `MachinePool` where one exists, so you can filter by any of those instead of grepping free text. The default log level (`-v=2`) already includes the manager’s own flow logging in addition to errors, so raising `-v` is rarely needed just to see that a reconcile failed; it helps to see *why* a decision was made before the error, not to see the error itself. `--diagnostics-address`’s endpoint can change the running level without a restart when `--insecure-diagnostics` is not set; see [Manager Flags](). ## 2\. Narrow down which object, and how often ```text sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m])) ``` matches the alert’s own query. If it is one object erroring repeatedly, its logs name it on every line; if it is spread across many objects of one kind, look for a cause common to the whole namespace or cluster (quota, RBAC, an API server problem) rather than the object’s own spec. ## Common causes | Cause | What you see | | --- | --- | | **The Kubernetes API server is unavailable or throttling requests** | A `Get`, `Create`, `Update` or `Patch` failed with a server-side error (`500`, `503`, or a `429` from client-side rate limiting). Transient: the error clears once the API server does; if it does not, check the API server’s own health. | | **The manager’s own patch was rejected by the validating webhook** | The manager itself writes a `Terraform*` object’s finalizer and a handful of spec fields (for example a `TerraformMachine`’s `providerID`, a `TerraformCluster`’s `controlPlaneEndpoint`) through the same admission path a user’s `kubectl apply` goes through. If the webhook cannot be reached, that patch fails and shows up here. See [Webhook Unavailable](). | | **The manager’s own RBAC is missing a permission** | The `ClusterRole` it runs as (`captf-manager-role`) was edited by hand, or an upgrade changed what it needs and the installed `ClusterRole` was not updated to match. The error names the verb, resource and group a `403` `Forbidden` response refused. See [RBAC]() for what the manager needs and creates. | | **A namespace `ResourceQuota` blocks a create** | The manager creates Jobs, Secrets, Leases, and a per-namespace runner `ServiceAccount` and `RoleBinding` on an object’s behalf. A quota on any of those object counts in the tenant namespace surfaces as a `Forbidden` create error here rather than anywhere on the `Terraform*` object’s own status. | | **A conflicting concurrent write** | Two updates to the same object raced (for example, the manager and an operator editing it at the same moment). The condition-patching helper already retries a conflicting status write itself; a `Conflict` that still reaches the log usually clears on the next reconcile, which controller-runtime’s own per-item backoff already schedules, distinct from a Job’s retry backoff (see [Failing Jobs]()), and not something to act on unless it repeats for the same object. | None of these are the same as an object’s Job failing: a Job failure is recorded on the object’s own `status` and conditions and never increments `controller_runtime_reconcile_errors_total` — only an error the reconciler itself returns does. See [Failing Jobs]() if what you are chasing is instead a failing apply, destroy, drift check or refresh. ## Confirm it worked > [!TIP] > > ```sh > kubectl logs -n captf-system deploy/captf-controller-manager --tail=50 --follow > ``` > > No further `Reconciler error` lines appear for the affected controller, and `sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m]))` returns to 0. > [!NOTE] > > **See also** > > - [Failing Jobs]() — a Job or an object’s status, as opposed to the controller itself. > - [Webhook Unavailable]() — a specific, common cause of reconcile errors. > - [RBAC]() — what the manager’s own `ClusterRole` grants. # Runbook: state that will not read `StateReadable` reports whether the controller could read a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`’s state this reconcile. While it is `False` or `Unknown`, no plan, apply, drift check or refresh Job runs for the object, and `Ready` follows it down. This runbook covers every reason `StateReadable` (and its True counterpart) can carry, what each means, and what to do. For the `CAPTFStateUnreadable` alert itself, see [Observability](). Replace ``, `` and `` below with the object’s namespace, kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`) and name. ## Find the reason ```sh kubectl get -n \ -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}: {.message}{"\n"}{end}' ``` Also check for a `Warning` event: a reason that turns `False` emits one, named either `StateLost`, `StateLocked` or, for the three read errors below, `StateUnreadable`. ```sh kubectl events --for / -n ``` ## Deleting while state is unreadable If the object ever applied, deletion is **held** while its state is missing or unreadable: the destroy cannot run safely without the state, and dropping the finalizer would leave the infrastructure running with no record of it. “Ever applied” means the object was provisioned, a state backup exists, the durable inputs Secret pins an image digest, or that Secret carries the `captf.io/applied: "true"` marker (see [the applied marker](<#the-applied-marker>)). The controller keeps the finalizer, `StateReadable` stays `False` with the reason `StateLost`, `StateEncrypted`, `StateCorrupt` or `StateInconsistent`, a `Warning` event is emitted, and every state backup is preserved. An object that never applied has nothing to lose and still deletes immediately. Two ways out: - **Restore, then destroy.** Set `captf.io/restore-state` to a serial from `status.stateBackups`, as in the [state restore runbook](). On a held delete the restore runs first, and the normal destroy follows once the state reads again. - **Abandon the infrastructure.** Set `captf.io/abandon-infrastructure` to the object’s `metadata.uid`. The controller checks it before restore and destroy, so it works even while a restore is pending. It removes the finalizer without running a destroy and emits a `Warning` event, `InfrastructureAbandoned`, naming the cause. > [!CAUTION] > > **Abandoning leaves the infrastructure running and untracked** > > The infrastructure keeps running and is no longer tracked by anything in Kubernetes; delete it through the cloud provider. The state backups are garbage-collected with the object, so copy out anything you want to keep first (see the [stuck destroy runbook]()). ```sh uid=$(kubectl get -n -o jsonpath='{.metadata.uid}') kubectl annotate -n captf.io/abandon-infrastructure="$uid" ``` The value must equal the UID exactly; the controller ignores any other value, so a stale annotation copied from another object does nothing. A held delete is one of four cases the annotation releases; the [stuck destroy runbook]() lists all of them. An object whose state reads and whose destroy can start is destroyed as usual, even with the annotation set. ### The applied marker The durable inputs Secret `captf-inputs--` carries `captf.io/applied: "true"` once an apply succeeded or a state with an inputs hash was read. Unlike status, it survives `clusterctl move`, so after a move an object whose state is lost still holds with `StateLost`, deleting or not, instead of being deleted without a destroy or applied again from scratch. One limit: deleting the namespace removes that Secret along with the state and backups, so a moved object in a deleted namespace cannot be told apart from one that never applied. ## StateRead The healthy, `True` reason: the state was read this reconcile. Nothing to do. ## StateNotFound `Unknown`. No state Secret exists yet, because the object has never completed an apply, or (for a mutable kind) it was deleted along with the state on a previous destroy. This is not an error: a new object reports it until its first successful apply, and it never fires `CAPTFStateUnreadable`. If an object that used to be provisioned reports this instead of `StateRead`, its state Secret is gone; see **StateLost** below, which is the reason an object that applied before, with a missing state Secret, carries instead. ## StateLost `False`. Either the state Secret of an object that applied before is missing (it need not still be provisioned), or its state has no recorded inputs hash although the object is provisioned (only possible for `TerraformMachine`, which is immutable and so never rebuilds inputs to re-apply from). Either way, applying again is not safe: for the missing-Secret case, a second apply would create a second set of resources next to the live ones instead of managing the ones the lost state recorded; for the missing-inputs-hash case, there is nothing to re-apply. No Job runs, and the condition does not clear on its own. Fix: restore the object’s newest usable state backup with the `captf.io/restore-state` annotation. See the [state restore runbook](), including how to list the backups in `status.stateBackups`. If no backup exists, `StateLost` never clears; the object’s resources still exist in the cloud but nothing in Kubernetes can manage them until a state is restored, or reconstructed out of band with [manual recovery](). The [Total State Loss and Import runbook]() gives the whole flow, including recreating the object with `import` blocks. ## StateLocked `False`. The state’s lock `Lease` is held by something other than this object’s own runner — for example a workstation running `terraform state rm`, or a Job that died without releasing it. The condition’s message names the holder, its operation and when it took the lock; every Job for the object waits `lockTimeoutSeconds` (`spec.jobs.lockTimeoutSeconds`, defaulting to 300 seconds — see [Job tuning]()) for it and then fails. The controller already force-unlocks a lock whose holder pod is gone, the next time it starts a Job for the object; most `StateLocked` conditions clear on their own. See the [stale-lock runbook]() for how that detection works, how to tell a live holder from a stale one by hand, and manual force-unlock. ## StateEncrypted `False`. The state Secret carries OpenTofu’s client-side state encryption envelope. CAPTF v1 cannot read an encrypted state at all — not the outputs, not the resource count, nothing — so this condition never clears on its own and the object gets no drift checks, health checks or further applies. Because an encrypted state fails the reader before it is ever parsed, the manager never took a backup of it either: there is no CAPTF-side backup to restore. Fix: disable client-side state encryption for this object’s module (drop the encryption configuration so future writes are plain state), and use a workstation that holds the decryption key to decrypt the current state and push the plaintext back with `terraform state push` (or `tofu state push`) against the same `kubernetes` backend configuration — the backend config, not the state format, is what CAPTF’s reader needs to match; see [manual recovery]() for the exact `secret_suffix`, `namespace` and `labels` the object’s backend uses. Once the pushed state is unencrypted, the next reconcile reads it normally. ## StateCorrupt `False`. The reader could not treat the Secret’s payload as gzip-compressed Terraform state JSON: the gzip or JSON decoding failed, the decompressed size exceeded the reader’s cap, or (for a multi-Secret chunked state) there were more chunks than the reader accepts. An unsupported state file version is reported the same way. The condition’s message names which of these it was. See [Chunking and size caps]() for the reader’s exact limits. A corrupt state is never backed up by the manager (backups copy only a state the reader could parse), so a backup taken before the corruption is your most recent recoverable copy. Fix: restore the newest backup from before the corruption with the [state restore runbook](). If nothing wrote a backup before the state became corrupt, there is no CAPTF-side recovery: rebuild the state out of band with [manual recovery](), using the resource list from the last known-good backup or the cloud provider’s own inventory as your guide. ## StateInconsistent `False`. The set of state Secrets for the object’s suffix does not add up to one complete, contiguous state: a chunk is missing or duplicated, a Secret in the set has an unexpected name, or the base Secret has no `tfstate` data key. The condition’s message names which of these it was. This can be transient: a runner Job writes a multi-chunk state one Secret at a time, so a reconcile that reads mid-write sees an incomplete set and the next reconcile, after the Job finishes, usually reads a complete one. If it persists past the Job finishing, something outside CAPTF edited or deleted one of the chunk Secrets by hand. As with `StateCorrupt`, an inconsistent state is never backed up, so restore the newest backup from before the inconsistency appeared with the [state restore runbook](); the reader’s own limits are at [Chunking and size caps](). ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n \ > -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}{"\n"}{end}' > ``` > > Expect `True/StateRead`. `status.observedStateSerial` moves to the restored or rebuilt state’s serial, and the object’s next drift check or apply proceeds from it. > [!NOTE] > > **See also** > > - [State restore]() > - [Stale state lock]() > - [Terraform State]() > - [Conditions reference]() > - [Observability]() # Stale State Lock Both Terraform and OpenTofu lock the `kubernetes` backend for the duration of a run and release the lock on a clean exit ([Terraform State]()). Only a `SIGKILL` or a node loss leaves the lock held with nothing left to release it — a stale lock. This page covers finding a held lock, telling a stale one from a live one, and clearing it. It applies the same way to a `TerraformCluster`, a `TerraformMachine` and a `TerraformMachinePool`: all three use the same backend and the same lock mechanics. > [!NOTE] > > **Before you begin** > > - `get` access to Leases and Pods, and `patch` access to the object, in its namespace, on the management cluster. > - Docker or Podman, and pull access to the object’s pinned runtime image, if the automatic force-unlock in [Diagnosis](<#diagnosis>) does not apply and you need [Fix](<#fix>)’s manual run. > - Replace ``, `` and `` below with the object’s namespace, name and Kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`; the lowercase singular works with `kubectl`). Run every command against the management cluster. ## Symptoms - The object reports `StateReadable=False`/`StateLocked`, naming the holder, its operation and when it took the lock. - A Job runs longer than usual, then fails: it waited out `lockTimeoutSeconds` (`spec.jobs.lockTimeoutSeconds`, default 300 s; see [Tuning Jobs]()) for a lock that never cleared. - The [`CAPTFForceUnlocks`]() alert fired: the controller already force-unlocked a stale lock and moved on: not itself a problem, but worth checking why the previous Job died. ## Cause The lock is a `coordination.k8s.io/v1` Lease named `lock-tfstate-default-` in the object’s namespace (`` is `status.stateSecretSuffix`). `spec.holderIdentity` carries the lock ID; the `app.terraform.io/lock-info` annotation carries the JSON lock info (`ID`, `Operation`, `Who`, `Version`, `Created`). `Who` is `@`; in a Job pod the hostname is the pod name. Job pods get a 600-second termination grace period, and on `SIGTERM` (a drain, an eviction, or the Job’s `activeDeadlineSeconds`) the runner gives the runtime up to 570 s to finish in-flight provider calls, write state and release the lock before it is killed. So an ordinary deletion or drain does not leave a stale lock; only a hard kill or a lost node does. ## Diagnosis Before creating a Job, the controller already checks the lock and acts on what it finds: - No holder: proceeds; nothing to do. - Holder present, and its pod is one of the object’s own runner Job pods and still exists (not finished): the lock is live; the next Job waits out `lockTimeoutSeconds` for it. - Holder present, and its pod is one of the object’s own runner Job pods but no longer exists: stale. The controller passes `--force-unlock` to the next Job, which force-unlocks it after `init`, emits an event, and continues. - Holder present, but it is not recognizable as one of the object’s own runner pods — the `Who` field has no `@`, or its hostname is not a pod of this object’s Jobs, for example a workstation running `terraform state rm`: **never force-unlocked automatically**, whatever else is true. The object reports `StateReadable=False`/`StateLocked` naming the holder, its operation and when it took the lock. In most cases waiting for the object’s next reconcile (it retries with backoff) resolves a stale lock without any manual step. Manual inspection and force-unlock below are for the last case, where the controller will never act on its own. To inspect the lock by hand: ```sh kubectl get lease -n lock-tfstate-default- -o yaml ``` Read `spec.holderIdentity` (the lock ID) and the `app.terraform.io/lock-info` annotation. The part of `Who` after the last `@` is the pod name; confirm whether it still exists and has finished: ```sh kubectl get pod -n ``` - **Pod exists and has not finished:** the lock is live. Do not force unlock; a concurrent run against the same state would be unsafe. Let the holder finish, or find out why it is still running. - **Pod is gone, or finished (`Succeeded`/`Failed`):** if it is one of this object’s own runner pods, the controller force-unlocks it automatically the next time it creates a Job (trigger a reconcile, for example by waiting for the next requeue, or by editing an annotation). If it is not one of this object’s own runner pods, force-unlock by hand below. - **`Who` has no `@` (holder unknown):** the controller never force-unlocks this automatically. Confirm independently that no process still holds the lock, then force-unlock by hand. ## Fix > [!CAUTION] > > **Force-unlock only a lock whose holder is gone** > > Never force-unlock a live lock: a concurrent run against the same state would be unsafe. Confirm independently that no process still holds the lock first. Run the pinned image’s own `terraform`/`tofu` binary against the backend, the same pattern as the [stuck-destroy runbook]()’s manual recovery: ```sh docker run --rm --entrypoint /captf/runtime \ -v "$PWD/root:/captf/work/root" -w /captf/work/root \ @ init -input=false -no-color \ -backend-config=secret_suffix= -backend-config=namespace= \ -backend-config=in_cluster_config=true -backend-config=labels= docker run --rm --entrypoint /captf/runtime \ -v "$PWD/root:/captf/work/root" -w /captf/work/root \ @ force-unlock -force ``` `@` comes from the object’s durable inputs Secret (`captf.io/image`/`captf.io/image-digest`; drop any `:tag` from `image`, append `@` and the digest). `` must be the same HCL object the Job itself would pass, or `init` reads an empty state: read it off the state Secret’s own `.metadata.labels` the same way as the [stuck-destroy runbook](). `force-unlock` touches no resources, so `./root/main.tf.json` needs only a backend declaration, not the full rendered module: ```json { "terraform": { "backend": { "kubernetes": {} } } } ``` (an empty `./root/terraform.tfvars.json`, `{}`, avoids a missing-file warning, though force-unlock never reads it). `force-unlock` needs an initialized backend first, unlike `apply` or `destroy`, so `init` runs first with the same `-backend-config` flags the Job itself would pass — see [Job Environment]() for the exact set, and note that `in_cluster_config=true` needs this run from inside the cluster. This clears the holder and the `lock-info` annotation; it does not delete the Lease object itself — only a workspace delete removes it, and the `default` workspace (the only one CAPTF ever uses) cannot be deleted. ## Confirm > [!TIP] > > ```sh > kubectl get lease -n lock-tfstate-default- \ > -o jsonpath='{.spec.holderIdentity}' > ``` > > Empty output means the lock is clear. The object’s next reconcile starts a Job normally; `StateReadable` clears to `True`/`StateRead` once the reconcile after that Job finishes reads the state. > [!NOTE] > > **See also** > > - [Terraform State]() — the backend, its Secret naming and how the lock fits in. > - [Stuck destroy runbook]() — the same manual-run pattern, for a destroy that cannot succeed. > - [Conditions reference]() — every `StateReadable` reason. # Runbook: state or inputs near a size limit CAPTF stores an object’s state and its rendered module inputs in Kubernetes Secrets, which cap out well before an unbounded Terraform or OpenTofu project would. Two alerts warn before that cap: `CAPTFStateNearSecretLimit` and `CAPTFInputsNearLimit`; a third condition reason, `InputsTooLarge`, marks the point where the cap is already hit. This runbook covers all three: what each limit is, how to find the object tripping it, and how to shrink what it is trying to store. Replace ``, `` and `` below with the object’s namespace, kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`) and name. ## CAPTFStateNearSecretLimit: the state is approaching a Secret’s size cap **The limit.** A Kubernetes Secret holds at most 1 MiB. OpenTofu’s `kubernetes` backend writes the whole state into a single Secret and fails to save a state larger than that; Terraform’s splits an oversized state across additional `-part-N` Secrets instead, up to the reader’s own cap of 32 chunks. Either way, `captf_state_bytes` (the compressed size) crossing 900 KiB fires the alert — for OpenTofu that is a warning before a hard failure, for Terraform a warning about growth. See [Chunking and size caps]() for the exact numbers and how CAPTF reads a chunked state. **Find it.** ```sh kubectl get -n -o jsonpath='{.status.stateSecretSuffix}{"\n"}' kubectl get secret -n -l tfstate=true,tfstateSecretSuffix= -o name ``` One result means an unchunked (or not yet chunked) state; more than one means Terraform has already split it into `-part-N` Secrets. For the exact compressed size, read `captf_state_bytes{namespace="",name=""}` from Prometheus, or sum the chunks by hand: ```sh kubectl get secret -n -l tfstate=true,tfstateSecretSuffix= \ -o jsonpath='{range .items[*]}{.data.tfstate}{"\n"}{end}' | while read -r c; do echo "$c" | base64 -d | wc -c done ``` Compare the total against `captf_state_resources` (also per object) to see how many managed resources it is spread across. **Shrink it.** The state holds every managed resource’s full attribute set, not just what you set in configuration: - Split a large module into more than one `TerraformCluster` or `TerraformMachine` (or use a `TerraformMachinePool` for repeated instances instead of one resource per machine): each gets its own state. - Drop attributes you do not need from the state by not managing them: large inline blocks such as embedded certificates, rendered cloud-init or user-data templates, or full API responses stored in a resource’s computed attributes are common culprits. Move that content to a source the module reads at apply time (a Secret, a bucket object) instead of a resource argument Terraform tracks verbatim. - For a resource with many similar instances (`count` or `for_each`), each instance’s full attribute set is stored once; reducing the count reduces the state proportionally. ## CAPTFInputsNearLimit and InputsTooLarge: the rendered inputs are approaching or over the cap **The limit.** The rendered `main.tf.json` plus `terraform.tfvars.json` (what the module actually runs against) is capped at 1,000,000 bytes, reported per object on `captf_inputs_bytes`. `CAPTFInputsNearLimit` warns above 900,000 bytes; at or above the cap, no Job starts and `ApplyJobSucceeded` turns `False` with reason `InputsTooLarge`. See [Limits]() for how the size is computed and what counts toward it. **Find it.** ```sh kubectl get -n \ -o jsonpath='{range .status.conditions[?(@.type=="ApplyJobSucceeded")]}{.status}/{.reason}: {.message}{"\n"}{end}' ``` `InputsTooLarge`’s message names the operation that could not render. The `captf_inputs_bytes` gauge is set to the rendered size even when it was too large to run, so it is set for this object right up to the cap. **Shrink it.** The three contributors are bootstrap data, cluster exports and user variables: - **`bootstrap_data`** (`TerraformMachine` and `TerraformMachinePool` only): the bootstrap provider’s cloud-init or ignition content, carried base64-encoded (roughly 4/3 its raw size). Trim what the bootstrap provider generates — fewer files, less embedded content — through its own configuration (KubeadmConfig, RKE2Config, or your bootstrap provider’s equivalent), not through CAPTF. - **`captf_cluster_outputs`** (machines and pools): whatever the cluster module’s `exports` output publishes, handed to every machine and pool of the cluster. See [What a cluster hands to its machines and pools](). Publish only what the machine or pool module actually needs from `exports`, not the cluster module’s full internal state. - **User variables** (`spec.variables`, `spec.variablesFrom`): trim large inline values, especially ones duplicated across many objects that could instead reference a `ConfigMap` or `Secret` your module reads directly at run time rather than passing through tfvars. See [Module Variables]() for how variables are set and merged. ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n \ > -o jsonpath='{.status.conditions[?(@.type=="ApplyJobSucceeded")].reason}{"\n"}' > ``` > > For `InputsTooLarge`, the next reconcile after the inputs shrink re-renders and, once it fits, starts the Job; the alerts clear once `captf_state_bytes` or `captf_inputs_bytes` drops back under 900 KiB / 900,000 bytes for their `for` window. > [!NOTE] > > **See also** > > - [Terraform State]() > - [Job Inputs]() > - [Module Variables]() > - [Conditions reference]() > - [Observability]() # My Object Will Not Delete An object with a `deletionTimestamp` that does not go away is waiting on one of a short list of things. Read the object first: ```sh kubectl get -n -o yaml ``` Look at `metadata.finalizers` (CAPTF’s must be the one holding it), the `Deleting` condition’s message, `DeletionBlocked`, `StateReadable` and `ApplyJobSucceeded`, then follow the chart. ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD S[Object has a deletionTimestamp
and the finalizer stays] --> P{Paused=True?} P -- yes --> P1[Unpause the Cluster or remove
the paused annotation] P -- no --> D{DeletionBlocked=True?} D -- yes --> D1[Delete the machines and pools
it counts, then wait] D -- no --> J{A Job is running?} J -- yes --> J1[Wait: the destroy starts after it] J -- no --> W{ApplyJobSucceeded
is a lease wait?} W -- yes --> W1[Wait or inspect the lease holder] W -- no --> R{StateReadable
False?} R -- "StateLocked" --> R1[Stale lock runbook] R -- "Lost, Corrupt,
Encrypted, Inconsistent" --> R2[Restore a backup,
or abandon] R -- no --> C{Deleting message says
it waits for credentials?} C -- yes --> C1[Fix the named condition,
or abandon] C -- no --> F{ApplyJobSucceeded
False?} F -- "IdentityNotAllowed" --> F1[Allow the namespace again,
or abandon] F -- "DestroyFailed" --> F2[Fix the failure,
or abandon] F -- "JobDeadlineExceeded" --> F3[Raise the deadline] F -- no --> Z[No condition explains it:
check the manager] ``` ## Symptom to action | What you see | Meaning | Action | | --- | --- | --- | | `Paused=True`, or the Cluster has `spec.paused: true`; `Deleting` says `Deletion waits until the object is unpaused` | A paused object runs only bookkeeping, so it never destroys; `clusterctl move` relies on this | Unpause. See [Order]() | | `kubectl delete terraformmachine` is refused: `delete the Machine instead` | A live Machine references it | Delete the Machine; see [Order]() | | `DeletionBlocked=True`/`DependentsExist` on a `TerraformCluster` | Machines or pools with its cluster label still exist | List them with `-l cluster.x-k8s.io/cluster-name=`; they delete through their Machines. If one is stuck, work on that object first | | A Job is running (`status.activeJob`) | The destroy waits for it | Wait; see [Slow Jobs]() | | `ApplyJobSucceeded=Unknown`/`WaitingForRunLease`, `WaitingForClusterOperation` or `WaitingForMachineOperations` | The destroy waits for a lease | Wait; see [Leases]() | | `StateReadable=False`/`StateLost`, `StateCorrupt`, `StateEncrypted` or `StateInconsistent` | Held on the state | [Restore]() or [abandon](). See [Unreadable State]() | | `StateReadable=False`/`StateLocked` | A foreign holder has the state lock; the destroy Job will wait and fail | [Stale State Lock]() | | `Deleting` message `The destroy Job waits for its credentials: …` | Credentials cannot be prepared, often in a terminating namespace | Fix the named condition ([identities]()), or [abandon](). See [Terminating namespaces]() | | `ApplyJobSucceeded=False`/`IdentityNotAllowed` | The identity no longer allows the namespace, or is gone | Allow the namespace again, or abandon | | `ApplyJobSucceeded=False`/`DestroyFailed` with a Job | The destroy failed; it retries with backoff forever | [Failing Jobs](), then [Stuck Destroy]() | | `DestroyFailed`, message `The durable inputs Secret is missing` | The destroy cannot be rendered | [Restore the Secret](), or abandon | | `ApplyJobSucceeded=False`/`JobDeadlineExceeded` | The destroy ran out of time | Raise `activeDeadlineSeconds`; see [Tuning Jobs]() | | `ApplyJobSucceeded=False`/`ImagePullFailed` or `ImageInvalid` | The pinned image cannot run | [Failing Jobs]() | | Nothing explains it | The manager is not reconciling | [Reconcile Errors](), [Webhook Unavailable]() | > [!NOTE] > > **A message that names the abandon annotation** > > It tells you the controller sees no way to proceed alone. The [abandon]() page lists exactly which cases the annotation releases, and an object whose state reads and whose destroy can start is destroyed anyway. > [!CAUTION] > > **Abandoning or stripping leaves cloud resources running** > > A destroy that cannot be completed and an object that must go anyway end the same way: back up what you need, clean up the cloud resources, then either [abandon]() or [strip the finalizer](). > [!NOTE] > > **See also** > > - [Stuck Destroy](). > - [Runbooks](). # Held Deletions A destroy needs the state. When the state is missing or unreadable, the controller cannot tell what the module created, and removing the finalizer would leave that infrastructure running with nothing tracking it. So it holds the deletion instead: no destroy runs, the finalizer stays, and the state backups stay with it. This page covers what triggers a hold, what ends one, and the abandon annotation, which ends it without a destroy. ## Ever applied A missing state means two different things. For an object that never applied there is nothing to destroy, and the finalizer comes off at once. For one that did, the state was lost. The controller tells them apart with `everApplied`, which is true when **any** of these holds: | Signal | Where it lives | Survives `clusterctl move` | | --- | --- | --- | | `status.initialization.provisioned` is true | The object’s status | No | | The `captf.io/applied: "true"` marker | The durable inputs Secret `captf-inputs-*`, set at the first successful apply or restore, never cleared | Yes, it moves with the Secret | | A pinned image digest | The same Secret; only a successful apply pins one, but an image change clears it | Yes | | Any state backup exists | The `captf-state-backup-*` Secrets | Yes (owned by the object) | With `--state-backups=0` no backup is ever taken, so the last signal is absent; the other three still hold. The marker exists because status does not move: see [`clusterctl move`](). The same test guards the [ownerless release](): a deleting object with no owner is dropped without a look at the state only if it never applied. ## What holds a deletion `StateReadable` becomes `False` and the deletion is held for these reasons: | Reason | State | Notes | | --- | --- | --- | | `StateLost` | No state Secret, and the object applied before | The Secret was deleted by hand, or the namespace was | | `StateCorrupt` | The payload cannot be decoded, exceeds the reader’s caps, or has an unsupported version | Restore a backup taken before it broke | | `StateEncrypted` | The state carries OpenTofu’s client-side encryption envelope; CAPTF cannot read it | No backup exists (an unreadable state is never backed up) | | `StateInconsistent` | The chunks do not form one complete state | A missing or duplicated chunk, or a stray Secret in the set | The condition message adds how to leave the hold: the restore annotation, and the abandon annotation with this object’s UID. Each reason’s cause and repair are in [Unreadable State](). `StateLocked` is **not** a hold. The state reads, so the destroy starts, and its Job waits `lockTimeoutSeconds` for a lock someone else holds and then fails. See [Stale State Lock](). ## What ends a hold ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A[Deleting, state lost or unreadable] --> B{abandon annotation
equals the UID?} B -- yes --> X[Abandon: cleanup, finalizer off,
InfrastructureAbandoned Warning] B -- no --> C{restore-state names
a complete backup?} C -- yes --> R[Restore Job] R --> S{state reads again?} S -- yes --> D[Destroy Job, then cleanup] S -- no --> A C -- no --> H[Held: requeue every minute] ``` The order inside one pass matters: the controller checks the abandon annotation **before** the restore and the destroy decisions, so a pending restore that cannot start does not hide it. ### Restore, then destroy Setting `captf.io/restore-state=` to a serial in `status.stateBackups` during a hold starts a restore Job. Once the state reads again, the normal destroy follows. This is the only time a restore runs on a deleting object: with readable state the annotation is ignored and the destroy goes ahead. The restore takes the same leases as any other, is never gated, and is not retried for the same serial after a failure. See [Backups and restore]() and the [State Restore runbook](). ### Abandon `captf.io/abandon-infrastructure=` removes the finalizer without a destroy. The value must equal the object’s UID exactly; any other value is ignored, and a held object says so in its `StateReadable` message. The annotation releases a deletion in these cases, and the controller records which one in the event: | Case | How the controller sees it | Cause in the event | | --- | --- | --- | | Held on the state | `StateReadable` is `False` | `StateReadable ` | | The last destroy failed | `status.lastRun` is a destroy and `ApplyJobSucceeded` is `False` (also after a restore) | `the last destroy failed: ApplyJobSucceeded ` | | The durable inputs are gone | `ApplyJobSucceeded=False`/`DestroyFailed` with no Job, because the destroy cannot be rendered | `the destroy cannot start: …` | | The identity does not allow the namespace | `ApplyJobSucceeded=False`/`IdentityNotAllowed` | `the destroy cannot start: …` | | The credentials cannot be prepared | The `Deleting` message says the destroy waits for its credentials | `the destroy cannot start: it waits for its credentials: …` | > [!WARNING] > > **An object whose state reads and whose destroy can start is destroyed anyway** > > Even with the annotation set. The annotation would only skip a teardown that may well succeed. It takes effect once that destroy fails or turns out unable to start. The cases that are decided only when the destroy is about to start (the last three) are checked where it stops, not up front. What the abandon does: - Runs the same [cleanup]() as a successful destroy: the state Secrets, the lock Lease, the durable inputs, the plan key, the leases and the mirror ownership go. **Back up the state first** if you need it: the [stuck destroy runbook]() shows how. The backups are not deleted by cleanup but are garbage collected with the object. - Waits while a Job still holds a live run lease (five-second retries), like every other release. - Emits a `Warning` event `InfrastructureAbandoned`: `Removed finalizer without a destroy (): captf.io/abandon-infrastructure names this object's uid. Whatever the module created keeps running and is no longer managed; the state backups go with the object`. - Leaves the infrastructure running and untracked. Clean it up through the cloud provider. > [!CAUTION] > > **Abandoning leaves the infrastructure running and untracked** > > Whatever the module created keeps running and is no longer managed, and the state backups go with the object. Back up the state first if you need it. ```sh uid=$(kubectl get -n -o jsonpath='{.metadata.uid}') kubectl annotate -n captf.io/abandon-infrastructure="$uid" ``` > [!NOTE] > > **See also** > > - [Unreadable State](). > - [Stuck Destroy: abandon instead](). > - [Other manual actions](). # Stuck Destroy CAPTF has no skip-destroy annotation, only an abandon one (see [Abandon instead](<#abandon-instead>)): if a `destroy` Job keeps failing, the object stays with `ApplyJobSucceeded=False`/`DestroyFailed` (so `Ready=False`) and the controller retries with backoff forever. This applies the same way to a `TerraformCluster`, a `TerraformMachine` and a `TerraformMachinePool`. This page walks through recovering: back up what the destroy would remove, clean up the cloud resources another way, then remove the object’s finalizer by hand. > [!CAUTION] > > **Removing the finalizer deletes the only record of your cloud resources** > > Removing the finalizer garbage-collects the state Secrets, the state backups and the durable inputs Secret through their owner references, which are **the only record of the live cloud resources**. Back them up, or un-own them, before you do that. If the destroy is not failing but never starts because the state is missing or unreadable, deletion is held on purpose; see [deleting while state is unreadable]() before removing anything by hand. > [!WARNING] > > **A stuck control-plane destroy blocks the control plane** > > A stuck destroy on a control-plane machine also blocks KubeadmControlPlane/RKE2ControlPlane remediation, scale and upgrade until it is resolved. ## Abandon instead `captf.io/abandon-infrastructure` releases a deleting `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` whose destroy cannot run. It covers four cases: - The deletion is held because the state is missing or unreadable (`StateReadable` is `False`). - The last destroy Job failed (`status.lastRun` is a destroy and `ApplyJobSucceeded` is `False`). - The destroy cannot be rendered because the durable inputs Secret is gone (`ApplyJobSucceeded` is `False`/`DestroyFailed`). - The destroy cannot start because the identity no longer allows the namespace or was deleted (`IdentityNotAllowed`), or the runner credentials cannot be prepared (the `Deleting` condition says the destroy waits for its credentials). Set it to the object’s `metadata.uid`; any other value is ignored. The controller removes the finalizer without a destroy and records an `InfrastructureAbandoned` `Warning` event naming the cause. Whatever the module created keeps running, untracked, and the state backups are deleted with the object, so back up what you need and clean the cloud resources up yourself (steps 2 and 4 below). The controller checks the annotation before restore and destroy, so it also works while a restore is pending. An object whose state reads and whose destroy can start is destroyed as usual, even with the annotation set: it takes effect once that destroy fails or turns out unable to start. For the commands, see [deleting while state is unreadable](). ## Terminating namespaces A deleting object needs credentials only when it runs a Job. With the namespace terminating, an object that never applied finishes at once, and a held one can be abandoned. A destroy or restore waiting on credentials shows the `Deleting` condition message `The Job waits for its credentials: …` and retries. > [!NOTE] > > **Before you begin** > > - `get`, `patch` and `delete` access to Secrets and to the stuck object, in its namespace, on the management cluster. > - Replace ``, `` and `` below with the object’s namespace, name and Kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`; the lowercase singular works with `kubectl`). Run every command against the management cluster. ## 1\. Find the state suffix State Secret and backup Secret names are keyed by a suffix the controller derives from the object’s namespace, kind and name. Read it back from the object’s own status rather than recomputing it: ```sh kubectl get -n -o jsonpath='{.status.stateSecretSuffix}' ``` Save the output as `` for the commands below. ## 2\. Back up the state and inputs Secrets ```sh kubectl get secret -n -l tfstate=true,tfstateSecretSuffix= -o yaml > backup-state.yaml kubectl get secret -n -l captf.io/state-backup=true,captf.io/state-backup-suffix= -o yaml > backup-state-backups.yaml kubectl get secret -n captf-inputs-- -o yaml > backup-inputs.yaml ``` `` is `c` for a TerraformCluster, `m` for a TerraformMachine, `mp` for a TerraformMachinePool. The first selector matches the base state Secret and every `-part-N` chunk of a large Terraform state in one call; the second matches every backup Secret (and its chunks) taken of this object — see [Terraform State]() and the [state restore runbook](). Both kinds of Secret are owned by the object and are garbage-collected along with it once its finalizer is removed. `backup-inputs.yaml`’s `captf.io/image-digest` annotation (see [Annotations, labels and finalizers]()) names the exact image that ran the last successful apply — you will need it in step 4. ## 3\. Un-own the Secrets, if you want them to survive finalizer removal `kubectl patch` takes names, not a label selector, so patch each Secret by name: ```sh for s in $(kubectl get secret -n -l tfstate=true,tfstateSecretSuffix= -o name); do kubectl patch -n "$s" --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]' done for s in $(kubectl get secret -n -l captf.io/state-backup=true,captf.io/state-backup-suffix= -o name); do kubectl patch -n "$s" --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]' done kubectl patch secret -n captf-inputs-- \ --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]' ``` Skip this if you would rather rely on step 2’s `backup-*.yaml` files: with those in hand you don’t need the live Secrets to survive finalizer removal. ## 4\. Clean up the cloud resources Either clean up out of band (the cloud console, or the module’s own tooling), or run the pinned image’s `terraform`/`tofu` binary by hand against the backed-up state and inputs. The image’s own binary lives at `/captf/runtime` ([Image Contract]()); the Job’s own use of it, including its exact `-backend-config` flags, is in [Job Environment](). ```sh docker run --rm --entrypoint /captf/runtime \ -v "$PWD/root:/captf/work/root" -w /captf/work/root \ @ init -input=false -no-color \ -backend-config=secret_suffix= -backend-config=namespace= \ -backend-config=in_cluster_config=true -backend-config=labels= docker run --rm --entrypoint /captf/runtime \ -v "$PWD/root:/captf/work/root" -w /captf/work/root \ @ destroy -auto-approve -input=false -no-color \ -var-file=terraform.tfvars.json ``` - `@` is the repository from `backup-inputs.yaml`’s `captf.io/image` annotation (drop any `:tag`) plus `@` and the digest from its `captf.io/image-digest` annotation. - `` must be the same HCL object the Job itself would pass, or `init` reads an empty state: the backend selects state by this whole map. Copy it from `backup-state.yaml`’s base Secret `.metadata.labels`, dropping `tfstate`, `tfstateSecretSuffix` and `tfstateWorkspace` (the backend sets those itself), as `{"key"="value",...}`; see [Annotations, labels and finalizers]() for what each remaining key means. - `./root` needs `main.tf.json` and `terraform.tfvars.json` from `backup-inputs.yaml`’s data (the durable inputs Secret’s rendered root module and tfvars), and the identity’s credentials as environment variables (`-e =` for each key of the mirrored credentials Secret, or a file for a file-based one — see [Identities and credentials]()). - `-backend-config=in_cluster_config=true` reads the pod’s own ServiceAccount token, so this only works run from inside the cluster (for example, a debug pod using the `captf-runner` ServiceAccount). Running it from a workstation needs a kubeconfig-based backend configuration instead. - Restoring `backup-state.yaml` first (recreate the Secrets it holds) and letting the `kubernetes` backend read that live state also works, and skips reconstructing the backend configuration by hand. ## 5\. Remove the finalizer > [!CAUTION] > > **Stripping the finalizer skips the controller** > > No destroy, no cleanup and no event run. Finish steps 2 and 4 first: the cloud resources the object created keep running, untracked, and the state backups go with the object. ```sh kubectl patch -n --type=json \ -p '[{"op":"remove","path":"/metadata/finalizers"}]' ``` Use the lowercase, plural CRD resource name (for example, `terraformmachines`) if your `kubectl` version needs it instead of the Kind. This removes the whole `metadata.finalizers` array: safe for a `Terraform*` object, which carries only CAPTF’s own finalizer, but check `.metadata.finalizers` first if something else may have added one, and remove that entry by index instead. ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n > kubectl get secret -n -l tfstate=true,tfstateSecretSuffix= > ``` > > The first command reports `NotFound` once the object is gone. The second shows nothing unless you un-owned the Secrets in step 3, in which case they are exactly what you chose to keep. ## The durable inputs Secret is missing If the object reports `ApplyJobSucceeded=False`/`DestroyFailed` with the message “The durable inputs Secret is missing, so destroy cannot be rendered; see [https://captf.io/docs/operator-guide/runbooks/stuck-destroy.html]()”, the durable inputs Secret (`captf-inputs--`: the rendered root module, tfvars and pinned image from the last successful apply — see [Job Inputs]()) was deleted or never written. `TerraformMachine` is the persistent case: it is immutable and never falls back to re-rendering current inputs for a destroy, since the Machine and its bootstrap Secret a rebuild would need are usually already gone by the time destroy runs, so once its durable Secret is gone the condition never clears on its own. `TerraformCluster` and `TerraformMachinePool` (both mutable) fall back to building current inputs instead; they show the same message only while that build is gated (for example, waiting on a dependency that is itself being deleted), and it usually clears once the gate does. The controller retries forever either way; it never invents inputs to destroy with. To recover: - **If you have a backup** (`backup-inputs.yaml` from a previous run of step 2 above, or any earlier copy of the durable inputs Secret), recreate it with `kubectl apply -f backup-inputs.yaml` after removing `metadata.uid` and `metadata.resourceVersion` from the YAML: reapplying the exact object restores its owner reference, labels and annotations, including the pinned `captf.io/image` and `captf.io/image-digest`. The next reconcile reads it and starts the destroy Job. - **If you have no backup**, the controller cannot destroy the object’s resources. Back up the state (step 2), then clean up out of band (step 4): without the rendered `main.tf.json` and `terraform.tfvars.json`, running the module by hand needs reconstructing them, but the state lists every resource the module created. Instead of removing the finalizer by hand, [abandon](<#abandon-instead>) releases the object without a destroy and records the cause. `status.source.image` and `status.source.imageDigest` still name the image the last Job ran. Then remove the finalizer (step 5). **Removing the finalizer without cleaning up the cloud resources first abandons them**: they stay in the cloud, unmanaged, with nothing in Kubernetes recording that they ever existed. > [!NOTE] > > **See also** > > - [Terraform State]() — how state, its backups and their owner references work. > - [State restore runbook]() — restoring a state instead of destroying it. > - [Job Inputs]() — the durable inputs Secret. # Stripping a Finalizer by Hand Removing the finalizer yourself (`kubectl patch ... remove /metadata/finalizers`) skips the controller entirely: no destroy, no cleanup, no event. Sometimes that is the last resort, for example after a destroy that can never succeed and a manual cloud cleanup. Prefer the [abandon annotation](), which releases the same cases but runs the cleanup and records why. This page lists what a bare strip causes, so you can decide what to preserve first. The commands are in the [stuck destroy runbook](). > [!CAUTION] > > **A bare strip leaves your cloud resources running and untracked** > > The state was the only record of them, and the state backups are collected with the object. Back up and clean up first; see “Before you strip” below. ## What happens when it is gone The object is deleted as soon as the finalizer goes. Kubernetes then garbage-collects everything the object owns, and the controller can no longer act for it. | What | Result | | --- | --- | | The cloud resources | **Keep running, untracked.** The state was the only record of them | | The state Secrets | Collected if owned (owner reference). A chunk written since the last reconcile, or any chunk of an object paused since, has only the backend labels and is **left behind** | | The state backups | Collected: they are owned by the object. The newest copy of a lost state goes with them | | The durable inputs Secret | Collected, with the pinned image and the rendered module needed to destroy by hand | | The plan key Secret, the Jobs and their pods | Collected | | The state lock Lease, the run lease, the cluster write lease | **Left behind.** They carry no owner reference | | The credential mirror | Collected once every owner is gone; the mirror’s owner list is not updated | | The runner ServiceAccount and RoleBinding | Stay until the [sweep]() finds the namespace empty | Two of those bite later: - **Untracked state.** A leftover state Secret has no owner and nothing that lists it. The state suffix is a hash of the namespace, kind and name, never the UID, so a new object with the same three derives the same suffix and reads that old state as its own. - **Leftover leases.** A run or cluster write lease whose holder Job is gone is taken over by the next Job after the one-minute grace. The sweep deletes them once the namespace holds no `Terraform*` object. ## Before you strip 1. **Back up** the state Secrets, the backups and the durable inputs Secret, or remove their owner references so they survive: stuck destroy runbook, steps 2 and 3. 2. **Clean up the cloud resources.** Run the pinned image’s `destroy` against the backed-up state, or use the cloud console. 3. **Check `metadata.finalizers`.** Something else may have added one. Remove only CAPTF’s, by index. 4. **For a `TerraformCluster`, delete its machines first.** The destroy’s dependents wait is skipped when you strip, and a machine left behind loses its cluster. > [!NOTE] > > **See also** > > - [Held deletions](). > - [Cleanup and garbage collection](). > - [Stuck Destroy](). # Runbook: identity, credentials or RBAC not ready Before a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` can run a Job, the controller must resolve its [identity](), mirror that identity’s credentials into the object’s namespace, and make sure a runner `ServiceAccount` exists and is bound to the runner `ClusterRole`. Three conditions cover those steps: `IdentityAllowed`, `CredentialsMirrored` and `RunnerRBACReady`. While any of them is not `True`, no Job starts. This runbook covers every `False` and `Unknown` reason they carry, and how to fix each. Replace ``, `` and `` below with the object’s namespace, kind (`terraformcluster`, `terraformmachine` or `terraformmachinepool`) and name. ## Find the reason ```sh kubectl get -n -o jsonpath='{range .status.conditions[?(@.type=="IdentityAllowed" || @.type=="CredentialsMirrored" || @.type=="RunnerRBACReady")]}{.type}={.status}/{.reason}: {.message}{"\n"}{end}' ``` ## IdentityAllowed ### IdentityNotFound `False`. Either the object (and, for a machine or pool, its `TerraformCluster`’s `spec.defaults.identityRef`) sets no `identityRef` at all, or `identityRef.name` names a `TerraformClusterIdentity` that does not exist. Fix: set `identityRef.name` (on the object, or on the cluster’s `spec.defaults` for a machine or pool that should inherit it) to an existing `TerraformClusterIdentity`, or create the one it already names. See [Reference it](). ### NamespaceNotAllowed `False`. The named identity exists, but its `spec.allowedNamespaces` does not include this object’s namespace. Fix: add the namespace to the identity’s `allowedNamespaces` (its `list`, or a `selector` matching one of the namespace’s labels). See [Create the identity]() for the field’s exact semantics, including the empty-list-versus-empty- selector distinction. ### SecretNotFound `False`. The identity exists and allows this namespace, but its `spec.secretRef` Secret does not exist (or was deleted after the identity was created and admission’s `SubjectAccessReview` check passed). Fix: create the credentials Secret at the namespace and name `spec.secretRef` names, or point `spec.secretRef` at one that exists. See [Create the credentials Secret](). ### IdentityCheckFailed `Unknown`. A transient error while checking the identity or its Secret — reading the `TerraformClusterIdentity`, evaluating a `selector` against the namespace’s labels, or reading the credentials Secret all failed for a reason other than not-found (an API server error, for example). Unlike the three `False` reasons above, this is not a configuration problem: the reconcile itself returns an error and retries with backoff, so it usually clears on its own. If it persists, read the manager’s logs for the underlying error. ## CredentialsMirrored ### MirrorPending `Unknown`. Set whenever `IdentityAllowed` is `False` — no identity resolved yet, the namespace is not (or no longer) allowed, or the credentials Secret is missing — and before the first mirror is ever written. Not an error condition in itself: fix the `IdentityAllowed` reason above, and `CredentialsMirrored` follows it. ### MirrorFailed `False`. Either the mirror Secret `captf-creds-` could not be created or updated (an API error, named in the message), or a Secret with that exact name already exists in the namespace but is not a mirror of this identity: it lacks the `captf.io/mirrored` label, or its `captf.io/identity` annotation names a different identity. > [!WARNING] > > **A name conflict does not clear on its own** > > The controller never overwrites a Secret it does not recognize as its own mirror. Fix the conflict case by renaming or removing whatever created the conflicting Secret — most often a Secret created by hand or by another tool using the same name CAPTF would mirror to. See [How credentials reach a Job]() for the exact mirror name (`captf-creds-`, or a truncated hash form for a very long identity name) and what the mirror carries. For any other `MirrorFailed` message, it names the underlying API error; retry after fixing that (for example, a namespace quota or a webhook rejecting the write). ## RunnerRBACReady ### RBACFailed `False`. Creating or updating the runner `ServiceAccount` or the `captf-runner` `RoleBinding` failed. The message names the error. One specific cause: a `RoleBinding` named `captf-runner` already exists in the namespace without `captf.io/managed=true` — the controller never modifies a `RoleBinding` it does not own, since a binding it did not create could carry subjects or a `RoleRef` from something else. Fix that case by renaming or removing the conflicting `RoleBinding`; CAPTF then creates its own. For any other message, it is an API error (permissions, quota, a webhook); the manager’s own RBAC to manage these objects is set up as part of [installation](). See [RBAC]() for what the controller creates and why. ### ServiceAccountNotOptedIn `False`. The object’s effective `spec.jobs.serviceAccountName` names a `ServiceAccount` other than the default `captf-runner`, and that `ServiceAccount` either does not exist or does not carry `captf.io/runner=true`. CAPTF treats that label as the namespace’s consent to bind the `ServiceAccount` to the runner `ClusterRole`; it is never inferred. Fix: label the `ServiceAccount` (`kubectl label serviceaccount -n captf.io/runner=true`), create it if it does not exist, or remove the override from `spec.jobs.serviceAccountName` (and the cluster’s `spec.defaults.jobs.serviceAccountName`, if that is where it came from) to use the default `captf-runner` instead. See [Job tuning]() for `spec.jobs.serviceAccountName` and its default inheritance. ## Confirm it worked > [!TIP] > > ```sh > kubectl get -n -o jsonpath='{range .status.conditions[?(@.type=="IdentityAllowed" || @.type=="CredentialsMirrored" || @.type=="RunnerRBACReady")]}{.type}={.status}/{.reason}{"\n"}{end}' > ``` > > Expect `IdentityAllowed=True/IdentityAllowed`, `CredentialsMirrored=True/Mirrored` and `RunnerRBACReady=True/RBACReady`. The next reconcile after all three are `True` starts a Job if one is otherwise due. > [!NOTE] > > **See also** > > - [Identities and Credentials]() > - [RBAC]() > - [Job tuning]() > - [Conditions reference]() > - [Events reference]() # Webhook Unavailable CAPTF validates every `TerraformCluster`, `TerraformClusterIdentity`, `TerraformClusterTemplate`, `TerraformMachine`, `TerraformMachineTemplate`, `TerraformMachinePool` and `TerraformMachinePoolTemplate` write with an admission webhook, and every one of those webhook rules has `failurePolicy: Fail`: if the webhook cannot be reached, the write is refused rather than let through unchecked. This page covers recognizing that, and getting the webhook serving again. > [!NOTE] > > **There is no dedicated alert** > > It shows up as errors on `kubectl apply` (or on anything else writing a `Terraform*` object) and, because the manager’s own writes go through the same path, often as [CAPTFReconcileErrors]() as well. > [!NOTE] > > **Before you begin** > > - `kubectl` access to the `captf-system` namespace (or wherever CAPTF is installed): its Deployment, Service, Endpoints and the `cert-manager` `Certificate` and `Secret` it depends on. ## What is blocked while the webhook is down Every `CREATE` and `UPDATE` of the seven kinds above is blocked: nothing new can be created, and no existing one can be changed — including by Cluster API’s own controllers. In practice that reaches further than a person running `kubectl apply`: - a `MachineDeployment` or `KubeadmControlPlane`/`RKE2ControlPlane` scaling up cannot create the `TerraformMachine`s for the new replicas; - `clusterctl move` fails, since it creates and updates `Terraform*` objects on the target cluster; - the manager’s own reconciles that write to a `Terraform*` object’s metadata or spec — adding the finalizer to a new object, removing it once deletion finishes, or writing back a resolved `providerID` or `controlPlaneEndpoint` — fail the same way, and surface as [CAPTFReconcileErrors](). - deleting a `TerraformClusterIdentity` or most `TerraformMachine`s is also blocked: those two kinds validate `DELETE` as well as `CREATE`/`UPDATE`. The identity’s check normally refuses a delete while it is still in use or its credentials are still mirrored somewhere; the machine’s normally redirects a direct delete through its owner `Machine` instead, to go through drain. Either check can also be the thing *allowing* a delete that would otherwise be refused (an identity no longer in use, a machine whose owner `Machine` is already gone), so while the webhook is down those deletes are blocked outright rather than falling back to permissive. What keeps working: everything that only patches an object’s `status` subresource — a Job finishing, `status.lastRun`, conditions, drift and health sampling — since none of the webhook rules cover the `status` subresource. Existing, already-running Jobs finish normally; only changes to an object’s metadata or spec are affected. ## 1\. Confirm the webhook, not something else, is the cause `kubectl apply` against a `Terraform*` object returns an error naming the webhook by name (`captf-validating-webhook-configuration`). Map the error text to a cause: | Error text | Likely cause | | --- | --- | | `... failed calling webhook ...: no endpoints available for service "captf-webhook-service"` | No manager pod is `Ready`; go to step 2. | | `... failed calling webhook ...: context deadline exceeded` or `connection refused` | The Service or pod is reachable but not serving on the expected port, or a `NetworkPolicy` blocks it; go to step 3. | | `... x509: certificate signed by unknown authority` | The webhook’s `caBundle` is empty or stale; go to step 4. | ## 2\. Check the manager pod is Ready ```sh kubectl get pods -n captf-system -l control-plane=controller-manager kubectl get endpoints -n captf-system captf-webhook-service ``` The manager’s readiness probe includes its webhook server: a pod that has not finished starting the webhook server, or whose serving certificate failed to load, reports `NotReady` and drops out of the `captf-webhook-service` `Endpoints` — which is exactly why `kubectl get endpoints` shows nothing while this is the cause. Read the pod’s own logs and events for why it is not ready (a crash, an image pull failure, or the certificate problem in step 4). ## 3\. Check the Service and NetworkPolicy ```sh kubectl get svc -n captf-system captf-webhook-service -o yaml kubectl get networkpolicy -n captf-system ``` The `Service` forwards port 443 to the manager container’s `:9443`; a `Ready` pod with a populated `Endpoints` list but a webhook that still cannot be reached from the API server usually means a `NetworkPolicy` (none is installed by default; see [Installation]()) blocking traffic to that port, or the Service’s `selector` no longer matching the pod’s labels after a manual edit. ## 4\. Check the certificate ```sh kubectl get certificate -n captf-system captf-serving-cert kubectl describe certificate -n captf-system captf-serving-cert kubectl get secret -n captf-system captf-webhook-service-cert ``` `cert-manager` issues the webhook’s serving certificate as the `Secret` `captf-webhook-service-cert`, mounted into the manager pod, from the `Certificate` `captf-serving-cert`. The `ValidatingWebhookConfiguration`’s `caBundle` is kept current by `cert-manager`’s CA injector, driven by the `cert-manager.io/inject-ca-from: captf-system/captf-serving-cert` annotation on `captf-validating-webhook-configuration` itself — check that annotation is still present (a `kubectl apply` of a stripped-down copy of the manifest can remove it) and that the `Certificate` reports `Ready`. `cert-manager` not running at all, or its CRDs missing, leaves the `Certificate` object present but never issued. See [Installation]() for how these names and the `cert-manager` dependency fit together, and confirm `cert-manager` itself is healthy in its own namespace if the `Certificate` never becomes `Ready`. ## Confirm it worked > [!TIP] > > ```sh > kubectl apply --dry-run=server -f - <<'EOF' > apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 > kind: TerraformClusterIdentity > metadata: > name: webhook-check > spec: > secretRef: > name: webhook-check > namespace: default > EOF > ``` > > A server-side dry run still goes through admission. It succeeding (or failing with a validation error about the object’s own content, not a webhook connectivity error) confirms the webhook is reachable again; a blocked create or scale-up elsewhere in the cluster starts progressing on its own once it does. > [!NOTE] > > **See also** > > - [Reconcile Errors]() — the manager’s own writes failing for the same reason. > - [Installation]() — what `clusterctl init` installs, and the exact resource names used above. # How It Works # The Reconcile Lifecycle `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` share one reconcile flow. This page walks through a single pass of it: what runs before any decision, how the next operation is chosen, how a failed operation backs off, how two objects of the same cluster stay off each other’s feet, and what each kind does once the shared flow is done. It is for anyone who wants to know why CAPTF started (or did not start) a Job, or why a delete is taking a while. ## One pass, one status patch Every reconcile writes the object’s status once, in a single patch at the end of the pass. Two things are written earlier, directly, because a crash between them and the deferred patch must not leave a gap: - the [clusterctl move block](<#the-clusterctl-move-block>), set just before a Job is created; - the preamble’s own writes, described below, which can stop the pass before there is anything else to decide. A reconcile that reaches the end of the pass runs through, in order: 1. **Externally managed.** An object carrying CAPI’s externally-managed annotation is left alone entirely; nothing below runs. 2. **Owner lookup.** Without an owner reference of the expected kind yet, the object waits (`DependenciesReady=Unknown`) unless it is being deleted with no state and no running Job, in which case the finalizer is dropped at once: there is nothing to destroy. An owner reference whose target is gone does not stop deletion; a destroy still runs from the durable inputs. An owner gate such as the owning `Cluster`’s infrastructure reference not naming a `TerraformCluster` sets `DependenciesReady` and stops the pass, unless the object is being deleted, in which case deletion proceeds anyway. The lookup also requires the owner to reference this object back (its own `infrastructureRef`, and, when set, a matching ownerRef UID and cluster-name label); an owner that fails this check is treated the same as one that is gone: `DependenciesReady=False`/`OwnerMismatch`, no Job, and deletion still proceeds. See [Owner references are checked against the owner]() for the full rule and why it exists. 3. **Finalizer.** Added if missing; the pass stops for this reconcile so the write is visible before anything else happens. 4. **Pause.** A paused object, or one whose `Cluster` is paused, runs only the paused branch described below. See [Conditions]() for every condition this flow sets and what each reason means, and [TerraformClusterIdentity]() for the identity and credential-mirror step that follows the preamble. ## The paused branch A paused object still processes its finished Jobs (bookkeeping, below) and deletes a Job that can never start, but starts nothing new. Once no Job is active, the clusterctl move block is cleared. This is what `clusterctl move` waits for after it pauses an object; see [the move runbook](). ## Bookkeeping Unpaused, the reconcile reads the object’s Jobs and processes every one that finished since the last pass it was seen: it records the run’s steps, error and drift summary in `status.lastRun`, sets the `ApplyJobSucceeded` and `DriftJobSucceeded` conditions, pins the image digest of a newly succeeded apply, prunes history beyond the configured limits, and finds whether a Job is active now. A stale state lock left by a Job whose pod is gone is flagged for the next Job to force-unlock; a lock held by something else (a workstation, another tool) is reported instead. A finished Job’s per-run inputs Secret is deleted once, which is also the point at which its run lease (and, for a `TerraformCluster`, its cluster write lease) is given back. > [!NOTE] > > **A Job is bookkept only after the status patch succeeds** > > What a finished Job’s pod reported is only durable once the status patch at the end of the pass has succeeded: the reconcile marks a Job “bookkept” only after that patch, so a failed patch leaves the Job to be read again next time instead of losing its result. While a Job is active, state is not read and no new operation is decided: the reconcile deletes the Job if it looks stuck (its per-run Secret is missing and no pod ever started, checked only once the Job is at least a minute old) and otherwise waits, holding the clusterctl move block. ## Choosing the next operation Once no Job is active, the reconcile reads state, builds the current inputs (for a mutable kind, or an immutable one not yet provisioned), and looks for a requested state restore. It then picks the next operation in this order: 1. A restore named by the `captf.io/restore-state` annotation, unless the object is being deleted: it is what was asked for, and checking or applying against the state about to be replaced would be wasted or worse. 2. Being deleted: `destroy`, or drop the finalizer at once if there is no state to destroy. 3. No state yet: `apply` (first apply). 4. State exists but carries no inputs hash: `apply`. 5. A mutable kind whose current inputs hash differs from the one state carries: `apply`. 6. A mutable kind whose newest finished apply failed (and was not blocked, was not stopped by a plan mismatch, and was not a drift remediation): `apply` again, even if the current inputs already match state, because the failed run may have changed reality without recording it. 7. A mutable kind with drift found and drift action `Remediate`, while the object has not yet reached its remediation-failure cap since the last successful drift check: `apply`. Drift detection and remediation are [Drift and health]()’s topic. Whichever of those an apply reason names but the first (no state yet, with nothing to break), an object under `applyPolicy` `Manual` does not apply directly: it waits for a plan and its approval first, or backs off behind the destructive-plan guard if neither policy applies. Both are [Plan approval]()’s topic; this page only fixes where that wait sits in the order above. When none of those apply reasons matched, the reconcile falls through to its refresh, membership and drift schedule, checked in this order and stopping at the first one due: 1. A successful apply not yet followed by a `refresh`, for kinds that ask for one (a machine or machine pool, so its addresses and health reach status without waiting for the next tick). 2. A `TerraformMachinePool` whose membership is still converging: `refresh` every 30 seconds, plus jitter, fixed rather than backed off. 3. Otherwise, while health reads pending: `refresh` after a delay that starts at 30 seconds and doubles (30s, 1m, 2m, 4m) up to a 5-minute ceiling, plus jitter. 4. A machine pool’s membership refresh interval elapsed: `refresh`. 5. The health-check interval elapsed: `refresh`. 6. The drift-check interval elapsed: `drift`. 7. Otherwise, requeue at whichever of the above is soonest. Every interval above adds a small deterministic jitter derived from the object’s UID (up to a tenth of the interval), so objects created together, such as a `MachineDeployment`’s machines or everything moved by one `clusterctl move`, do not all check in lockstep. Configuring these intervals is [Drift and health]()’s and [Configuring drift]()’s topic; machine pool membership convergence is [Machine pools]()’s. ## Retry backoff A failed Job’s operation is not retried immediately: the delay after `n` consecutive failures of that operation is one minute, doubling each time, capped at ten minutes. Once `n` reaches the failed-Jobs history limit (three by default), the delay jumps straight to the ten-minute cap instead of continuing to double, so the sequence with the default limit is one minute, two minutes, then ten minutes and no shorter. A Job killed by its `activeDeadlineSeconds` is a failure and does count. A Job stopped from outside (its pod was interrupted, or an apply was blocked before a destructive plan, or an approved apply’s plan changed since approval, or the run never started because a lease was held) does not count toward this backoff at all: those wait on a lease, an approval, or the object’s own next reconcile instead. > [!WARNING] > > **A failed state restore is never retried for the same serial** > > A failed state restore is a special case: it is never retried for the same backup serial, only for a new one or after the failed Job is deleted. ## Run leases and the cluster operation gate Before a Job is created, the reconcile takes the object’s own run lease, so two managers (or two reconciles racing after a crash) never start two Jobs for the same object at once. If another Job already holds it, the operation waits instead: `ApplyJobSucceeded`, `DriftJobSucceeded` or `RestoreJobSucceeded` (matching the operation) goes `Unknown` with a message naming the Job to wait for. With the manager’s `--cluster-operation-gate` flag (default enabled; see [Manager configuration]()), an `apply`, `destroy` or `restore` Job of an object that names a cluster additionally crosses a second gate meant to keep a cluster’s own apply or destroy from running at the same time as its machines’ and machine pools’: - A `TerraformCluster` takes the cluster’s write lease, then checks whether any machine or machine pool of the cluster is currently applying or destroying. If so, it keeps both leases and waits for them to finish before it starts: new machine and machine pool operations wait behind it meanwhile, so a stream of machine creates cannot starve the cluster’s own operation. - A `TerraformMachine` or `TerraformMachinePool` checks whether the cluster’s write lease is currently held. If it is, the machine or pool gives back the run lease it just took and waits for the cluster’s operation to finish first. ## Deletion order A `TerraformCluster` being deleted waits for every `TerraformMachine` and `TerraformMachinePool` carrying its cluster name in the namespace to be gone before its own destroy runs; machines and machine pools carry no such wait of their own. An owner reference whose target has already been removed does not block a destroy: deletion still runs from the durable inputs. Once a destroy succeeds, or deletion finds no state to destroy at all, the reconcile deletes the object’s state and durable inputs, releases its leases, drops it from its credential mirror’s owners, and removes the finalizer. See [Secrets]() for what those Secrets are and [RBAC]() for the namespace RBAC sweep that follows a finalizer’s removal. ## The clusterctl move block Clusterctl’s block-move annotation is set on the object just before a Job is created and stays set for as long as a Job is active; it is cleared, paused or not, once no Job is active. `clusterctl move` pauses every object first and then waits for the annotation to clear, which is why the paused branch above still runs bookkeeping and the stuck-Job check: a Job that can never start would otherwise hold the block indefinitely. ## Job names and history A Job’s name is deterministic: `captf----a-`, where `` is `c`, `m` or `mp`, `` is one more than the highest attempt number retained for that operation, and `` is six hex characters derived from the inputs hash, the operation, the attempt and, for `refresh` and `drift`, a tick that changes each time the operation last succeeded (so a retry finds the same Job instead of starting a second one); a drift-remediation `apply` and a `restore` carry a tick of their own too, so neither collides with an unrelated Job of the same inputs hash and attempt. A name that would exceed 57 characters replaces the object name with a hash instead. History beyond the configured successful- and failed-Jobs limits is pruned per operation, oldest first, while always keeping each operation’s newest success and, while no success is newer, its newest failure, since those are what retry backoff and drift remediation’s failure cap read. Tuning those limits, and the rest of a Job’s resources, deadlines and identity, is [Job tuning]()’s topic. ## Jobs per bring-up Bringing up one `TerraformCluster` with three control-plane and three worker `TerraformMachine`s runs eight Jobs when each machine’s apply outputs give a definite health reading: one cluster apply, one apply per machine (six), and one no-change re-apply once the cluster’s `control_plane_initialized` input flips (the cluster’s inputs hash changes, so the shared flow applies again even though nothing else about the cluster changed). A machine whose apply outputs read pending instead keeps the extra post-apply refresh described above; with all six machines reading pending that adds six more Jobs, back up to fourteen. Count Jobs over a bring-up with `sum(captf_jobs_total)` rather than listing Jobs: a finished Job is pruned once it falls outside the history limits above. ## After the shared flow Each kind does a little more once the shared flow has read state and decided: - A **`TerraformCluster`** decides its `control_plane_endpoint` input while the shared flow builds inputs, and writes an apply’s endpoint output back to `spec.controlPlaneEndpoint` while state is read: both are part of the shared flow, not a separate step. Each happens at most once. - A **`TerraformMachine`** keeps `cluster.x-k8s.io/remediate-machine` on its owning `Machine` in step with its instance health once the shared flow finishes, skipped for a paused or externally-managed machine. See [Machine health and remediation](). - A **`TerraformMachinePool`** writes its observed replica count back to its owning `MachinePool`’s `spec.replicas` when autoscaling is enabled, skipped for a paused, externally-managed or deleting pool. See [Machine pools](). ## One pass, visually ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A["Preamble: externally managed?
owner lookup, finalizer, pause"] A -->|paused| P["Paused branch:
bookkeeping, delete a stuck Job,
clear the move block"] A -->|not paused| B["Credentials: identity,
credential mirror, runner RBAC"] B --> C["Bookkeeping: finished Jobs,
digest pin, lock check"] C -->|a Job is active| D["Delete it if stuck,
else wait for it"] C -->|no Job is active| E["Clear the move block"] E -->|deleting, destroy succeeded| F["Cleanup: remove the finalizer"] E -->|deleting, machines or pools remain| G["Wait: deletion blocked"] E -->|otherwise| H["Read state, build inputs,
look for a requested restore"] H --> K["Decide the next operation"] K -->|start a Job| L["Take leases, start the Job"] K -->|drop the finalizer| F K -->|nothing to start| M["Requeue at the decided delay"] D --> N["Patch status once"] L --> N M --> N P --> N F --> N G --> N ``` > [!NOTE] > > **See also** > > - [Conditions]() — every condition this flow sets. > - [Events]() — the events it emits. > - [Drift and health]() — how drift checks and health work. > - [The state backend]() — reading and adopting state, and backups. > - [Plan approval]() — the Manual apply policy and the destructive-plan guard. > - [Jobs, Retries and Concurrency]() — the Job machinery in depth: names, retries, leases, cache lag and failover. > - [Deletion and Teardown]() — the delete path in depth. # Terraform State CAPTF stores every object’s Terraform or OpenTofu state in Kubernetes Secrets, alongside the object, and reads it back for outputs, drift and health. This page explains where that state lives, how CAPTF backs it up, and what happens to it when an object is deleted. For an inventory of every Secret CAPTF reads or writes, including state, see [Secrets](). ## The backend Every `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` gets its own Terraform or OpenTofu run, backed by the `kubernetes` backend configured to write into the object’s own namespace, in the `default` workspace (the only one CAPTF ever uses). The backend’s `secret_suffix` is derived from the object’s namespace, kind and name, which survive a `clusterctl move`, and never from its UID, so state follows the object across a move instead of being orphaned. ## Secret names and the suffix The suffix is the first 16 hex characters of `hex(sha256(//))`, a hyphen, and a short kind code: `c` for `TerraformCluster`, `m` for `TerraformMachine`, `mp` for `TerraformMachinePool`. It never ends in `-`, which the backend would otherwise try to parse as a chunk index. The base state Secret is named `tfstate-default-`. `status.stateSecretSuffix` records the suffix; it is informational only, since the controller derives it deterministically and never reads it back. The state lock is a `coordination.k8s.io/v1` Lease named `lock-tfstate-default-` in the same namespace. See [Locks](<#locks>) below. ## Chunking and size caps A Kubernetes Secret holds at most 1 MiB, and a large Terraform state can exceed that once compressed. Terraform splits an oversized state across additional Secrets named `tfstate-default--part-1`, `-part-2` and so on; OpenTofu does not chunk state at all and always writes a single Secret. CAPTF reads whichever shape is present: it lists every Secret carrying the backend’s own labels for the suffix, orders them by chunk index, and concatenates their payloads before decompressing. CAPTF caps what it is willing to read: at most 32 chunks and 64 MiB of decompressed state. Real state compresses 10-20x, so these limits are far beyond any plausible cluster or machine state; a state that exceeds them, or whose chunk set is incomplete, duplicated or names an unexpected Secret, is reported corrupt or inconsistent rather than partially read. [`CAPTFStateNearSecretLimit`]() warns before a state’s compressed size approaches the 1 MiB Secret limit. ## Locks Both runtimes hold the lock for the duration of a run and release it on a clean exit. A runner Job waits up to `lockTimeoutSeconds` (see [Job tuning]() for the field and its default) for a held lock before failing. Before starting a Job, the controller checks the lock Lease and reads its holder from the backend’s own lock info: a lock whose holder is one of the object’s own runner pods, and that pod either no longer exists or has already exited (a finished Job keeps its pod object until it is pruned), is stale, and the controller has the next Job force-unlock it automatically. A lock whose holder is unknown, or is not one of the object’s own runner pods, is left alone and reported as `StateReadable=False/StateLocked`. See the [stale state lock runbook]() to force-unlock one by hand. ## What CAPTF reads from state CAPTF parses only the fields it needs from the Terraform state v4 file: the serial, lineage, Terraform/OpenTofu version, the root module’s outputs, and a count of managed resources (a resource with `count` or `for_each` counts once). It also tracks the compressed size of the concatenated chunks and, from an annotation on the base Secret, the inputs hash of the last successful apply. Resource instance attributes are never parsed. Outputs may be sensitive and are never logged. After a successful apply or restore, the controller sets an owner reference to the object on every chunk Secret (so state moves with a `clusterctl move` and is garbage collected with the object) and records the applied inputs hash on the base Secret, since Terraform’s own chunk writes carry only the backend’s labels. A chunk that lacks the reference, or names an earlier UID of the object (after a restore), is owned again on the next reconcile with no Job running. ## State backups The manager keeps versioned copies of an object’s state so a lost, corrupted or wrongly overwritten state Secret can be recovered. Whenever the controller reads a state serial it has not observed before for the object, it copies the state Secrets’ data verbatim into a backup set named `captf-state-backup--` (with the same `-part-N` chunking as the source), owned by the Terraform\* object itself rather than by the state, so a backup survives the state Secret being deleted by hand and is garbage collected only when the object is. An unchanged serial is not backed up again. `--state-backups` (default 5; see [manager flags]()) sets how many backups per object the manager keeps; it prunes older ones in the same pass. `--state-backups=0` takes no new backups but leaves existing ones in place and restorable. `status.stateBackups` lists the newest backups (see [Common Fields]() for its fields). A state that cannot be parsed with the reader’s own limits is never backed up: encrypted, corrupt, inconsistent, or beyond the chunk and size caps above; nor is one the manager failed to copy, for example on a transient API error. Either way the manager logs why and counts it in [`captf_state_backups_total`]() with `result="skipped"`. A state that cannot be parsed is left that way and the reconcile continues; a failed copy instead retries on the next reconcile, since that serial is not recorded as observed. No backups are taken while an object is being deleted. ## Restore An annotation asks the controller to push a listed backup’s content back into the backend as a new state, through a restore Job that runs `state push -force`. See the [state restore runbook]() for the procedure, what a restore does and does not undo, and how to verify one. ## OpenTofu state encryption OpenTofu’s client-side state encryption wraps the state file in an envelope with no `version` field of its own. CAPTF detects that envelope and reports the state as encrypted (`StateReadable=False/StateEncrypted`) rather than misreading it as corrupt: reading an encrypted state’s outputs is not supported. ## State on deletion Neither backend deletes its own state: a `destroy` only empties the managed resources it recorded, and the `default` workspace cannot be deleted. The controller removes the state Secrets and the lock Lease itself, once it is safe to do so: after a destroy Job succeeds, or immediately on deletion of an object that was never applied and so has no state. That same cleanup also deletes the durable inputs Secret, the object’s run lease and, for a `TerraformCluster`, its cluster write lease. State backups are not deleted by this cleanup; they are owned by the object and are garbage collected when Kubernetes removes it after its finalizer is gone. Whether an object ever applied is recorded in the durable inputs Secret’s `captf.io/applied: "true"` marker, which survives `clusterctl move`; see [the applied marker](). If the finalizer is removed by hand before the state is cleaned up, see the [stuck destroy runbook](). > [!WARNING] > > **Deletion is held when an applied object’s state is missing or unreadable** > > Deletion is held, rather than proceeding, when the object ever applied and its state is missing or unreadable; see [deleting while state is unreadable]() for the two ways out. > [!NOTE] > > **See also** > > - [Secret Management](), the mechanics behind this page: labels, adoption, backups and the other Secrets. > - [Secrets]() for the full Secret inventory, including state and its backups. > - [State restore runbook](). > - [Stale state lock runbook](). > - [Stuck destroy runbook](). > - [Unreadable state runbook](). # Secret Management CAPTF keeps everything it needs between Jobs in Kubernetes Secrets and Leases: the Terraform state, the backups of it, the rendered inputs of each run, the cloud credentials, and a key for the plan approval. This chapter follows each of them from creation to deletion. It explains how they connect, how they reach the Terraform runtime inside a Job, and what an operator needs to provide, protect and back up. The chapter has these pages: - **Terraform State Secrets** --- Storage, naming, chunking, the integrity checks, the inputs hash and adoption. - **Backups and Restore** --- When a backup is taken, how many are kept, and how a restore runs. - **Credentials** --- An identity’s source Secret, the per-namespace mirror, rotation and revocation. - **Run Inputs and the Plan Key** --- The durable and per-run inputs Secrets and the key behind the plan hash. - **Inside the Job** --- Volumes, environment, RBAC, the runner’s steps and lock handling. - **Lifecycle Walkthroughs** --- Create, apply, drift, change, delete, `clusterctl move`, namespace deletion and lost state. - **Operator Files and Settings** --- The manifests, flags and backups you own. - **Security Considerations** --- The trust boundaries, stated plainly. This page is the inventory of every Secret and Lease, and how they relate. Existing pages cover parts of this ground from other angles, and are linked rather than repeated: [Terraform State](), [Job Inputs](), [Identities and Credentials](), [Plan Approval](), the [Security Model](), [Secrets]() (which also lists the Secrets CAPTF only reads, such as bootstrap data) and the [runbooks](). ## The objects at a glance `` is the state suffix: the first 16 hex characters of `sha256(//)`, a hyphen and `c`, `m` or `mp` (see [Terraform state]()). `` is `c`, `m` or `mp` for a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`. “The object” below is that `Terraform*` object. | Name pattern | Kind | Created by | Owner | Contents | Moves with `clusterctl move` | Deleted when | | --- | --- | --- | --- | --- | --- | --- | | `tfstate-default-` and `-part-N` | Secret | The Terraform or OpenTofu `kubernetes` backend, inside the Job | The object, as a non-blocking owner reference, set after the first successful apply. A chunk without one (written by a refresh or drift Job, or restored without references) or naming an earlier UID is owned again on the next reconcile | Gzipped state under the key `tfstate` | Yes: it carries the move label | Cleanup after a destroy, a delete with no state, or an abandon | | `lock-tfstate-default-` | Lease | The backend, when a run takes the lock | None | The lock holder’s information | No: a Lease is not discovered; the backend recreates it | Cleanup, with the state | | `captf-state-backup--` and `-part-N` | Secret | The manager, after it sees a new serial | The object (non-controller reference); re-owned after a restore | A verbatim copy of the state chunks | Yes: it carries the move label | With the object, by garbage collection; pruned beyond `--state-backups` | | `captf-inputs--` | Secret | The manager, when an apply Job starts | The object; re-owned after a restore | The rendered root module and tfvars, plus the image, digest, identity and applied marker | Yes, by following the owner reference | Cleanup | | `captf-run-` | Secret | The manager, right after it creates the Job | The Job | The same two rendered files | Not applicable: it lives only while the Job does | The first time the controller sees the Job finished | | `captf-plankey--` | Secret | The manager, before a plan or approved apply Job | The object; re-owned after a restore | 32 random bytes under the key `key` | Yes, by following the owner reference | Cleanup | | `captf-creds-` (the mirror) | Secret | The manager, on a reconcile of an object that uses the identity | Each object that uses it (non-controller references); a using object’s reference to its earlier UID is replaced | A copy of the identity’s source Secret data | Yes, and it is rewritten from the source on the target | When its last user is gone, or when the namespace stops being allowed | | The identity’s source Secret (any name) | Secret | The operator | None | Cloud credentials | No: copy it to the target yourself | By the operator | | `captf-run-` | Lease | The manager, before it starts a Job | None | The run lease: one Job at a time per object | No: a Lease is not discovered | When the Job finishes, at cleanup, and by the namespace sweep | | `captf-cluster-` | Lease | The manager, for a `TerraformCluster`’s apply, destroy or restore | None | The cluster write lease: machines and pools wait on it | No | When the Job finishes, at cleanup, and by the namespace sweep | The labels and annotations on each are in [Annotations, Labels and Finalizers](); this chapter names the ones that matter to the behavior it describes. Every Secret the manager creates carries `captf.io/managed=true`, which is also what its cache selects on (see [What the manager caches]()). ## How they connect ``` flowchart LR subgraph ops[Operator] SRC["Identity source Secret"] ID["TerraformClusterIdentity"] end subgraph mgr[Manager] MIR["captf-creds-identity
(mirror)"] DUR["captf-inputs-kind-name
(durable inputs)"] RUN["captf-run-job
(per-run inputs)"] KEY["captf-plankey-kind-name"] BAK["captf-state-backup-...
(backups)"] end subgraph job[Job pod] R["Runner"] TF["Terraform or OpenTofu"] end ST["tfstate-default-suffix
(state)"] LK["lock-tfstate-default-suffix
(Lease)"] ID --> SRC SRC -->|"copied on every reconcile"| MIR MIR -->|"envFrom and files"| R DUR -->|"copied at Job start"| RUN RUN -->|"/captf/config"| R KEY -->|"/captf/plan-key"| R R --> TF TF -->|"reads and writes"| ST TF -->|"holds"| LK ST -->|"new serial seen"| BAK BAK -.->|"restore Job"| ST ``` The flow, in words. An operator creates the identity’s source Secret and the `TerraformClusterIdentity` that names it. When an object runs, the manager mirrors the source into the object’s namespace and renders the object’s inputs into the durable Secret, then copies them into a per-run Secret that belongs to one Job. The Job’s runner receives the per-run Secret, the mirror and, for a plan, the plan key; it runs Terraform, which reads and writes the state Secrets and holds the lock Lease through the cluster API. The manager reads the state back, takes a backup when the serial is new, and records the outcome in status. Two rules explain most of what follows: - **State is the source of truth for what exists.** Status is rebuilt from state, which is why status is not restored by `clusterctl move`, and why a [marker on the durable inputs Secret]() is needed to remember that an object ever applied. - **Everything a Job reads is a copy.** The per-run Secret is written once and never rewritten, so a Job never sees inputs change under it, and the mirror is rewritten only by the manager, never by the Job. > [!NOTE] > > **See also** > > - [Secrets](), the inventory including the Secrets CAPTF only reads. > - [Terraform State]() and [Job Inputs](). > - [Security Model](). # Terraform State Secrets The state of every `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` lives in Secrets in the object’s own namespace, written by Terraform’s or OpenTofu’s `kubernetes` backend from inside the Job. This page covers how those Secrets are named, found, read and checked, and how the manager attaches them to the object. The higher-level description is in [Terraform State](); this page adds the mechanics. ## The backend The rendered root module declares an empty `terraform.backend.kubernetes` block. Everything the backend needs arrives on the command line of `init`: ```text init -backend-config=secret_suffix= \ -backend-config=namespace= \ -backend-config=in_cluster_config=true \ -backend-config=labels= ``` The backend therefore authenticates with the Job’s own ServiceAccount (see [Inside the Job]()), and the workspace is always `default`: the runner drops `TF_WORKSPACE` from the environment. ## Names and the suffix The suffix is the first 16 hex characters of `sha256(//)`, a hyphen and `c`, `m` or `mp`. It depends on names, never on the UID, so it survives `clusterctl move`, and it never ends in `-`, because the backend parses a chunk index from a trailing number. | Secret or Lease | Name | | --- | --- | | Base state Secret | `tfstate-default-` | | Further chunks | `tfstate-default--part-N` | | Lock Lease | `lock-tfstate-default-` | `status.stateSecretSuffix` records the suffix for you to read; the controller derives it again every time and never reads it back. ## Labels The `labels` backend setting makes the backend stamp a fixed set of labels on every chunk and on the lock Lease, next to its own `tfstate=true`, `tfstateSecretSuffix` and `tfstateWorkspace=default`: | Label | Value | | --- | --- | | `captf.infrastructure.cluster.x-k8s.io/owner-kind` | The object’s kind | | `captf.infrastructure.cluster.x-k8s.io/owner-name` | The object’s name; a 16-hex-character hash of it if it is over 63 characters | | `cluster.x-k8s.io/cluster-name` | The owning Cluster’s name, with the same hashing | | `captf.io/managed` | `true` | | `clusterctl.cluster.x-k8s.io/move` | Empty: the move marker | > [!WARNING] > > **The label map must never change for an existing object** > > The backend lists its chunks with a selector made of the whole map, so a changed map would stop matching the Secrets already there and the state would appear to vanish. ## Chunking and compression Each Secret holds the state gzip-compressed under the data key `tfstate`. Terraform splits a state that compresses past about 1 MiB into further Secrets; OpenTofu 1.12 writes a single Secret and does not chunk. The reader accepts either shape. ## How the manager reads state The reader (the manager’s own, not Terraform’s) works in this order, and each failure maps to a `StateReadable` reason: 1. It lists Secrets by the backend selector: `tfstate=true`, the suffix and the `default` workspace. State Secrets are read uncached, straight from the API server. 2. It orders the base Secret and the `-part-N` chunks. A gap, a duplicate, an unexpected name or a base Secret without the `tfstate` key is `StateInconsistent`. 3. It refuses more than 32 Secrets, or more than 64 MiB once decompressed: `StateCorrupt`. The decompression stops at the first gzip member, so a trailing chunk left behind when the state shrank is ignored rather than misread. 4. If the document carries an `encryption_version`, it is OpenTofu’s client state encryption, which CAPTF cannot read: `StateEncrypted`. This check comes first. 5. The state `version` must be 4. Any other value, or none, is reported as `StateCorrupt`. 6. It reads the serial, the lineage, the Terraform or OpenTofu version, the root outputs and the number of managed resources. Resource attributes are not parsed. No Secret at all is either “no state yet” (`StateNotFound`, Unknown) or `StateLost`, depending on whether the object ever applied; see [lost state](). Outputs come from the state, never from a Job’s result, and `status.observedStateSerial` records the serial they were read from. For what each reason means and how to recover, see the [unreadable state runbook](). ## The inputs hash The base Secret carries the annotation `captf.io/inputs-hash`, a value of the form `h2:` over the canonical rendered inputs of the apply that wrote the state. The controller compares it with the hash of the inputs it would render now: - A mutable object whose current hash differs is re-applied (`InputsChanged`). - A state with no hash at all is `StateWithoutInputsHash`: the controller cannot tell what produced it. Terraform’s own writes carry only the backend labels, so the annotation is added by the manager, in the step below. ## Adoption After a successful apply or restore, the manager adopts the state: - It adds an owner reference to the `Terraform*` object on every chunk. The reference is not a controller reference and leaves `blockOwnerDeletion` unset, so it never blocks the object’s deletion; it makes Kubernetes garbage-collect the chunks with the object, and makes `clusterctl move` follow the object to the target. - It sets `captf.io/inputs-hash` on the base Secret, with an optimistic lock so a concurrent write is not overwritten. Adoption runs only when the Job’s hash is set and differs from the state’s. A `-part-N` chunk that Terraform creates later, during a refresh, a drift check or a retry with the same hash, therefore has no owner reference when it is written. The next reconcile that finds no Job running owns it again, as it does a chunk restored without references or naming an earlier UID of the object, and emits `OwnerReferencesRepaired`. It does not do so while the object is paused, or while a Job holds the run lease. The cleanup after a destroy finds every chunk by the label selector, not by ownership. > [!NOTE] > > **See also** > > - [Terraform State]() for caps, locks and what is read. > - [Backups and restore](). > - [Size limits runbook](). > - [Stale state lock runbook](). # Backups and Restore The backend keeps only the latest state. To let you recover from a lost or damaged state Secret, the manager copies the state aside whenever it changes. This page covers when a backup is taken, how backups are named, labeled and pruned, and how a restore moves one back. The procedure is the [state restore runbook](); this page explains what happens underneath it. ## When a backup is taken The manager takes a backup when a reconcile reads a state serial it has not seen for the object, which happens after an apply, a refresh, a drift Job or a restore. It does not take one before an operation, so a backup is always a snapshot of something that already happened, never a safeguard in front of a change. No backup is taken when: - `--state-backups` is `0` (existing backups stay restorable); - the object is being deleted; - the state is unreadable (encrypted, corrupt, inconsistent or of an unsupported version); - the serial is below 1. A copy that fails on a transient API error is retried on the next reconcile, since that serial is not yet recorded as seen. Each outcome is counted in `captf_state_backups_total` (see [Metrics]()). ## What a backup looks like A backup is a set of Secrets, one per state chunk, copied verbatim: - **Name:** `captf-state-backup--`, with `-part-N` for further chunks. If a Secret for the same serial already exists with different content, as happens after a restore, the name gains `-`, the first eight characters of the digest. - **Labels:** the five backend labels (see [Terraform state Secrets]()) plus `captf.io/state-backup=true` and `captf.io/state-backup-suffix=`. It carries none of the `tfstate*` labels, so Terraform never mistakes a backup for live state. - **Annotations:** `captf.io/state-backup-serial`, `-lineage`, `-taken-at`, `-source-job` (when known), `-digest` (SHA-256 of the concatenated compressed payload), `-resources` (managed resource count), `-set` (the base name), `-chunk` (this chunk’s index) and `-chunks` (the count), and `captf.io/inputs-hash` on the first chunk. - **Owner:** the `Terraform*` object, through a non-controller reference with `blockOwnerDeletion` unset. The state Secret is never the owner, which is why a backup survives the state being deleted by hand. After a restore, the next reconcile of the object replaces a missing or earlier-UID reference. A set is **complete** when it has between 1 and 32 chunks, every index is present and the serial is at least 1. Only complete sets are listed or restorable. ## Retention `--state-backups` (default 5) is how many complete sets the manager keeps per object. After each backup it prunes the rest, newest first: - Only complete sets count toward the number kept. - An incomplete set is deleted, except the newest set overall, which may still be written. - A set whose serial is named by a pending `captf.io/restore-state` annotation is never pruned and does not count toward the limit. `status.stateBackups` lists the complete sets, newest first, at most 16. Each entry has the `serial`, the `takenAt` time and the `bytes`, the compressed size summed over the chunks. > [!WARNING] > > **Backups are a convenience, not disaster recovery** > > A run of failed or wrong applies writes a new serial each time, so five backups can hold five bad states and no good one. See [Security considerations](). ## Restore A restore pushes a backup into the backend through a Job. ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A["Annotation captf.io/restore-state=serial"] --> B{"Names a complete backup,
not already restored?"} B -->|no| X["RestoreBackupNotFound,
or skipped"] B -->|yes| C["Take leases, start the restore Job"] C --> D["Runner: init, state push -force,
state list"] D -->|success| E["Remove the annotation,
adopt, record the serial"] D -->|failure| F["Keep the annotation,
record the serial"] ``` The sequence: 1. You set `captf.io/restore-state=` on the object. 2. The controller re-reads the annotation and `status.lastRestoredSerial` from the API server, not from its cache. A value that is not a serial, or that names no complete backup, sets `RestoreJobSucceeded=False`/`RestoreBackupNotFound` and starts nothing. A serial equal to `status.lastRestoredSerial` is skipped. 3. The controller takes the object’s run lease (and the cluster write lease for a `TerraformCluster`) and starts a restore Job. Its root module declares only the backend, and its `captf-run-` Secret carries an empty tfvars file. Its config volume is a projection of that Secret plus each backup chunk, mapped to `restore/`. 4. The runner reassembles the chunks under the reader’s caps, checks that the result is valid JSON, and writes `restore.tfstate` with mode `0600`. 5. It runs `init`, a force-unlock only if the controller handed it a stale lock’s ID, `state push -force`, then `state list`. The Job fails if `state list` shows no managed resource although the backup recorded some. 6. Success or failure records the serial in `status.lastRestoredSerial`. On success the controller removes the annotation, the restored state is adopted with the backup’s inputs hash, and the next reconcile reads it. On failure the annotation stays. To retry the same serial after a failure, remove the annotation, wait until `status.lastRestoredSerial` clears, then set the annotation again; deleting the failed Job does not retry. A deleting object does not restore, with one exception: when the deletion is held because the state is lost or unreadable, the restore runs first and the destroy follows (see [lost state]()). Restoring writes back exactly the backup’s serial, which is already backed up, so it adds no new backup. > [!WARNING] > > **Resources created after the restored serial are orphaned** > > Resources created after that serial remain in the cloud and are no longer in the state. > [!NOTE] > > **See also** > > - [State restore runbook](). > - [Terraform State](). > - [`--state-backups`](). # Credentials A Job never reads the operator’s credentials Secret. The manager copies it into the object’s namespace, and the Job gets that copy. This page covers the pieces and how a credential change propagates. For creating an identity and choosing its namespaces, see [Identities and Credentials](). ## The pieces - **The source Secret.** Any name, in any namespace, created and owned by the operator. CAPTF never writes it. It carries no move label, so `clusterctl move` does not carry it (see [`clusterctl move`]()). - **The `TerraformClusterIdentity`.** A cluster-scoped object that names the source Secret in `spec.secretRef` and the namespaces that may use it in `spec.allowedNamespaces`. It carries no credential data. - **The mirror**, `captf-creds-`, in each allowed namespace where an object uses the identity. It is an `Opaque` Secret with a copy of the source’s data and: - labels `captf.io/mirrored=true` and `captf.io/managed=true`; - annotations `captf.io/source-hash` (a SHA-256 over the source’s data) and `captf.io/identity` (the identity’s name); - one non-controller owner reference per `Terraform*` object that uses it, with `blockOwnerDeletion` unset. A reference to an earlier UID of an object is replaced on its next reconcile. A name over 253 characters is shortened to `captf-creds-` and 16 hex characters of a hash. ## Who may point an identity at a Secret Creating an identity, changing `spec.secretRef`, or widening `spec.allowedNamespaces` makes the admission webhook run a `SubjectAccessReview`: the requesting user must be able to `get` the named Secret. An unset `allowedNamespaces` allows no namespace, and `{}` is rejected as ambiguous; `selector: {}` allows every namespace. Without the check, a role that may manage identities but not read Secrets could use an identity to have the manager mirror a Secret into a namespace it can read. See [Identities and Credentials](). ## When the mirror is written The manager makes sure the mirror is current: - on every reconcile of every non-deleting object that uses the identity; - just before it creates a deletion Job, so a destroy has credentials. Each time, it reads the source Secret uncached, hashes its data and compares the hash with the mirror’s `captf.io/source-hash`. If they differ it rewrites the mirror’s data. That is the only way a rotation reaches the mirror. ## Rotation Nothing watches the source Secret. A rotated credential reaches the mirror on the next reconcile of any object that uses it, which is bounded by the manager’s `--sync-period` (default ten minutes). A Job that is already running keeps the credentials it started with; the next Job gets the new ones. The identity’s own controller only reports status: it re-reads the source every five minutes to set the identity’s `Ready` condition and `status.namespaces`, and does not copy anything. ## Revocation If the identity stops allowing a namespace, because you narrowed `allowedNamespaces` or relabeled the namespace, the manager deletes that namespace’s mirror on the next reconcile and the objects there report `IdentityAllowed=False`/`NamespaceNotAllowed`. A destroy that needs the identity then waits with `ApplyJobSucceeded=False`/`IdentityNotAllowed` and requeues; the way out is to allow the namespace again or to [abandon the infrastructure](). When the last object using the mirror is removed, cleanup removes its own owner reference and deletes the mirror. You cannot delete an identity while any object uses it: the delete webhook refuses until `status.namespaces` is empty. ## Conflicts If a Secret named `captf-creds-` already exists and is not a mirror, the manager never overwrites it. It reports `CredentialsMirrored=False` with `MirrorFailed`. Rename or remove the Secret that is in the way. ## Delivery to the runtime The Job gets the mirror two ways at once: - `envFrom`, so every key becomes an environment variable; - a read-only mount at `/var/run/captf/credentials`, mode `0440`, one file per key. > [!WARNING] > > **Every key in the source Secret reaches the Terraform process** > > Anything in the source Secret reaches the Terraform process as an environment variable, whether or not the module uses it. The runner removes keys starting with `TF_` or `KUBE_` from the environment, except the three the Job sets itself (`TF_IN_AUTOMATION`, `TF_INPUT` and `KUBE_NAMESPACE`). Otherwise a key such as `TF_WORKSPACE` or `TF_VAR_x` could move the state or override the rendered inputs. A dropped key is still present as a file. More on the environment is in [Inside the Job]() and the [runtime environment](). > [!NOTE] > > **See also** > > - [Identities and Credentials](). > - [Identities and credentials runbook](). > - [Security considerations](). # Run Inputs and the Plan Key Three more Secrets belong to each object’s runs: a durable copy of the last rendered inputs, a per-run copy handed to one Job, and the key behind the plan hash. What goes into the inputs, and how the hash over them is computed, is in [Job Inputs](); this page covers where the rendered files are kept and for how long. ## Durable inputs `captf-inputs--` is written when an apply Job starts. It holds: - the data keys `main.tf.json` and `terraform.tfvars.json`: the rendered root module and its variables. Bootstrap data and every variable value are in clear text here; - annotations, in the table below. | Annotation | Holds | | --- | --- | | `captf.io/image` | The image reference | | `captf.io/identity` | The identity | | `captf.io/image-digest` | The digest the image resolved to, pinned after a successful apply | | `captf.io/applied` | The applied marker, set after the first successful apply | It has one owner reference to the object, and the labels `captf.io/managed` and the owner-kind and owner-name labels. It has no move label: it travels with `clusterctl move` by following the owner reference. What the durable Secret is for: - **Destroy and drift of an immutable machine.** A `TerraformMachine` is never re-rendered once provisioned, so its destroy, refresh and drift runs use the files, image and identity saved here. - **Pinning the image.** Changing the image reference clears the pinned digest, which is set again after the next successful apply. - **The applied marker.** `captf.io/applied: "true"` is set after the first successful apply, or when the manager reads a state that carries an inputs hash. Only deleting the Secret removes it. Because the Secret moves and status does not, the marker is how an object’s history survives a move; see [lost state](). A name that would pass 253 characters is shortened to a prefix and 16 hex characters of a hash. Cleanup deletes the Secret after a destroy. ## Per-run inputs `captf-run-` is created right after the Job, from that reconcile’s rendered inputs; the pod waits for the volume. It holds exactly `main.tf.json` and `terraform.tfvars.json`, carries only the label `captf.io/managed=true`, and is owned by the Job, so Kubernetes removes it with the Job. The controller also deletes it the first time it sees the Job finished. The reason for a second copy is that a running Job must never see its inputs change. The durable Secret is rewritten whenever a later apply renders new inputs; the per-run Secret is written once. A restore Job’s per-run Secret holds a backend-only root module and an empty tfvars file, and the Job’s config volume adds the backup chunks (see [Backups and restore]()). > [!WARNING] > > **Both Secrets hold the same plain text** > > Anyone who can read Secrets in the namespace can read it. See [Security considerations](). ## The plan key `captf-plankey--` holds 32 random bytes under the key `key`. It is owned by the object, labeled `captf.io/managed` with the owner labels, created the first time a plan needs it and never rotated. Cleanup deletes it, and it moves by its owner reference. The runner uses it as the key of an HMAC when it fingerprints a plan, so the `p2:` plan hash in `status.plan.planHash` reveals nothing about the planned values (see [What the plan hash binds]()). The key is mounted, read-only at `/captf/plan-key/key` with mode `0440` (runner flag `--plan-key-file`), only on: - plan Jobs, which `applyPolicy: Manual` uses; - the apply Job that carries an approved plan. A destructive-plan guard run under `Automatic`, and every drift, refresh, destroy and restore Job, do not mount it. Because the key never rotates, a hash stays valid for the object’s life. > [!NOTE] > > **See also** > > - [Job Inputs](). > - [Plan Approval](). > - [Inside the Job](). # Inside the Job A Job has one init container, which copies the runner binary into a shared volume, and one main container, which runs the role image with the runner as its entry point. The runner prepares a working directory, then runs Terraform or OpenTofu steps against it. This page lists what the pod mounts, what environment it gets, which permissions it holds, and which steps run for each operation. Fields you can change are in [Tuning Jobs](); the full environment is in the [Job environment]() reference. ## Volumes | Volume | Mount | Access | Contents | | --- | --- | --- | --- | | `runner` | `/captf/bin` | Read-only in the main container; read-write in the init container | The runner binary | | `work` | `/captf/work` | Read-write | The working directory, `HOME` and `TF_DATA_DIR` | | `tmp` | `/tmp` | Read-write | Scratch space | | `config` | `/captf/config` | Read-only | The per-run Secret; for a restore, a projection of it with the backup chunks | | `creds` | `/var/run/captf/credentials` | Read-only | The credentials mirror, one file per key | | `plan-key` | `/captf/plan-key` | Read-only | The plan key, on plan and approved-apply Jobs only | The three Secret volumes use mode `0440`, and the pod’s `fsGroup` defaults to `65532`, so a non-root image user can read them through that group. Both containers have a read-only root filesystem by default; a user `securityContext` on the main container can change that, within the limits [Tuning Jobs]() lists. The init container mounts only `/captf/bin`. ## Environment The Job sets `TF_IN_AUTOMATION=1`, `TF_INPUT=0`, `HOME=/captf/work`, `TMPDIR=/tmp`, `KUBE_NAMESPACE` and `CHECKPOINT_DISABLE=1`, and the credentials mirror adds its keys through `envFrom`. Entries you put in `spec.jobs.env` that begin with `TF_` or `KUBE_` are dropped. The runner then builds the environment each step sees from what the pod has: it removes every remaining `TF_*` and `KUBE_*` variable, whatever its source (the identity Secret, the image’s `ENV`), forces `TF_DATA_DIR` to `/captf/work/.terraform` and `HOME` to the working directory, and adds `TF_CLI_CONFIG_FILE` when the image ships providers. So a module cannot be steered by a `TF_WORKSPACE`, `TF_VAR_*` or `TF_CLI_ARGS` that arrived with a credential. See [Credentials](). ## Preparing the working directory Before the first step, the runner: - copies the top-level files of `/captf/config` into `/captf/work/root` with mode `0600`: the rendered module and the tfvars; - when the image has `/captf/providers`, writes a CLI configuration at `/captf/work/cli.tfrc` with mode `0600` that installs providers only from that filesystem mirror (`direct` is excluded), so a Job never downloads a provider. ## ServiceAccount and RBAC The Job runs as the ServiceAccount `captf-runner` in the object’s namespace, or as an override ServiceAccount you label `captf.io/runner=true`. The manager creates the ServiceAccount and a RoleBinding named `captf-runner` to the ClusterRole `captf-runner`, per namespace, before the first Job. Its rules are shipped in `config/rbac/runner_clusterrole.yaml`: | Resource | Verbs | Why | | --- | --- | --- | | `secrets` | `get`, `list`, `create`, `update`, `delete` | The backend reads, writes, lists and deletes surplus state chunks | | `leases` (`coordination.k8s.io`) | `get`, `create`, `update` | The state lock, unlock and force-unlock; the backend never deletes the Lease, so there is no `delete` | | `events` (`events.k8s.io`) | `create` | Run progress events on the object | > [!WARNING] > > **The Secret verbs cover every Secret in the namespace** > > Kubernetes RBAC cannot restrict a Secret verb to labeled or named objects when the names are dynamic, so these verbs cover every Secret in the namespace. [Security considerations]() explains what that means. When a namespace has no `Terraform*` objects left, the manager deletes the managed ServiceAccounts, RoleBindings and Leases there, both after the last finalizer is removed and periodically. ## Steps by operation Every operation begins with `init`, which receives the backend settings (see [Terraform state Secrets]()). If the controller found a stale lock before starting the Job, a force-unlock step follows `init`. Then: | Operation | Steps after `init` | | --- | --- | | `apply` | `validate`, `apply -auto-approve` | | Guarded or approved `TerraformCluster` apply | `validate`, `plan -detailed-exitcode -out`, `show -json`, then `apply` of the saved plan | | `plan` (`applyPolicy: Manual`) | `validate`, `plan -out`, `show -json` | | `destroy` | `destroy` | | `refresh` | `apply -refresh-only` | | `drift` | `apply -refresh-only`, `plan -refresh=false`, `show -json` | | `restore` | `state push -force`, `state list` | The steps that touch state (`init`, `plan`, `apply`, `destroy`, `refresh` and `state push`) pass `-lock-timeout`, from `lockTimeoutSeconds` (default 300 seconds). The others, `validate`, `show`, force-unlock and `state list`, do not. ## Locks Terraform takes the state lock for the run, in the Lease `lock-tfstate-default-`. Before it starts a Job, the manager checks that Lease. A lock is force-unlocked only when its holder is one of the object’s own runner pods and that pod is gone or has finished. Any other holder leaves the lock alone and reports `StateReadable=False`/`StateLocked`, and you must unlock it yourself; see the [stale lock runbook](). CAPTF has two more Leases of its own, in the object’s namespace, apart from the backend’s: the run lease `captf-run-` keeps two Jobs from running for one object, and the cluster write lease `captf-cluster-<16 hex of sha256("/cluster")>` makes machines and pools wait while a `TerraformCluster` applies, destroys or restores (the cluster lease is on by default; `--cluster-operation-gate=false` turns it off, and the run lease is always on). They are released when the Job finishes, by cleanup and by the namespace sweep. See [run leases](). ## Redaction The runner replaces known secrets with `(sensitive)` in failure summaries, events and its own logs. This is best effort; see [What CAPTF keeps out of status, events and logs](). > [!NOTE] > > **See also** > > - [Runtime environment](), what a module sees. > - [RBAC](). > - [Runner CLI](). # Lifecycle Walkthroughs This page follows the Secrets and Leases through the events of an object’s life. It uses a `TerraformMachine` as the running example; a `TerraformCluster` and a `TerraformMachinePool` differ only where noted. Each section names the page that covers the detail. ## Create and first apply 1. You create the identity’s source Secret and the `TerraformClusterIdentity` (see [Credentials]()), then the Cluster API objects. Cluster API creates the `TerraformMachine` and makes the `Machine` its owner. 2. The manager’s reconcile resolves the identity, makes sure the mirror `captf-creds-` is current, and creates the `captf-runner` ServiceAccount and RoleBinding in the namespace if they are missing. 3. It renders the inputs, writes the durable `captf-inputs--` Secret, takes the run lease (and, for a `TerraformCluster`, the cluster write lease), creates the Job, and creates the per-run `captf-run-` Secret for it. 4. The Job’s runner runs `init` and `apply`. Terraform creates `tfstate-default-` and the lock Lease `lock-tfstate-default-`. Neither is owned by anything yet. 5. When the Job finishes, the manager deletes the per-run Secret and releases the leases, reads the state, and adopts it: it adds the owner reference to every chunk and sets `captf.io/inputs-hash`. It pins the image digest and sets `captf.io/applied` on the durable Secret, and takes the first backup, `captf-state-backup--`. ``` sequenceDiagram participant Op as Operator participant M as Manager participant J as Job Op->>M: identity, Secret, Cluster API objects M->>M: mirror, ServiceAccount, RoleBinding M->>M: durable inputs, leases M->>J: create Job, then per-run Secret J->>J: init, apply (creates state and lock) M->>M: delete per-run Secret, release leases M->>M: adopt state, pin digest, first backup ``` ## Refresh and drift A refresh or drift Job runs against the existing state: it takes the run lease, gets a new per-run Secret, and holds the state lock while it runs. If it writes a new serial, the manager sees it on the next reconcile and takes a backup. A refresh does not change the inputs hash, so it does not adopt the state again; any chunk it creates has no owner reference until the next adoption. ## A change For a mutable object, a change in the image or the inputs produces a new inputs hash. The manager rewrites the durable Secret (an image change also clears the pinned digest), starts an apply Job with a new per-run Secret, and adopts the state with the new hash when it succeeds. With `applyPolicy: Manual`, a plan Job runs first with the plan key mounted, and the apply waits for your approval (see [Run inputs and the plan key]()). Each new serial adds a backup, and the oldest complete set beyond `--state-backups` is pruned. An immutable `TerraformMachine` never re-renders after it is provisioned: a change means a new object, and the durable Secret keeps what the old one needs for its destroy. ## Delete 1. You delete the object, or Cluster API deletes it with its `Machine`. The manager refreshes the mirror just before it creates the destroy Job, so the Job has credentials, and takes the leases. 2. The Job’s runner runs `init` and `destroy`. 3. When the destroy succeeds, cleanup deletes the state Secrets one by one and the lock Lease, the durable inputs Secret, the plan key and the leases, removes the object’s reference from the mirror (the mirror is deleted if it was the last), and removes the finalizer. 4. The state backups are not deleted by cleanup. They are owned by the object, so Kubernetes garbage-collects them when the object is gone. 5. When the namespace has no `Terraform*` objects left, the manager deletes the managed ServiceAccounts, RoleBindings and Leases in it. An object that never applied has no state, so it is deleted at once, with the same cleanup. ## `clusterctl move` `clusterctl move` copies the objects it discovers to the target, then deletes them from the source with their finalizers stripped, so the manager’s own delete path never runs on the source. What arrives: - The state chunks, the backups and the durable inputs Secret, which carry the move label or an owner reference to a moved object. The plan key and the mirror also follow their owners. - Not the identity’s source Secret: you copy it yourself. - Not the Leases. The backend recreates the lock Lease at the target’s next `init`, and the manager takes new run leases. - Not status. The target rebuilds it from the state on its first reconcile, which is what `captf.io/applied` on the durable Secret is for. The procedure is the [`clusterctl move` runbook](). ## Namespace deletion Every Secret CAPTF keeps is namespaced, so deleting a namespace removes the state, the backups and the durable inputs together with the objects. > [!CAUTION] > > **Deleting a namespace orphans the cloud resources** > > The cloud resources are not touched by that. Two consequences follow: - An object whose namespace is terminating that never applied finishes at once. One that applied is held, and you can [abandon it](). A destroy or restore that waits on credentials shows that in the `Deleting` condition and retries. - After a move, an object in a namespace that is then deleted cannot be told apart from one that never applied, because the marker went with the Secret. Move first, then delete the source namespace only when the target has taken over. ## Lost state If the state Secrets disappear (someone deletes them by hand, or an etcd restore brings back an older cluster state without them), the manager decides from what remains whether the object ever applied. It counts it as applied when any of these is true: - the object is provisioned; - the durable inputs Secret carries `captf.io/applied: "true"`; - the durable inputs Secret pins an image digest; - a state backup exists. An applied object with no state reports `StateReadable=False`/`StateLost`. No Job runs. > [!WARNING] > > **Applying again would duplicate live resources** > > Applying again would create a second set of resources next to the live ones. A deleting object is held: its finalizer stays and the backups are preserved. You choose: - **Restore**, with `captf.io/restore-state=` from `status.stateBackups`. The restore runs first, then the normal destroy (see [Backups and restore]()). - **Abandon**, with `captf.io/abandon-infrastructure=`. The finalizer is removed without a destroy, the infrastructure stays running untracked, and the backups go with the object. An object that never applied has nothing to lose, so it is deleted immediately. The runbooks are [unreadable state]() and [state restore](). > [!NOTE] > > **See also** > > - [The reconcile lifecycle](). > - [Terraform State](). > - [Runbooks](). # Operator Files and Settings CAPTF creates most of its Secrets itself. This page lists the few things an operator provides or must keep right: the credentials Secret and identity, the runner’s ClusterRole, the manager’s environment and one flag, the encryption of Secrets at rest, and what to back up outside the cluster. ## Credentials Secret and identity One Secret and one identity per set of cloud credentials. The Secret’s keys are exactly what the module’s providers read from the environment (and, as files, from `/var/run/captf/credentials`). The identity names it, and lists the namespaces that may use it. ```yaml apiVersion: v1 kind: Secret metadata: name: aws-prod namespace: platform type: Opaque stringData: AWS_ACCESS_KEY_ID: AWS_SECRET_ACCESS_KEY: --- apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: aws-prod spec: secretRef: name: aws-prod namespace: platform allowedNamespaces: list: - tenant-a ``` Allow only the namespaces that need the credentials: whatever runs in an allowed namespace can read the mirror (see [Security considerations]()). [Identities and Credentials]() describes the other `allowedNamespaces` forms, and the [identities runbook]() the failure modes. ## The runner ClusterRole The manager creates a RoleBinding named `captf-runner` in each namespace that has an object, to the ClusterRole `captf-runner`. The release manifests install the ClusterRole; this is what `config/rbac/runner_clusterrole.yaml` defines, with the `captf-` prefix that `config/default` adds: runner\_clusterrole.yaml ```yaml apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: captf-runner rules: - apiGroups: [""] resources: ["secrets"] verbs: ["get", "list", "create", "update", "delete"] - apiGroups: ["coordination.k8s.io"] resources: ["leases"] verbs: ["get", "create", "update"] - apiGroups: ["events.k8s.io"] resources: ["events"] verbs: ["create"] ``` > [!NOTE] > > **These are the minimum the state backend uses; do not trim them** > > Reading and writing state needs `get`, `list`, `create` and `update` on Secrets, `delete` is needed when Terraform deletes surplus chunks after a state shrinks, and the Lease verbs cover lock, unlock and force-unlock. The `events` rule is unused with the manager flag `--runner-events=false`. [RBAC]() covers the manager’s own permissions. ## Manager environment The manager Deployment must set `POD_NAMESPACE` and `SERVICE_ACCOUNT_NAME` through the downward API. The shipped manifest does: ```yaml env: - name: POD_NAMESPACE valueFrom: fieldRef: fieldPath: metadata.namespace - name: SERVICE_ACCOUNT_NAME valueFrom: fieldRef: fieldPath: spec.serviceAccountName ``` The webhook uses them to recognize the manager’s own ServiceAccount, the only identity allowed to set `TerraformMachine.spec.providerID`. Without them the manager logs a warning and the webhook refuses its writes, so machines never finish provisioning. See [Manager environment](). ## The `--state-backups` flag `--state-backups` sets how many complete state backups the manager keeps per object. The default is `5`; `0` takes no new backups but leaves the existing ones restorable; a negative value is rejected at start. Raise it if you want a longer history to restore from (each backup is a full copy of the compressed state), and see [Backups and restore]() for how the count works. Set it as in [Changing a flag](). ## Encrypting Secrets at rest CAPTF adds no encryption of its own to any Secret it keeps, and CAPTF cannot read OpenTofu’s client-side state encryption. Every Secret in this chapter, including the state, the inputs with their bootstrap data and the mirrored credentials, is stored by the API server the way any Secret is. Turn on encryption at rest for Secrets in the management cluster’s API server. An `EncryptionConfiguration` that encrypts Secrets with a key held in the file: encryption-config.yaml ```yaml apiVersion: apiserver.config.k8s.io/v1 kind: EncryptionConfiguration resources: - resources: - secrets providers: - aescbc: keys: - name: key1 secret: - identity: {} ``` Pass it to the API server with `--encryption-provider-config`. > [!TIP] > > **Prefer a KMS provider when you have one** > > A KMS provider keeps the key outside the cluster and is the better choice when you have one; the file above keeps the key on the control plane’s disk, next to the data it protects. After enabling it, rewrite the existing Secrets so they are stored encrypted (`kubectl get secrets -A -o json | kubectl replace -f -`). The Kubernetes documentation on [encrypting confidential data at rest]() is the reference for key rotation and providers. ## What to back up outside the cluster The backups CAPTF keeps are in the same namespace and owned by the object (see [Security considerations]()), so they do not protect against losing the namespace, the object or the cluster. Keep these outside it: - **etcd snapshots** of the management cluster, or a Velero (or equivalent) backup of the tenant namespaces that includes Secrets. That captures the state, the backups and the durable inputs together. - **The identity source Secrets**, which `clusterctl move` does not carry, and the `TerraformClusterIdentity` manifests (`clusterctl move` carries the identity object itself, but a backup must hold it too). Store the Secrets wherever you store other credentials; keep the manifests in version control. - **The module images** your objects reference, by digest. The durable inputs Secret pins the digest of the last successful apply; a destroy needs that image to exist. > [!WARNING] > > **Treat any such backup as sensitive as the Secrets themselves** > > The backups hold the same state, inputs and credentials. > [!NOTE] > > **See also** > > - [Installation]() and [Configuration](). > - [RBAC](). > - [Secrets](). # Security Considerations This page states what the way CAPTF keeps Secrets does and does not protect, so you can decide where to run which module. The wider threat model, including images, RBAC for the Terraform objects and Pod Security Admission, is in the [Security Model](); this page is about the Secrets in this chapter. ## The runner can reach every Secret in its namespace The Job runs as `captf-runner`, whose ClusterRole grants `get`, `list`, `create`, `update` and `delete` on Secrets in the namespace (see [Inside the Job]()). The state backend needs those verbs, and Kubernetes RBAC cannot narrow them to the state Secrets, because the chunk names are dynamic. So the runner, and with it the module image and any provider or program the module runs, can read, overwrite and delete every Secret in that namespace: the state, the backups, the inputs and the plan keys of other objects, and the credential mirrors. > [!CAUTION] > > **The trust boundary is the namespace and the module image publisher** > > The trust boundary is therefore the **namespace** and the **publisher of the module image**. A module in a namespace is trusted with everything in it. Put modules you do not fully trust in their own namespaces, with their own identity that allows only that namespace, and do not share a namespace between tenants. ## The plan key does not protect against a hostile module (See also [Limits]() for the approval gates.) The plan key is readable by the module (it is mounted into the Job that plans). The `p2:` plan hash is a keyed fingerprint: it detects that the plan changed between the review and the apply, and it hides the planned values from anyone who can read `status.plan`. > [!WARNING] > > **Review the module you approve, not only its plan** > > The plan key does not defend against a module image that wants to lie about its plan, because that image can read the key and compute the hash it likes. ## Backups are not disaster recovery State backups are copies taken after each write, in the same namespace, owned by the object. They help with a state Secret that was lost or damaged. They do not help when: - **The object goes away.** Deleting the object, or removing its finalizer by hand, garbage-collects every backup with it. - **The namespace or cluster goes away.** They are Secrets in that namespace. - **Bad applies pile up.** Each new serial adds a backup and the oldest complete set beyond `--state-backups` (default 5) is pruned, so a run of bad applies can push the last good state out. - **Someone with Secret access acts.** The runner can delete them, as above, and so can anyone with the right in the namespace. Back the namespace up outside the cluster; see [What to back up outside the cluster](). ## No encryption at rest of its own CAPTF does not encrypt the Secrets it keeps. The state, the inputs with their bootstrap data and variable values, and the credential mirrors are stored as the API server stores any Secret. Enable encryption at rest for Secrets on the management cluster (see [Encrypting Secrets at rest]()) and limit who can read Secrets. OpenTofu’s client-side state encryption is not supported: an encrypted state reports `StateReadable=False`/`StateEncrypted`. ## Credentials arrive twice and in full The Job mounts every key of the credentials mirror as a file and also injects it as an environment variable, so the Terraform process, its providers and anything the module runs see all of it, whether or not the module uses each key. Put only what the module needs in an identity’s source Secret, and prefer short-lived credentials where the cloud supports them. The runner removes `TF_*` and `KUBE_*` keys from the environment so a credential Secret cannot steer Terraform itself (see [Credentials]()); that is a guard against a mistake, not a boundary against the module. ## Rotation is not instant A change to the source Secret reaches the mirror on the next reconcile of an object that uses it, which takes up to the manager’s `--sync-period` (ten minutes by default). A Job already running keeps its old credentials. After revoking a credential in the cloud, expect Jobs started in that window to fail with the old one, and to pick up the new one on the next run. Narrowing `allowedNamespaces` takes effect the same way, at the next reconcile. ## What CAPTF keeps out of logs The runner replaces known secrets with `(sensitive)` in failure summaries, events and logs, and CAPTF never logs state or inputs. This is best effort: it cannot recognize a secret it was not told about, so do not treat it as a reason to put sensitive data in places Secrets would not go. See [What CAPTF keeps out of status, events and logs](). > [!NOTE] > > **See also** > > - [Security Model](). > - [RBAC](). > - [Secrets](). # Jobs, Retries and Concurrency Every Terraform or OpenTofu run CAPTF does is one Kubernetes Job, started by the controller, watched to the end and counted. This chapter follows a Job from the decision to run it to the bookkeeping after it finishes, and explains the machinery that keeps those Jobs from colliding: the names that make creation idempotent, the retry rules, the deadlines, the leases, the cache-lag checks, the state lock and leader election. It is the model behind the “nothing is happening” questions that [Slow Jobs]() and [Failing Jobs]() answer one symptom at a time, and it goes deeper than [The Reconcile Lifecycle](), which it extends. The pages: - **Job names, attempts and history** --- Deterministic names, adoption and pruning. - **Choosing the Operation** --- The priority order of `DecideOp`. - **Retries and backoff** --- What counts as a failure and what does not. - **Deadlines and lock timeouts** --- The two clocks on a Job. - **Leases and the operation gate** --- Who may run when. - **Cache lag** --- The live checks that cover a stale cache. - **The state lock and stale locks** --- Terraform’s lock against CAPTF’s leases. - **Leader election and failover** --- One active manager, and what a failover does to Jobs in flight. Two more pages sit alongside them: - [Requeue intervals and schedules](). - [Nothing is happening](): from a wait reason to the action. This page covers the life of a Job, the operations and the flags. ## The life of a Job ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A[Reconcile: bookkeeping,
no active Job] --> B[DecideOp picks an operation] B --> C{Backoff,
approval or gate?} C -- wait --> W[Requeue] C -- run --> D[Take the run lease,
and the cluster lease] D --> E{Lease free?} E -- no --> W E -- yes --> F[Create the Job,
then its per-run Secret] F --> G[Job runs] G --> H[Bookkeeping: read the pod,
delete the per-run Secret] H --> I[Set conditions and lastRun,
release the leases, prune] I --> A ``` Four rules shape the diagram: - **At most one Job per object.** The reconcile does nothing else while a Job of the object runs, except to watch it. The run lease makes that hold across managers and across crashes. - **Creation is idempotent.** Names are deterministic, so a retry after a crash or a stale read finds the Job it already created. - **The controller owns retries.** A Job has `backoffLimit: 0`, a pod `restartPolicy: Never` and no TTL; every retry is a new Job, named with a new attempt number, started only after the backoff. - **Results are read once.** After a Job finishes, one pass reads its pod, deletes its per-run Secret, records the outcome on the object and marks the Job as bookkept. Later passes read the Job’s annotations, not its pod. ## The operations | Operation | What it runs | Chosen when | | --- | --- | --- | | `apply` | `validate`, then `apply` of the rendered inputs | No state, state without an inputs hash, changed inputs, a failed last apply, or drift remediation | | `plan` | `validate`, then `plan` for review | A `TerraformCluster` under `applyPolicy: Manual` needs a plan to approve | | `destroy` | `destroy` of the durable inputs | The object is deleting | | `refresh` | `apply -refresh-only` | Once after an apply, and on the health, membership and pending schedules | | `drift` | `apply -refresh-only`, then a plan without refresh | The drift interval is due | | `restore` | `state push` of a backup, then `state list` | `captf.io/restore-state` names a complete backup | Which of these wins when several are possible is [the priority order](). Every op is a Job of the same shape, differing in the runner’s `--op`. ## Manager flags and where the rest lives The manager has no flags for backoff or for the failure limit. The retry timing is compiled in, and the limit is the per-object `spec.jobs.failedJobsHistoryLimit`. The flags that matter here: | Flag | Default | What it does | Page | | --- | --- | --- | --- | | `--cluster-operation-gate` | `true` | A `TerraformCluster`’s apply, destroy or restore and its machines’ operations exclude each other | [Leases]() | | `--sync-period` | `10m` | The informers’ resync, and the orphan sweep’s interval | [Schedules]() | | `--drift-default-interval` | `30m` | Drift interval for objects that set none | [Schedules]() | | `--terraformcluster-concurrency` and the machine, template and pool counterparts | `10` each | Reconciles in flight per kind | [Leader election]() | | `--leader-elect` and its lease, renew and retry flags | `false`; `15s`, `10s`, `2s` | One active manager | [Leader election]() | | `--runner-events` | `true` | Runner progress events | [Events]() | | `--state-backups` | `5` | Backups kept per object | [Backups]() | The per-object knobs are in `spec.jobs`: `activeDeadlineSeconds`, `lockTimeoutSeconds`, `successfulJobsHistoryLimit` and `failedJobsHistoryLimit`. See [Tuning Jobs](). The full flag list is [Manager Flags](). > [!NOTE] > > **See also** > > - [The Reconcile Lifecycle](). > - [Deletion and Teardown]() for the destroy path. > - [Approvals and Gates]() for the plan flow. > - [Drift and Health](). # Job Names, Attempts and History A Job’s name is a function of what it does, so the controller can create the same Job twice without starting two. This page covers the name, the attempt number, what happens when the Job already exists, and how finished Jobs are kept and pruned. ## The name ```text captf----a- ``` | Part | Value | | --- | --- | | `kindshort` | `c` for a `TerraformCluster`, `m` for a `TerraformMachine`, `mp` for a `TerraformMachinePool` | | `name` | The object’s name; replaced by the first 16 hex characters of its SHA-256 when the whole name would pass 57 characters | | `op` | `apply`, `plan`, `destroy`, `refresh`, `drift` or `restore` | | `attempt` | One more than the highest attempt among the retained Jobs of the same op, whatever their outcome | | `hash6` | The first six hex characters of SHA-256 over the inputs hash, the op, the attempt and a tick | > [!NOTE] > > **The 57-character cap is deliberate** > > The Job controller names pods `-<5 characters>`, and the kubelet truncates a pod’s hostname at 63. At 57 the full pod name is the hostname Terraform records as the lock holder, so [stale-lock detection]() can find the pod by name. ### What goes into the hash The inputs hash and the tick differ by op, so that a repeated op gets a new name only when it should: | Op | Inputs hash | Tick | | --- | --- | --- | | `apply`, `plan` | The hash of the current rendered inputs | none | | A drift-remediation `apply` | The same | `remediate/` | | `refresh` | The hash recorded in state | The time of the last successful refresh | | `drift` | The hash recorded in state | The time of the last successful drift check | | `restore` | The backup’s hash | `restore/` | | `destroy` | The hash recorded in state | none | A refresh or drift retry therefore has the same name until one of them succeeds, and the next one has a new name. A remediation apply of unchanged inputs would otherwise collide with the retained first apply, so it carries its own tick. ## Attempts The attempt number is `max(attempt of retained Jobs of this op) + 1`, not the count of failures. Counting failures would reuse a name twice over: once pruning removed older failures (a1 to a5 failed, the newest three kept, a count of three would reuse a4), and again with a retained success when an apply repeats the same inputs. Taking the highest retained attempt makes a new name never equal a retained one. The attempt is also a label (`captf.infrastructure.cluster.x-k8s.io/attempt`) next to `.../op`, the owner kind and name, the cluster name and `captf.io/managed=true`, on the Job and its pods. ## Adopting a Job that exists The controller creates the Job, then its per-run Secret, then records `status.activeJob`. A crash between any two, or a pass whose Job cache has not yet seen the Job, repeats the sequence with the same name: 1. Creating the Job returns `AlreadyExists`. The controller reads the existing Job and carries on with it, so no second Job starts. 2. The per-run Secret, owned by the Job, is created if missing. One that already belongs to the same Job (by UID) is left as is, because the name embeds the inputs hash and so the content is identical. One owned by a different Job of the same name, left over from an earlier pruned Job whose garbage collection is pending, is deleted and recreated; adopting it would let that collection delete it under the new pod. Any other create error is returned. A **rejected** create (invalid, forbidden, unauthorized, bad request, too large or namespace gone) proves no Job exists, so the controller [gives the leases back](); a timeout or server error proves nothing and the leases are kept. A Job that can never start is deleted. If an active Job is more than a minute old, its per-run Secret is missing and every pod is still `Pending`, the controller deletes it and emits `StuckJobDeleted`; the next pass starts the operation again with its Secret. This also runs on a paused object, so a Job that can never start does not hold `clusterctl move`. ## The Job itself - `backoffLimit: 0`: a failed pod is a failed Job. The controller retries, with its own backoff. - `activeDeadlineSeconds`: the [deadline](). - Pod `restartPolicy: Never`, a termination grace period of 600 seconds, and no TTL: finished Jobs stay until pruned, because the controller derives backoff, digest pinning and conditions from them. - The pod mounts the per-run Secret (`captf-run-`) that holds the rendered files. It is deleted when the Job finishes, so rendered inputs do not outlive the run; see [Inside the Job](). ## Annotations the controller writes | Annotation | On | Meaning | | --- | --- | --- | | `captf.io/bookkept` | Any finished Job | Its pod was read, its Secret deleted and its completion counted | | `captf.io/interrupted` | A bookkept Job | The runner was stopped from outside, not by the deadline | | `captf.io/destructive-plan-blocked` | An apply | The runner stopped before a plan that deletes or replaces | | `captf.io/plan-changed` | An apply | The approved plan no longer matched | | `captf.io/plan-unreadable` | A plan | It succeeded without a readable plan, and counts as a failure | | `captf.io/drift-remediation` | An apply | The apply remediates drift | | `captf.io/approved-plan` | An apply under `Manual` | The plan hash it had to plan again | The marks exist because a bookkept Job’s pod is not read again: the result lives on in the annotation, the conditions and `status.lastRun`. ## History After bookkeeping, the controller prunes finished Jobs beyond `successfulJobsHistoryLimit` and `failedJobsHistoryLimit` (default 3 each), **per object and per op**, oldest first. Refresh and drift history never pushes out an apply’s. Two Jobs of an op stay whatever the limits say: - its newest success, so that an older failure cannot pose as the op’s latest outcome in `ApplyJobSucceeded`; - its newest failure while no newer success exists, since backoff and the remediation cap count it. Running Jobs are never touched. Pruned Jobs take their pods with them (background propagation). Because pruning caps the number of retained failures at the limit, the limit also bounds the retry count; see [Retries and backoff](). > [!NOTE] > > **See also** > > - [The Reconcile Lifecycle: Job names and history](). > - [Tuning Jobs: history limits](). # Choosing the Operation When no Job of the object is active, one pure function decides what runs next: `DecideOp`. It reads a snapshot (the Jobs, the state, the inputs hash, the annotations and the clock), returns an operation or a wait, and has no side effects. This page gives its order, the reasons it reports, and the exceptions. The shorter version is in [The Reconcile Lifecycle](); this is the full list. ## The order The first rule that applies decides the pass. | \# | Condition | Decision | Reason | | --- | --- | --- | --- | | 0 | A Job is active | Nothing; requeue in a minute | `JobActive` | | 1 | `captf.io/restore-state` names a complete, unconsumed backup, and the object is not deleting, or its deletion is held | `restore` | `RestoreRequested` | | 2 | Deleting, state lost or unreadable | Wait one minute | `DeletionHeld` | | 3 | Deleting, no state | Drop the finalizer | `DeletingWithoutState` | | 4 | Deleting | `destroy` | `Deleting` | | 5 | No state | `apply` | `NoState` | | 6 | State without an inputs hash | `apply` | `StateWithoutInputsHash` | | 7 | Mutable kind, and the current inputs hash differs from state’s | `apply` | `InputsChanged` | | 8 | Mutable kind, and the newest apply failed | `apply` | `LastApplyFailed` | | 9 | Drift found, action `Remediate`, and fewer than the failed-limit remediations failed since the last drift check | `apply` | `DriftRemediation` | | 10 | None of the above | The schedule below | | An apply reason from 5 to 9 can still be turned into something else: - **Under `applyPolicy: Manual`** (a `TerraformCluster` only), every apply reason but `NoState` becomes the plan flow: a `plan` Job when no plan of the current inputs is recorded, the apply with the approved plan hash once the approval annotation names it (or the plan changes nothing), else a wait. See [Approvals and Gates](). - **Otherwise**, an apply whose newest attempt was blocked before a destructive plan, for the same inputs hash, waits for the approval annotation or new inputs. See [The destructive-plan guard](). - **Backoff** replaces any `ActionJob` of an op that failed recently with a wait; see [Retries and backoff](). > [!NOTE] > > **An approval wait is bounded** > > The wait for an approval never requeues for longer than ten minutes, and approvals and input changes trigger a reconcile on their own. ### Why immutable kinds still retry Rules 7 and 8 apply to mutable kinds (a `TerraformCluster` and a `TerraformMachinePool`). A `TerraformMachine` is immutable: its inputs never change after creation, so it has no `InputsChanged` and no `LastApplyFailed`. It retries a failed apply through rules 5 and 6 instead: the state carries an inputs hash only after a successful apply, so a partly written state still reads as `StateWithoutInputsHash`, and the apply runs again after its backoff. ## The schedule With no apply to run, the reconcile checks these in order and stops at the first that is due. Each interval has a deterministic jitter of up to a tenth of the interval, derived from the object’s UID. | \# | Check | Operation | Reason | | --- | --- | --- | --- | | 1 | A successful apply not yet followed by a refresh, for kinds that ask for one (machine, pool) | `refresh` | `RefreshAfterApply` | | 2 | A pool’s membership is converging | `refresh` every 30 s, fixed | `MembershipConverging` | | 3 | Health reads pending | `refresh` after 30 s, doubling to 5 min | `HealthPending` | | 4 | The membership interval (pools) | `refresh` | `MembershipRefreshDue` | | 5 | The health-check interval (when remediation is on) | `refresh` | `HealthCheckDue` | | 6 | The drift interval | `drift` | `DriftDue` | | 7 | Otherwise | Requeue at the soonest deadline | `UpToDate` | Notes: - Rule 1 is skipped when the apply’s own output reading is definite (neither pending nor unknown): the reading stands in for the refresh and `status.lastRefresh` advances to the apply’s finish. - Rules 3 and 2 are exclusive: while converging, the fixed 30 seconds replaces the doubling pending delay. - The base of a periodic check is its own last run, else the last successful apply (so the first check comes one interval after provisioning), else the object’s creation time. After a `clusterctl move` there is no history, so the creation time is the base and the jitter keeps moved objects from checking in lockstep. - A drift or refresh Job runs against the current inputs for a mutable kind, and the durable ones for an immutable kind. A waiting input change pauses the whole schedule, the way a backoff does: a refresh would otherwise render the unapplied inputs and report the waiting change as drift. ## What each op reads and writes | Op | Inputs it runs against | Writes the durable Secret | | --- | --- | --- | | `apply`, `plan` | The current inputs | `apply` only | | `refresh`, `drift` | Current (mutable) or durable (immutable) | No | | `destroy` | The durable inputs | No | | `restore` | None: a backend-only root and the backup chunks | No | Only `apply` and `plan` run the spec’s image as written. Every other op runs the digest pinned after the last successful apply, when one exists; without one it falls back to the spec reference and emits `DigestUnknown`. > [!NOTE] > > **See also** > > - [Drift and Health]() for what the drift and refresh results mean. > - [Machine Pools]() for membership convergence. # Retries and Backoff CAPTF retries a failed operation forever, with a delay that grows to a cap. Nothing limits the number of attempts; what the failure limit bounds is the delay and the remediation retries. This page defines what counts as a failure, the delay, and the cases that deliberately do not back off. ## Counting failures For each operation separately, the controller walks the object’s finished Jobs newest first and counts the failed ones until it meets a success of that same op. A failing drift check therefore never delays an apply. | Outcome of a finished Job | Counts as a failure | | --- | --- | | Failed: a step failed, the image would not pull, the image broke the contract | Yes | | Killed by `activeDeadlineSeconds` | **Yes**, even if the runner reported itself interrupted | | Stopped from outside: a drain, an eviction, a Job deletion | No (`JobInterrupted`) | | An apply the runner blocked before a destructive plan | No (`DestructivePlanBlocked`) | | An approved apply whose plan no longer matched | No (`PlanChanged`) | | A plan Job that succeeded but whose plan could not be read | Yes: the next plan backs off instead of looping | | A lease wait | No: no Job existed | Two details explain the table: - **Interrupted versus the deadline.** On a deadline, Kubernetes sends the runner SIGTERM, as it does for a drain, and the runner may report the run as `interrupted`. The controller still counts it, because a step that always hangs would otherwise retry at once forever and never reach the cap. Only an interruption that is not a deadline kill is free. - **Blocked and plan-changed.** Both stopped before changing anything, and both are waiting for a person, not a timer. They count toward no backoff and no retry number; `DecideOp` waits for the approval or for new inputs instead (see [Choosing the operation]()). ## The delay The delay after `n` consecutive failures of an op is ```text min(1m × 2^(n-1), 10m) ``` counted from when the newest failed Job finished. Once `n` reaches the failed-Jobs history limit, it is the 10-minute cap straight away. The limit is `spec.jobs.failedJobsHistoryLimit`, default 3, and at least 1 (a limit of 0 counts like 1, because pruning always keeps the newest unresolved failure). The timing is compiled in: one minute base, ten minute cap. | Limit | Delay after failure 1, 2, 3, 4, 5 | | --- | --- | | 0 or 1 | 10m, 10m, 10m, 10m, 10m | | 3 (default) | 1m, 2m, 10m, 10m, 10m | | 5 | 1m, 2m, 4m, 8m, 10m | Pruning keeps at most the limit’s worth of failures, so `n` cannot exceed it; reaching the limit means “at least that many”. A success of the op resets the count. During the delay the decision is a requeue with the reason `Backoff` (for example `InputsChangedBackoff`); there is no way to skip it by hand. ## Retries by reason | What failed | What retries it | Delay | | --- | --- | --- | | An apply of a mutable kind | `LastApplyFailed`, until an apply succeeds | Backoff | | An apply of a `TerraformMachine` | `NoState` or `StateWithoutInputsHash` (state has no inputs hash until an apply succeeds) | Backoff | | A drift remediation apply | The remediation cap, below | Backoff | | An apply Job deleted while it ran (cluster, pool) | An apply of the current inputs stays due; see [below](<#an-apply-job-deleted-while-it-ran>) | Backoff | | `refresh`, `drift` | The schedule: still due, so due again | Backoff | | `destroy` | `Deleting`, forever | Backoff; see [Destroy]() | | `plan` | The plan flow | Backoff | | `restore` | Nothing: not retried for the same serial | None, see [below](<#restore>) | ### The remediation cap An apply that remediates drift (`captf.io/drift-remediation` on the Job) is not retried as an unconverged apply: `LastApplyFailed` ignores it. Instead, `DecideOp` counts the failed apply Jobs that finished after the last successful drift check, stopping at an apply success. Remediation starts only while that count is below the failed limit (at least 1). At the limit, the drift stays pending (`DriftDetected` stays `True`) and no more remediation runs until the next successful drift check resets the count. Blocked and plan-changed applies do not count. See [Drift](). ### An apply Job deleted while it ran If an apply Job of a `TerraformCluster` or `TerraformMachinePool` is deleted while it runs, it never finishes and bookkeeping never reads its result, but it may have applied part of its change. CAPTF confirms with live reads (the Job is `NotFound` and the object’s live `status.activeJob` still names it) and records `captf.io/interrupted-apply=` on the durable inputs Secret. An apply of the current inputs then stays due, even when the inputs equal the state’s, and is guarded where the destructive guard applies. `ApplyJobSucceeded` is `False`/`ApplyFailed`: `Job : disappeared while it ran and may have applied part of its change; an apply of the current inputs is due`. It clears when an apply started afterwards succeeds. A stuck Job that CAPTF deleted itself, and a `TerraformMachine`, record nothing. See [Machine pools](). ### Restore > [!WARNING] > > **A failed restore is not retried for the same serial** > > A failed restore is not retried for the same backup serial and counts toward no backoff. To try again, remove the annotation, wait for `status.lastRestoredSerial` to clear, and set it again. See [Other manual actions](). ## What does not back off These end the pass with a requeue, not a failure: | Wait | Requeue | | --- | --- | | A lease (`WaitingForRunLease`, `WaitingForClusterOperation`, `WaitingForMachineOperations`) | 30 s | | Credentials, dependencies or an owner not ready | 30 s | | An unreadable or lost state | 1 min | | A Job still running | 1 min fallback; the Job watch wakes the reconcile sooner | | The Job cache behind the API server | 5 s | | An approval for a blocked apply or a plan | At most 10 min; an annotation wakes the reconcile | | `JobPolicyInvalid`, `InputsTooLarge`, missing durable inputs | 10 min; a change wakes the reconcile | ## Metrics and events Each finished Job is counted once. `captf_job_attempts` records its retry number (one plus the earlier failed Jobs of the op since the last success, not counting blocked and plan-changed ones). The events are `JobFailed`, `JobInterrupted`, `JobDeadlineExceeded` and `DestructivePlanBlocked`, each once per Job. See [Metrics]() and [Events](). > [!NOTE] > > **See also** > > - [Failing Jobs](). > - [The Reconcile Lifecycle: retry backoff](). # Deadlines and Lock Timeouts A Job runs against two clocks. `activeDeadlineSeconds` bounds the whole Job. `lockTimeoutSeconds` bounds how long one runner step waits for the Terraform state lock. They are set per object in `spec.jobs`, inherited by machines and pools from their cluster’s `spec.defaults.jobs` field by field, and the second must stay below the first. | Setting | Default | Range | Enforced by | | --- | --- | --- | --- | | `activeDeadlineSeconds` | 3600 | 1 to 86400 | Kubernetes, on the Job | | `lockTimeoutSeconds` | 300 | 0 to 3600 | Terraform or OpenTofu, as `-lock-timeout` | Neither is a manager flag. See [Tuning Jobs]() for how to set them. ## The deadline Kubernetes counts `activeDeadlineSeconds` from the Job’s start. When it passes, the pod is sent SIGTERM and then, after the grace period, SIGKILL. The Job sets a termination grace period of **600 seconds** so that a SIGTERM is not a SIGKILL in 30 seconds: the runner interrupts the runtime, which finishes the provider calls in flight (a VM being created), writes the state and releases the lock. The runner’s own stop timeout is the grace period less a 30 second margin, which leaves time to write its result. A deadline kill is a failure, with three condition reasons by op: | Op | Condition | Reason | | --- | --- | --- | | `apply`, `plan`, `destroy` | `ApplyJobSucceeded=False` | `JobDeadlineExceeded` | | `refresh`, `drift` | `DriftJobSucceeded=False` | `DriftJobDeadlineExceeded` | | `restore` | `RestoreJobSucceeded=False` | `RestoreFailed` | > [!WARNING] > > **A deadline kill counts toward backoff** > > It **counts toward backoff** like any other failure, even though the runner reports the interruption: see [Retries](). An image that cannot be pulled also ends on the deadline, but its more specific reason, `ImagePullFailed`, is checked first. > [!TIP] > > **Raise the deadline for a slow module** > > Raising it also lengthens the [lease backstop](), because a lease is held up to the holder’s deadline plus 15 minutes. ## The lock timeout `lockTimeoutSeconds` is passed to the runner as `--lock-timeout`, and the runner adds `-lock-timeout=s` to the steps that take or wait for the state lock: | Step | Gets `-lock-timeout` | | --- | --- | | `init` | Yes: it succeeds while the lock is held | | `apply`, `plan`, `destroy` | Yes | | `apply -refresh-only` (refresh, and the first step of drift) | Yes | | `plan -refresh=false` (the second step of drift) | Yes | | `state push` (restore) | Yes | | `force-unlock`, `validate`, `show -json`, `state list` | No | A step that cannot get the lock within the timeout fails, and the Job fails with it. That counts toward backoff. The lock itself is the Terraform state lock, which is a different mechanism from CAPTF’s leases: see [The state lock and stale locks](). A timeout of 0 means do not wait: fail at once on a held lock. ## The merged-policy check The deadline must be longer than the lock wait, or a Job could reach its deadline while still waiting for a lock. The admission webhook checks one policy at a time against the built-in default of the field the policy does not set. It cannot see the field-wise merge of a machine’s policy over its cluster’s defaults, so the controller checks the **merged** policy again before it starts a Job: ```text lockTimeoutSeconds (default 300) must be less than activeDeadlineSeconds (default 3600) ``` When it fails, the object reports `ApplyJobSucceeded=False`/`JobPolicyInvalid` with a message naming both effective values and whether each is configured or the built-in default, no Job starts, and the reconcile requeues in ten minutes (a change to the object triggers it sooner). The message is on `ApplyJobSucceeded` whichever op was due: a refresh or drift that cannot start reports there too. Two exceptions: - **A destroy never waits on this check**, so a teardown cannot wedge on a policy error. See [The destroy Job](). - **A restore does not run it either**: the restore path starts its Job without the check. To fix it, lower `lockTimeoutSeconds` or raise `activeDeadlineSeconds` in the policy that sets it, or in the cluster’s `spec.defaults.jobs`. > [!NOTE] > > **See also** > > - [Slow Jobs](). > - [Job Environment]() for the Job’s exact arguments. # Leases and the Operation Gate Two Kubernetes Leases decide who may start a Job. They are CAPTF’s own locks, separate from the [Terraform state lock](), and they are taken **before** a Job exists, so a wait costs no pod and no failure. This page covers what each lease is, how a lease is taken and freed, the writer preference between a cluster and its machines, the wait reasons, and the one known gap. ## The two leases | Lease | Name | Taken by | When | | --- | --- | --- | --- | | **Run lease** | `captf-run-` | Every object, for every op | Always; no flag turns it off | | **Cluster write lease** | `captf-cluster-` | A `TerraformCluster` only | `--cluster-operation-gate` on, the op is `apply`, `destroy` or `restore`, and the cluster name is known | The run lease makes “one Job at a time per object” hold across managers and across a crash between taking the lease and creating the Job. `` is the object’s state suffix, so the name is the same length whatever the object is called. The cluster lease is keyed by a hash of `/`: there is one per Cluster. A lease is a `coordination.k8s.io` Lease with the labels `captf.io/lease=run` or `cluster`, `captf.io/managed=true`, the owner kind and name and the cluster name, and the annotations `captf.io/lease-op` and `captf.io/lease-acquired-at`. Its `spec.holderIdentity` is **the name of the Job about to be created**. Names are deterministic ([Job names]()), so the holder is known before the Job exists. ```sh kubectl get lease -n -l captf.io/lease=run kubectl get lease -n -l captf.io/lease=cluster ``` ## Taking a lease The controller takes a lease by creating it, or, when it exists, by updating it with the `resourceVersion` it just read from the API server (never a cache): 1. If the holder is this Job already, it refreshes the acquire time. 2. If the lease is **free**, it takes it over. 3. Otherwise it is not acquired, and the operation waits. A conflict or a concurrent create means another writer won; the loser waits for the next reconcile, never retrying in the same one. ### Free A lease may be taken over when any of these holds: - it has no holder; - its holder Job has finished (a read through the API server); - its holder Job does not exist and the lease is older than the **grace**, one minute: this covers a manager that crashed between taking the lease and creating the Job, and a reader racing the create; - it is older than its **backstop**, whatever its holder. ### Grace and backstop The backstop is the holder’s `activeDeadlineSeconds` plus 600 seconds (the pod’s termination grace period) plus five minutes: with the default deadline of one hour, 4500 seconds. The lease records it as `leaseDurationSeconds`. It frees a lease whose Job is stuck in a way Kubernetes did not end, for example a pod that cannot terminate. ## The operation gate A cluster’s apply, destroy or restore and its machines’ and pools’ applies, destroys and restores must not overlap: a cluster apply can change what the machines read, and a machine apply can run against a cluster in flux. Plans, refreshes and drift checks mutate nothing, so they take the run lease only. The two sides wait for each other in a fixed order, with **writer preference**: the cluster is not starved by a stream of machine creates. ``` sequenceDiagram participant K as TerraformCluster participant L as Leases participant M as Machine or pool K->>L: take run lease, then cluster write lease K->>L: list live machine and pool apply, destroy, restore alt machine operations in flight K-->>K: WaitingForMachineOperations, keep both leases M->>L: take own run lease M->>L: read cluster write lease: live M-->>M: WaitingForClusterOperation, give back own run lease else none K->>K: create the Job end ``` - **A cluster** takes its run lease, then the cluster write lease, then lists the live run leases of its machines and pools that are applying, destroying or restoring. If any are live, it **keeps both leases** and waits (`WaitingForMachineOperations`, naming up to five Jobs). Keeping the cluster lease is what makes new machine operations wait behind it. - **A machine or pool** takes its own run lease, then reads the cluster lease. If a live Job holds it, the machine gives its run lease back and waits (`WaitingForClusterOperation`). The write precedes the read on both sides, so on a linearizable store at least one of the two sees the other: they cannot both proceed. A machine that sees the cluster gives its lease back, so the cluster is not held up by an operation that never started. With `--cluster-operation-gate=false`, only the run lease is taken, and the cluster and its machines may run at once. See [Configuration](). ## The wait reasons A wait sets a condition to `Unknown`, requeues in 30 seconds, creates no Job, and counts toward no backoff or remediation cap. The event is `Normal` and is emitted once per wait, not on every requeue. | Reason | Meaning | Holder named in the message | | --- | --- | --- | | `WaitingForRunLease` | Another live Job holds this object’s run lease, or the Cluster’s write lease (a second cluster operation) | The Job; or “another manager took it first” when a concurrent writer won the race | | `WaitingForClusterOperation` | A machine’s or pool’s op waits for its cluster’s op | The cluster’s Job | | `WaitingForMachineOperations` | A cluster’s op waits for machine and pool ops in flight | Up to five Jobs and a count | The condition depends on the op: `ApplyJobSucceeded` for apply, destroy and plan; `DriftJobSucceeded` for refresh and drift (which can only wait for the run lease); `RestoreJobSucceeded` for restore. A wait that outlasts the holder’s deadline means the holder cannot finish. Look at that Job. ## Release A lease is given back at these points, always by checking the holder first and deleting with the UID and `resourceVersion` it read, so a lease another Job took in the meantime stays: - **When its Job finishes.** The pass that counts a finished Job (deletes its per-run Secret, once per Job) releases the run lease and, for a cluster, the cluster lease that Job held. - **When a create was rejected and no Job exists.** The create, or a pass that waited at the gate, proves no Job holds the lease only if the API server has no Job of that name. A rejection (invalid, forbidden, unauthorized, too large, namespace gone) releases both leases at once. A timeout or a server error may hide a create that went through, so the leases stay. Job names are deterministic, so a pass with a lagging cache can rebuild the name of a Job an earlier pass created, which is why the controller reads the Job live before it releases. - **When a machine sees the cluster operation,** its own run lease is given back at once (see above). The reverse is deliberate: a cluster that waits for machines keeps its leases. - **When the object is cleaned up** after a destroy, every run and cluster lease carrying its owner labels is deleted. See [Cleanup](). A failure to release is only logged: a lease whose Job finished is free anyway, and the next Job takes it over. ## The known gap A cluster that waits for machine operations holds both leases under the name of the Job it intends to create. That name embeds the inputs hash. If the inputs change **while it waits**, the next pass computes a different Job name and cannot refresh the old holder: its leases are held by a Job that does not exist. The new name waits for the grace, up to a minute after the last pass that refreshed them, before it takes the lease over. The cluster’s machines see the cluster lease as live for the same time. > [!WARNING] > > **Do not delete the Leases by hand** > > It resolves itself; no manual step is needed, and **do not delete the Leases by hand**. > [!NOTE] > > **See also** > > - [The Reconcile Lifecycle: run leases](). > - [Slow Jobs: waiting for a lease](). > - [Cache lag](): the lease as evidence of a Job the cache has not shown yet. # Cache Lag The controller lists an object’s Jobs through the manager’s informer cache. That cache can trail the API server by a moment, in particular behind a Job this controller created a second ago. A pass that reads “no Job runs” from a stale list would do harm: - clear the `clusterctl.cluster.x-k8s.io/block-move` annotation, letting `clusterctl move` start while a Job runs; - clear `status.activeJob`; - drop a deleting object’s finalizer while a Job still runs. So before any of those, the controller asks the API server directly. This page explains what it checks, when, and what it does on a mismatch. ## What is cached and what is not | Read | Through | | --- | --- | | The object’s Jobs (`List`) | The cache, scoped to the managed Jobs | | A single Job, to confirm `status.activeJob` or a lease holder | The uncached API reader | | The run and cluster leases | The uncached API reader | | Pods of a Job | The uncached API reader; pods are never cached | | Secrets (state, durable inputs, per-run, backups) | Live reads | | The state lock Lease | The uncached API reader | ## The two checks Both run when the pass found no active Job in the cached list, before it clears block-move and `status.activeJob`. Both run on the paused branch too. **`status.activeJob` against the list.** If `status.activeJob` names a Job the cached list lacks, the controller reads that Job from the API server. Found means the cache has not caught up: the pass waits. Not found means the Job is gone, and clearing it is right. **The live run lease.** Block-move is written to the object before a Job exists, and `status.activeJob` in a later patch, so a pass can see the first and not the second. The run lease is written before the Job is created and read live, so its holder is the evidence. If the lease is live and its holder Job is **not** in the cached list, the controller reads the holder from the API server and asks whether it belongs to this object (its owner-kind and owner-name labels). Yes means a running Job the caches have not shown: wait. A holder the API server does not have, or one of another object, is no lag: that wait is the lease gate’s own, and starting a Job is refused there. ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A["No active Job in the cached list"] --> B{"status.activeJob names
a Job the list lacks?"} B -->|"found on the API server"| W["Cache lag: wait"] B -->|"not found, or no name"| C{"Live run lease with a holder
not in the list?"} C -->|"holder is this object's Job"| W C -->|"no holder, or another object"| D["No lag: clear block-move
and status.activeJob"] ``` ## When the lease is read Two reads per reconcile for every idle object would be wasteful, so the lease check runs only when its answer can change what the pass does (`leaseMatters`): - the cached object carries block-move (which “no lag” would clear); - `status.activeJob` names a Job (which “no lag” would clear); - the object is deleting (a pass may drop the finalizer). An idle object with none of these reads neither. If it then wants to start a Job, the [run lease]() is the gate: a Job the cache has not shown yet holds the lease, so the new one waits for it. ## What a lag does The pass ends with a requeue after **5 seconds** (`LagRequeue`) and changes nothing else. The Job watch normally triggers the next reconcile as soon as the cache catches up. The same 5 seconds is used when removing a consumed annotation conflicts with a concurrent change, and when a deleting object’s [cleanup]() finds a live Job holding the run lease. The ownerless delete path uses the same evidence: an object with no owner reference is dropped only if there is no state Secret, no active Job, no lag on `status.activeJob` and no live run lease. See [Order and finalizers](). > [!NOTE] > > **See also** > > - [The Reconcile Lifecycle: the clusterctl move block](). > - [Leases and the operation gate](). # The State Lock and Stale Locks CAPTF has two kinds of lock, and they answer different questions. | | CAPTF’s leases | The Terraform state lock | | --- | --- | --- | | Question | May this object start a Job? | May this runner step read and write the state? | | Taken | Before the Job exists | By Terraform or OpenTofu, per step | | Held by | The name of the Job about to run | The runner’s pod, by hostname | | Objects | `captf-run-`, `captf-cluster-` | `lock-tfstate-default-` | | Released by | The controller | The runtime, or the next Job’s `force-unlock` | | Waiting costs | Nothing: no Job exists | A pod, and `lockTimeoutSeconds` | Because the run lease allows one Job per object, two CAPTF Jobs never contend for one object’s state lock. What can hold it is a **leftover** of a runner that died, or a **foreign** holder such as a workstation. This page covers how the controller tells those apart and what it does about each. ## Where the lock lives The Kubernetes backend keeps the state in `tfstate-default-` Secrets and the lock in the Lease `lock-tfstate-default-`. The lock’s details are in the annotation `app.terraform.io/lock-info`, a JSON document with `ID`, `Operation`, `Who`, `Version` and `Created`. `Who` is `user@hostname`, and inside a Job the hostname is the pod’s name. A Job’s pod name is `-<5 characters>`, and Job names are capped at 57 characters, so the hostname is the **whole** pod name and never truncated. That is what lets the controller recognize its own runner. ## The stale-lock rules The controller checks the lock only when no Job of the object is active, in the same bookkeeping pass that reads the finished Jobs. A held lock is **stale** only when all of these hold: 1. The Lease has a holder, and the lock info parses. 2. The part of `Who` after the last `@` names a pod of **this object’s own Jobs**: it matches the Job-name pattern for this kind and object, for any op and attempt, followed by five characters. Any other holder is “unknown” and is never stale. 3. That pod, read live from the API server, does not exist, or has reached a terminal phase (`Succeeded` or `Failed`). A finished Job keeps its pod object until pruned, and its containers have exited, so an OOM-killed or evicted runner leaves exactly this. A pod that is only terminating still counts as alive, because it may be running the runtime inside its grace period. The lock ID to release is `ID` from the lock info, else the Lease’s holder identity. ``` %%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%% flowchart TD A["Lock held, no Job active"] --> B{"Holder is a pod of this
object's own Jobs?"} B -->|no| F["Foreign: report StateLocked,
never unlock"] B -->|yes| C{"Pod gone, Succeeded
or Failed?"} C -->|no| L["Alive: leave the lock"] C -->|yes| S["Stale: next Job runs force-unlock"] ``` ## What happens to a stale lock The controller does not unlock it itself. It hands the ID to the **next Job** it starts, which runs `init` and then `force-unlock -force ` before its own operation: - A stale lock that is already gone by then, or one a different holder took since (`state is already unlocked`, `does not match existing lock`), is not an error. - The controller emits a `ForceUnlocked` Warning event naming the Job when it starts it, and counts the force unlock in its metrics. So a stale lock is cleared by whichever Job runs next: an apply, a refresh, a drift check or a destroy. If nothing is due to run (an apply waiting for approval, for example), the lock stays until something does. See [Stale State Lock](). ## A foreign lock If the lock is held and its holder is **not** one of the object’s own runner pods, the controller never unlocks it. > [!WARNING] > > **A foreign lock may be mid-write** > > A lock held from a workstation running `terraform state rm` could be mid-write, and the controller cannot know. It sets `StateReadable=False`/`StateLocked` with the holder, the operation and the time from the lock info, emits `StateLocked`, and keeps starting Jobs as usual. Every such Job waits `lockTimeoutSeconds` for the lock, then fails, and the failures back off. Release the lock with `force-unlock` once the holder is gone, and the next Job proceeds. `StateLocked` is not a deletion hold: the state reads, so the destroy starts and waits like any other Job. See [Held deletions](). ## Where the deadline fits The lock wait is bounded by `lockTimeoutSeconds`, which must be less than the Job’s deadline so a wait alone cannot consume it. See [Deadlines and lock timeouts](). > [!NOTE] > > **See also** > > - [Terraform State: locks](). > - [Slow Jobs](). > - [Unreadable State](). # Leader Election and Failover Only one manager reconciles at a time. This page covers how the election is configured, what runs only on the leader, and what a failover looks like to the Jobs that were in flight, which is where the leases, the deterministic names and the cache-lag checks pay off. ## Configuration | Flag | Default | Meaning | | --- | --- | --- | | `--leader-elect` | `false` | Turn the election on | | `--leader-elect-lease-duration` | `15s` | How long non-leaders wait before they force-acquire | | `--leader-elect-renew-deadline` | `10s` | How long the leader retries a renewal before it gives up | | `--leader-elect-retry-period` | `2s` | How often a candidate tries | The default is off, but the shipped Deployment sets `--leader-elect` and runs one replica, ready for a scale-up. > [!WARNING] > > **Run more than one replica only with leader election on** > > Turn the election on whenever you run more than one replica. The lock is a `coordination.k8s.io` Lease named `controller-leader-election-captf` in the manager’s namespace; the manager needs the [leader-election Role]() for it. The manager does not release the lease when it stops, so after a shutdown a replacement waits for the lease to expire, up to 15 seconds by default. See [Configuration](). ## What runs only on the leader | Component | On every replica | Leader only | | --- | --- | --- | | The `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool`, template and identity controllers | | Yes | | The orphan sweep (start, then every `--sync-period`) | | Yes | | The admission webhooks | Yes | | | Health and readiness endpoints | Yes | | A replica that is not the leader serves webhooks and idles. Per-kind concurrency (`--terraformcluster-concurrency` and its machine, template and pool counterparts, default 10 each) is the number of reconciles in flight on the leader. One reconcile never runs two Jobs for one object, so concurrency affects throughput across objects, not safety. ## What a failover does to a Job A Job is a Kubernetes object. It keeps running when the manager goes away, and the new leader finds it. Everything the old leader knew is in the API server: | Where the old leader stopped | What the new leader sees | What resolves it | | --- | --- | --- | | Before taking the run lease | Nothing | The decision is made again | | After taking the lease, before creating the Job | A lease whose holder Job does not exist | The same Job name is derived and acquired, or the lease is free after the one-minute [grace]() | | After creating the Job, before its per-run Secret | A Job whose pod waits for a missing volume | The next pass creates the Secret (the Job exists: `AlreadyExists` is adopted); after a minute a Job that cannot start is deleted and recreated | | After creating the Job, before the status patch | A Job and a block-move annotation, no `status.activeJob` | The Job is in the list or [the lease shows it]() | | While the Job runs | A running Job | Watched to the end | | After the Job finished, before bookkeeping | A finished Job, not bookkept | Bookkeeping reads its pod, once | | After bookkeeping, before the status patch persisted | The Job is not marked bookkept | The marker is written only after the patch, so the Job is read again | The last row is by design: the controller marks a Job bookkept only after the status patch that records its result succeeded, and counts it once (its per-run Secret is deleted once). A failure to mark is logged and retried. The new leader reconciles every object once as its caches fill, then on events and requeues. Nothing is lost by a failover. What it costs is time: up to the lease duration before a new leader acts, plus a reconcile of every object. ## Two leaders Election failure modes (a long pause, a partitioned leader) can briefly give two managers the idea they lead. The run lease is the answer: whichever creates or updates the lease first, using a `resourceVersion` it just read from the API server, wins; the other waits as `WaitingForRunLease` with the message “another manager took it first”. See [Leases](). > [!NOTE] > > **See also** > > - [Leases and the operation gate](). > - [Requeue intervals and schedules](). > - [Reconcile Errors](). # Reference # Custom Resources CAPTF adds seven kinds to the `infrastructure.cluster.x-k8s.io/v1alpha1` API group. Three run a Terraform or OpenTofu module, one per role of the [module contract](); three are templates Cluster API clones them from; and one holds the cloud credentials the others use. Each has its own page here: what it is, a minimal and a full YAML example, and every field of its spec and status, with its type, default, validation and what it does. - **TerraformCluster** --- The infrastructure around one workload cluster: runs the cluster role. - **TerraformMachine** --- The infrastructure of one `Machine`: runs the machine role. - **TerraformMachinePool** --- One native scaling group for a `MachinePool`: runs the machinepool role. - **TerraformClusterIdentity** --- Cloud credentials, and the namespaces allowed to use them. - **Templates** --- [TerraformClusterTemplate](), [TerraformMachineTemplate]() and [TerraformMachinePoolTemplate](): what Cluster API clones the three kinds above from. - **Common Fields** --- The source, identity, Job settings, variables and status fields every module-running kind shares. ## The kinds | Kind | Scope | Runs | Referenced by | | --- | --- | --- | --- | | [`TerraformCluster`]() | Namespaced | The cluster role | `Cluster.spec.infrastructureRef` | | [`TerraformMachine`]() | Namespaced | The machine role | `Machine.spec.infrastructureRef` | | [`TerraformMachinePool`]() | Namespaced | The machinepool role | `MachinePool.spec.template.spec.infrastructureRef` | | [`TerraformClusterTemplate`]() | Namespaced | Nothing; cloned into a `TerraformCluster` | `ClusterClass.spec.infrastructure.templateRef` | | [`TerraformMachineTemplate`]() | Namespaced | Nothing; cloned into each `TerraformMachine` | `MachineDeployment` and `MachineSet` `spec.template.spec.infrastructureRef`, the control plane’s machine template, and `ClusterClass` machine infrastructure `templateRef`s | | [`TerraformMachinePoolTemplate`]() | Namespaced | Nothing; cloned into a `TerraformMachinePool` | `ClusterClass.spec.workers.machinePools[].infrastructure.templateRef` | | [`TerraformClusterIdentity`]() | Cluster | Nothing; its Secret feeds the others’ Jobs | The others’ `spec.identityRef` | ``` flowchart LR Cluster --> TC[TerraformCluster] Machine --> TM[TerraformMachine] MachinePool --> TMP[TerraformMachinePool] TCT[TerraformClusterTemplate] -. cloned into .-> TC TMT[TerraformMachineTemplate] -. cloned into .-> TM TMPT[TerraformMachinePoolTemplate] -. cloned into .-> TMP TC --> ID[TerraformClusterIdentity] TM --> ID TMP --> ID ``` Every kind is in the `cluster-api` category, so `kubectl get cluster-api -A` lists them with the Cluster API objects. None has a short name. For how the kinds relate to Cluster API’s objects in practice, see [The Kinds](). ## Reading these pages - **Field paths** are written in full from the object’s root, such as `spec.drift.intervalSeconds`. An item of a list is written `[]`: `spec.variablesFrom[].secretRef.name`. - **Each field’s description** says what it does, then whether it is **Required**, its **Default** and where that default comes from (the admission webhook, the controller, or a [manager flag]()), whether it is **Immutable**, and its **Allowed values** or **Range**. - **Validation** sections list what the admission webhooks reject. The CRDs themselves declare no defaults; every default named here is applied by CAPTF. - **Shared fields.** `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` share their module source, identity, Job settings, variables and most of their status. Each kind’s page lists those fields and links to [Common Fields](), which covers them in full. These pages are kept in step with the provider’s API: a check fails when the CRDs gain or lose a field a page does not name (see [Writing Documentation]()). # TerraformCluster A `TerraformCluster` is the Cluster API infrastructure object of a CAPTF cluster. It runs the cluster-role module image you name in `spec.source` as a Job, keeps the module’s Terraform state in a Secret, and reports the module’s outputs (the control-plane endpoint, the failure domains and the exports that machines and pools inherit) in its spec and status. A `Cluster` references it through `Cluster.spec.infrastructureRef`, and Cluster API sets the owner reference from the `Cluster` to the `TerraformCluster`. You normally create one in one of two ways: a clusterctl template renders it next to the `Cluster`, or a ClusterClass topology creates it from a [`TerraformClusterTemplate`](). See [The kinds]() for how it relates to the other kinds. | Property | Value | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Scope | Namespaced | | Module role | `cluster` (see [Cluster role]()) | | Referenced by | `Cluster.spec.infrastructureRef` | | Finalizer | `terraformcluster.infrastructure.cluster.x-k8s.io` | | Short names | None | | Categories | `cluster-api` | | Status subresource | Yes | ## Example The smallest valid object sets a module image and an identity. A `TerraformCluster` does nothing until a `Cluster` references it. terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo namespace: team-a spec: source: image: ghcr.io/captf-io/noop-cluster:v0.1.0-opentofu identityRef: name: aws ``` The `Cluster` that uses it: cluster.yaml ```yaml apiVersion: cluster.x-k8s.io/v1beta2 kind: Cluster metadata: name: demo namespace: team-a spec: infrastructureRef: apiGroup: infrastructure.cluster.x-k8s.io kind: TerraformCluster name: demo ``` ## Full example This object sets every spec field. Real clusters rarely need all of them. terraformcluster.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: demo namespace: team-a spec: source: image: ghcr.io/captf-io/aws-cluster:v0.1.0-opentofu imagePullPolicy: IfNotPresent identityRef: name: aws jobs: activeDeadlineSeconds: 7200 # (1)! lockTimeoutSeconds: 600 variables: region: eu-west-1 vpc_cidr: 10.0.0.0/16 variablesFrom: - secretRef: name: demo-cluster-secrets controlPlaneEndpoint: # (2)! host: demo-api.example.com port: 6443 drift: intervalSeconds: 900 # (3)! action: Remediate applyPolicy: Manual # (4)! defaults: # (5)! identityRef: name: aws jobs: activeDeadlineSeconds: 3600 drift: intervalSeconds: 1800 ``` 1. `lockTimeoutSeconds` must stay below `activeDeadlineSeconds`. See [Validation](<#validation>). 2. Optional. Set it only when you own the endpoint (a fixed DNS name or VIP). Omit it and the controller copies the module’s `control_plane_endpoint` output here once. It is immutable after that. 3. This cluster is checked every 15 minutes and drift is reverted automatically. 4. Every change is planned first and waits for `captf.io/approve-plan`. 5. Inherited by this cluster’s machines and pools. It never applies to the `TerraformCluster` itself. ## Spec `spec` must set at least one property, and `spec.source` and `spec.identityRef` are required in practice (see [Validation](<#validation>)). The fields shared with the other Job-running kinds are documented on [Common fields](); the table links each one. | Field | Type | Description | | --- | --- | --- | | `spec.source` | object | The module image to run and how to pull it. See [Source](). **Required.** `spec.source.image` must be a valid image reference. | | `spec.identityRef` | object | The `TerraformClusterIdentity` that supplies cloud credentials to the Jobs. See [Identity reference](). **Required.** | | `spec.jobs` | object | Job policy: deadline, lock timeout, resources, service account, security context and history limits. See [Jobs](). **Mutable.** | | `spec.variables` | object | Inline module variables, a JSON object. See [Variables](). **Mutable.** | | `spec.variablesFrom` | array | ConfigMaps and Secrets that supply module variables, at most 16. See [Variable sources](). **Mutable.** | | `spec.controlPlaneEndpoint` | object | The API server endpoint. See [Control-plane endpoint](<#control-plane-endpoint>). | | `spec.drift` | object | Drift detection policy. See [Drift](<#drift>). **Mutable.** | | `spec.applyPolicy` | string | When a change is applied. See [Apply policy](<#apply-policy>). **Default:** `Automatic`. **Mutable.** | | `spec.defaults` | object | Values this cluster’s machines and pools inherit. See [Defaults](<#defaults>). **Mutable.** | ### Control-plane endpoint `spec.controlPlaneEndpoint` is the host and port of the workload cluster’s API server. A value you set is passed to the module as its `control_plane_endpoint` input, which tells the module to use it instead of creating a load balancer. When you leave it unset and the module outputs an endpoint, the controller writes that output here once and records `module` in the `captf.io/endpoint-source` annotation. Cluster API then copies the endpoint to `Cluster.spec.controlPlaneEndpoint`. The [module contract]() describes the input rules, including the fallback to `Cluster.spec.controlPlaneEndpoint`. | Field | Type | Description | | --- | --- | --- | | `spec.controlPlaneEndpoint` | object | Optional. **Default:** unset. The controller fills it from the module’s output when you do not. **Mutable** until it has a host and a port, then **Immutable.** | | `spec.controlPlaneEndpoint.host` | string | The API server hostname or IP address. **Range:** 1 to 512 characters. Set together with `port`. | | `spec.controlPlaneEndpoint.port` | integer | The API server port. **Range:** 1 to 65535. Set together with `host`. | > [!WARNING] > > **A complete endpoint cannot be changed or cleared** > > Every Machine and kubeconfig of the cluster points at the endpoint, so the webhook rejects any change once both `host` and `port` are set. To use a different endpoint, create a new cluster. ### Drift `spec.drift` controls the periodic check that compares the real infrastructure with the module. See [Drift and health]() for what runs and when, and [Drift]() for how to use it. | Field | Type | Description | | --- | --- | --- | | `spec.drift` | object | Optional. **Default:** unset; both fields below then take their defaults. **Mutable.** | | `spec.drift.intervalSeconds` | integer | Seconds between drift checks. **Default:** the manager’s `--drift-default-interval` (30 minutes), applied at reconcile; see [Manager flags](). **Range:** 0 or more. `0` disables drift checks, and with them every health sample after provisioning. **Mutable.** | | `spec.drift.action` | string | What happens when a check finds changes. **Default:** `Report`, applied at reconcile. **Allowed values:** `Report` (record the finding only), `Remediate` (apply the current inputs and revert out-of-band changes). **Mutable.** | > [!WARNING] > > **Disabling drift also stops health sampling** > > With `spec.drift.intervalSeconds: 0` the `InfrastructureHealthy` condition stops updating after provisioning, because health is read only from a completed refresh or drift Job. See [Disable drift checks](). ### Apply policy | Field | Type | Description | | --- | --- | --- | | `spec.applyPolicy` | string | **Default:** `Automatic`, applied at reconcile. **Allowed values:** `Automatic`, `Manual`. **Mutable.** | With `Automatic`, every change to the inputs, and every drift remediation, applies as soon as the controller sees it. Only a plan that deletes or replaces resources waits, for the `captf.io/approve-destructive-plan` annotation. With `Manual`, the controller runs a plan Job first, publishes the result in [`status.plan`](<#status>), and applies it only when the `captf.io/approve-plan` annotation names `status.plan.planHash`. The apply runs only if it plans exactly the same changes again. The first apply of a new cluster, which has no state yet, is never gated. See [Plan approval]() and [Manual approval](). ### Defaults `spec.defaults` holds values the cluster’s `TerraformMachine` and `TerraformMachinePool` objects inherit field by field: a value a machine or pool sets wins, an unset one comes from here. The values never apply to the `TerraformCluster` itself, and there is no `source` because every role names its own image. | Field | Type | Description | | --- | --- | --- | | `spec.defaults` | object | Optional. **Mutable.** | | `spec.defaults.identityRef` | object | The identity of machines and pools that set no `identityRef`. **Default:** `spec.identityRef`. Same shape as [Identity reference](). | | `spec.defaults.identityRef.name` | string | The name of the `TerraformClusterIdentity`. | | `spec.defaults.jobs` | object | A Job policy merged field by field under each machine’s or pool’s own `jobs`. Same type as `spec.jobs`; see [Jobs](). It is validated like `spec.jobs`. | | `spec.defaults.drift` | object | A drift policy merged field by field under each machine’s or pool’s own `drift`. It has no `action`: a machine always reports, and a pool sets its own. | | `spec.defaults.drift.intervalSeconds` | integer | **Default:** the manager’s `--drift-default-interval`. **Range:** 0 or more. `0` disables a machine’s drift checks but not a pool’s, which then uses the manager default. | See [Templates and ClusterClass]() and [Identities]() for the usual setup. ## Status Nothing in `status` is load-bearing: the controller rebuilds every value from the spec, the state Secret, the durable inputs Secret or the Job list, and `clusterctl move` does not carry status over. | Field | Type | Description | | --- | --- | --- | | `status.conditions` | array | Conditions of the object, at most 32, keyed by `type`. See [Conditions](<#conditions>). | | `status.conditions[].type` | string | The condition type, such as `Ready`. | | `status.conditions[].status` | string | `True`, `False` or `Unknown`. | | `status.conditions[].reason` | string | A CamelCase reason; see [Conditions](). | | `status.conditions[].message` | string | A human-readable detail. It can name a Job and plan counts but never a variable value. | | `status.conditions[].lastTransitionTime` | time | When `status` last changed. | | `status.conditions[].observedGeneration` | integer | The `metadata.generation` the condition was computed for. | | `status.initialization` | object | Cluster API’s initialization status; holds `provisioned`. See [Initialization](). | | `status.observedGeneration` | integer | The generation this status was computed for. See [Observed generation](). | | `status.activeJob` | object | The Job running for this object now, if any. See [Active job](). | | `status.lastRun` | object | The result of the most recent completed Job. See [Last run](). | | `status.lastDriftCheck` | time | When the last drift check completed. See [Drift checks and refreshes](). | | `status.lastRefresh` | time | When the last refresh or drift check completed. See [Drift checks and refreshes](). | | `status.pendingRefreshes` | integer | Consecutive refreshes that read health as pending. See [Drift checks and refreshes](). | | `status.observedStateSerial` | integer | The state serial the outputs were read from. See [State](). | | `status.lastRestoredSerial` | integer | The serial of the last restore Job consumed. See [State](). | | `status.stateSecretSuffix` | string | The backend `secret_suffix` of this object’s state. See [State](). | | `status.stateBackups` | array | The state backups kept, newest first, at most 16. See [State](). | | `status.source` | object | What the last Job actually ran. See [Image in use](). | | `status.failureDomains` | array | The failure domains from the module’s `failure_domains` output. Between 1 and 100 entries, keyed by `name`. Cluster API uses them to spread control-plane Machines. | | `status.failureDomains[].name` | string | **Required.** The failure domain name. **Range:** 1 to 256 characters. | | `status.failureDomains[].controlPlane` | boolean | Whether control-plane machines may use this domain. | | `status.failureDomains[].attributes` | map | Free-form string attributes the module reports for the domain. | | `status.plan` | object | The change waiting for approval under `spec.applyPolicy: Manual`. Empty when none waits. Approve it by setting `captf.io/approve-plan` to `status.plan.planHash`. | | `status.plan.inputsHash` | string | **Required** when a plan is set. The hash of the inputs the plan was made for. **Range:** 1 to 128 characters. | | `status.plan.job` | string | **Required** when a plan is set. The Job that made the plan: a plan Job, or an approved apply that found the plan changed. **Range:** 1 to 63 characters. | | `status.plan.planHash` | string | **Required** when a plan is set. A fingerprint of the plan’s changes, and the value that approves it. **Range:** 1 to 128 characters. | | `status.plan.add` | integer | Resources the plan creates. **Range:** 0 or more. | | `status.plan.change` | integer | Resources the plan updates in place. **Range:** 0 or more. | | `status.plan.destroy` | integer | Resources the plan destroys, counting replacements. **Range:** 0 or more. | | `status.plan.outputChanges` | integer | Root module outputs the plan changes. An output change alone needs approval, because cluster exports feed every machine and pool module. **Range:** 0 or more. | | `status.plan.resources` | array of strings | `
()` of each changed resource, sorted by address, at most 50, each 1 to 600 characters. The labels are the action (`create`, `update`, `delete`, `replace`, `read` or `forget`), or `import` or `move` for an otherwise unchanged resource, comma-separated: `aws_instance.a (import)`. Values are never shown. | | `status.plan.truncated` | boolean | `true` when `status.plan.resources` lists fewer resources than the plan changes. | | `status.plan.createdAt` | time | When the plan was made. | `status` must set at least one property when present. > [!NOTE] > > **Example status** > > ```yaml > status: > conditions: > - type: Ready > status: "True" > reason: Ready > observedGeneration: 3 > lastTransitionTime: "2026-10-02T09:14:07Z" > - type: ApplyJobSucceeded > status: Unknown > reason: PlanAwaitingApproval > message: Plan awaits approval; see status.plan > observedGeneration: 3 > lastTransitionTime: "2026-10-02T11:40:12Z" > - type: EndpointAvailable > status: "True" > reason: EndpointAvailable > observedGeneration: 3 > lastTransitionTime: "2026-10-02T09:14:07Z" > initialization: > provisioned: true > observedGeneration: 3 > failureDomains: > - name: eu-west-1a > controlPlane: true > - name: eu-west-1b > controlPlane: true > lastRun: > job: demo-plan-4f7c2 > operation: plan > lastRefresh: "2026-10-02T11:30:01Z" > source: > image: ghcr.io/captf-io/aws-cluster:v0.1.0-opentofu > plan: > inputsHash: 9a1c0f3e6d5b > job: demo-plan-4f7c2 > planHash: 7be2d41a0c93 > add: 1 > change: 2 > destroy: 0 > outputChanges: 0 > resources: > - aws_security_group_rule.api (create) > - aws_lb.api (update) > - aws_lb_listener.api (update) > truncated: false > createdAt: "2026-10-02T11:40:12Z" > ``` ## Conditions A `TerraformCluster` sets the conditions below. Each section of [Conditions]() lists every reason and what it means; this page does not copy them. | Type | `True` | `False` | `Unknown` | | --- | --- | --- | --- | | [`Ready`]() | Every input is healthy. Cluster API mirrors it into the `Cluster`’s `InfrastructureReady`. | An input is `False`. | An input is `Unknown`. | | [`DependenciesReady`]() | The `Cluster` and everything the object waits for exist. | The owner `Cluster` is missing or does not reference this object back, the `Cluster` references another provider, or a `variablesFrom` source is missing or invalid. | Still waiting for the owner. | | [`IdentityAllowed`]() | The identity exists and allows the namespace. | No identity is set, the identity is missing, does not allow the namespace or has no Secret. | The identity check could not be completed. | | [`CredentialsMirrored`]() | The identity’s Secret is copied for the Job. | It cannot be copied. | The mirror is pending: it has not been made yet, or no identity resolved. | | [`RunnerRBACReady`]() | The runner’s ServiceAccount and RBAC exist. | They cannot be created, or an override ServiceAccount lacks the `captf.io/runner=true` label. | Never `Unknown`. Until the controller first reaches it, the condition is absent and counts as `Unknown` in `Ready`. | | [`ApplyJobSucceeded`]() | The last apply or destroy succeeded. | It failed, hit its deadline or could not start, or a destructive plan is blocked. | No apply has completed yet, the run waits for a lease or for the cluster’s machines and pools, or a plan awaits approval or changed. A running apply keeps the last result. | | [`StateReadable`]() | The Terraform state can be read. | It is encrypted, corrupt, inconsistent, lost or locked by another holder. | No state exists yet. | | [`RestoreJobSucceeded`]() | The last state restore succeeded. | It failed, or the requested backup does not exist. | The restore waits for a run lease or for other operations. Set only once a restore is requested with `captf.io/restore-state`. | | [`OutputsValid`]() | The module’s outputs satisfy the contract. | An output is missing or invalid. | Required outputs are still `null`. | | [`InfrastructureHealthy`]() | The module’s health output reads healthy. | It is provisioning, or it reads pending, unhealthy, degraded, stopped or terminated. | Before the first apply, or health is unknown. | | [`DriftJobSucceeded`]() | The last drift or refresh Job succeeded. | It failed or hit its deadline. | No check has completed, a Job is running or waits for a lease, or the durable inputs Secret is missing. | | [`DriftDetected`]() | The last drift check found changes: reported, pending remediation or being remediated. | It found none. | No check has completed. | | [`DeletionBlocked`]() | Machines or pools of the cluster still exist, so the destroy waits. | Nothing blocks the destroy. | Never `Unknown`. | | [`EndpointAvailable`]() | A valid control-plane endpoint is known. | The cluster is provisioned but has no endpoint. | Never `Unknown`; the condition is not set before provisioning. | | [`Paused`]() | The object or its `Cluster` is paused. | It is not paused. | Never `Unknown`. | | [`Deleting`]() | The object is being deleted. It makes `Ready` `False`. | It is not. | Never `Unknown`. | `DriftDetected`, `DeletionBlocked` and `DriftJobSucceeded` have negative or informational meaning and never feed `Ready`; `DeletionBlocked` and `DriftDetected` read `True` when something needs attention. After provisioning, `Ready` summarizes only `InfrastructureHealthy` and `Deleting`, so a failed re-apply does not flip the `Cluster`’s `InfrastructureReady` and suspend its MachineHealthChecks. See [Ready summarization](). ## Printer columns `kubectl get terraformclusters` shows: | Column | Source | Notes | | --- | --- | --- | | `CLUSTER` | The `cluster.x-k8s.io/cluster-name` label | The owning `Cluster`. | | `READY` | `status.conditions` entry of type `Ready` | Its `status`. | | `PROVISIONED` | `status.initialization.provisioned` | | | `ENDPOINT` | `spec.controlPlaneEndpoint.host` | | | `IMAGE` | `spec.source.image` | Only with `-o wide`. | | `AGE` | `metadata.creationTimestamp` | | ## Validation The validating webhook (`validation.terraformcluster.infrastructure.cluster.x-k8s.io`) and the CRD schema enforce these rules on create and update. The webhook reports every violation it finds at once. **Required fields** - `spec` must set at least one property. - `spec.source.image` must be set and a valid image reference. The image’s content and registry are not checked. - `spec.identityRef.name` must be set. `spec.defaults.identityRef` is not a substitute: it only applies to machines and pools. **Immutable fields** - `spec.controlPlaneEndpoint` is mutable until it has both a host and a port. After that, any change or removal is rejected. A half-set endpoint stored earlier stays completable, and unrelated updates to it are not rejected. - Every other field is mutable. A new `spec.source.image` is a new module version, applied against the existing state. **Cross-field and range rules** - `spec.controlPlaneEndpoint` must set both `host` and `port`, or neither. `host` is 1 to 512 characters and `port` is 1 to 65535. - `spec.applyPolicy` must be `Automatic` or `Manual`. - `spec.drift.action` must be `Report` or `Remediate`. `spec.drift.intervalSeconds` and `spec.defaults.drift.intervalSeconds` must be 0 or more. - `spec.jobs` and `spec.defaults.jobs` are validated only when they change, so an older policy never blocks a finalizer patch: - The container and pod security contexts may not weaken the hardened defaults: no privileged mode, privilege escalation, added capabilities, writable root filesystem, unmasked `/proc`, `runAsNonRoot: false`, `runAsUser: 0`, `Unconfined` seccomp profile or Windows host process. - `lockTimeoutSeconds` must be less than `activeDeadlineSeconds`. When only one is set, it is compared with the other’s built-in default (a 300 second lock timeout, a 3600 second deadline). - `spec.variables` must be a JSON object of at most the contract’s maximum number of keys. Each key must be a Terraform identifier that is not a `captf_` name, a cluster-role contract input or a module meta-argument. - Each `spec.variablesFrom[]` entry must set exactly one of `configMapRef` and `secretRef`; there are at most 16 entries. Keys in the referenced sources are checked at reconcile, not at admission. - `status.conditions` holds at most 32 entries and `status.failureDomains` at most 100. **Defaulting.** The CRD declares no defaults. The controller applies `Automatic` for `spec.applyPolicy`, `Report` for `spec.drift.action` and the manager’s `--drift-default-interval` for `spec.drift.intervalSeconds` at reconcile; the stored object stays as you wrote it. Deleting a `TerraformCluster` is always admitted. On a deleting object the Job policy rules are skipped, so its finalizer stays removable. ## Lifecycle On create, the controller waits for the owning `Cluster`, resolves the identity, mirrors credentials, prepares the runner’s RBAC and runs an apply Job. After the apply it validates the module’s outputs, writes the endpoint and failure domains, and sets `status.initialization.provisioned`. Machines and pools wait for the cluster’s apply to finish. See [The reconcile lifecycle]() and [Job inputs](). On change, the controller compares the rendered inputs with the last applied ones and runs a new apply. A plan that deletes or replaces resources waits for approval, and under `Manual` every change does. See [Approvals](). Between changes it runs refresh and drift Jobs on the interval in `spec.drift`; see [Drift and health](). On delete, the controller holds the destroy while any machine or pool of the cluster still exists (`DeletionBlocked`), then runs a destroy Job and removes the finalizer. A missing or unreadable state, or a failed destroy, holds the deletion until you restore the state or abandon the object. See [Deletion]() and [Deletion order](). > [!NOTE] > > **See also** > > - [Common fields]() > - [TerraformClusterTemplate]() > - [TerraformMachine]() > - [TerraformMachinePool]() > - [TerraformClusterIdentity]() > - [Conditions]() > - [Annotations and labels]() > - [Manager flags]() > - [Cluster role contract]() # TerraformClusterTemplate A `TerraformClusterTemplate` holds the metadata and spec of the [`TerraformCluster`]() that a ClusterClass creates. It has no status and runs nothing itself. A `ClusterClass` references it in `spec.infrastructure.templateRef`. When you create a `Cluster` with `spec.topology`, Cluster API clones `spec.template` into a `TerraformCluster`, applies the class’s patches to the clone, and points the `Cluster`’s `spec.infrastructureRef` at it. You create a template once per class, and every `Cluster` on the class gets its own `TerraformCluster`. Without a ClusterClass you create the `TerraformCluster` directly and never need a template. See [Templates and ClusterClass](). | Property | Value | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Scope | Namespaced | | Module role | `cluster`, through the `TerraformCluster` it creates | | Referenced by | `ClusterClass.spec.infrastructure.templateRef` | | Finalizer | None | | Short names | None | | Categories | `cluster-api` | | Status subresource | No | ## Example terraformclustertemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterTemplate metadata: name: noop namespace: team-a spec: template: metadata: labels: team: platform spec: source: image: ghcr.io/captf-io/noop-cluster:v0.1.0-opentofu identityRef: name: aws drift: intervalSeconds: 900 ``` A ClusterClass often leaves `identityRef` out and sets it, and the image, with a patch. The template then triggers a warning on create, not an error; see [Validation](<#validation>). ## Spec `spec.template` is required, and it must set at least one property. | Field | Type | Description | | --- | --- | --- | | `spec.template` | object | **Required.** The `TerraformCluster` to create. **Immutable** in its `spec`; see [Validation](<#validation>). | | `spec.template.metadata` | object | Optional. Metadata copied onto the created `TerraformCluster`. **Mutable.** | | `spec.template.metadata.labels` | map | Labels copied onto the created object. Keys and values must be valid Kubernetes labels. | | `spec.template.metadata.annotations` | map | Annotations copied onto the created object. Keys must be valid annotation keys. | | `spec.template.spec` | object | **Required.** The spec of the created `TerraformCluster`: exactly the [`TerraformCluster` spec](). **Immutable.** | | `spec.template.spec.source` | object | The module image. **Required** by the webhook. See [Source](). | | `spec.template.spec.identityRef` | object | The identity for the Jobs. Optional here; a patch or the `Cluster` can set it, but the created object is rejected without one. See [Identity reference](). | | `spec.template.spec.jobs` | object | Job policy. See [Jobs](). | | `spec.template.spec.variables` | object | Inline module variables. See [Variables](). | | `spec.template.spec.variablesFrom` | array | Variable sources. See [Variable sources](). | | `spec.template.spec.controlPlaneEndpoint` | object | The API server endpoint. See [Control-plane endpoint](). | | `spec.template.spec.drift` | object | Drift policy. See [Drift](). | | `spec.template.spec.applyPolicy` | string | `Automatic` or `Manual`. See [Apply policy](). | | `spec.template.spec.defaults` | object | Values the cluster’s machines and pools inherit. See [Defaults](). | Defaults are the same as for a `TerraformCluster`: the CRD declares none, and the controller applies them to the created object at reconcile. ## Printer columns `kubectl get terraformclustertemplates` shows: | Column | Source | | --- | --- | | `IMAGE` | `spec.template.spec.source.image` | | `AGE` | `metadata.creationTimestamp` | ## Validation The validating webhook (`validation.terraformclustertemplate.infrastructure.cluster.x-k8s.io`) checks every create and update. - `spec.template.metadata` must hold valid labels and annotations. - `spec.template.spec` follows the `TerraformCluster` rules for `source`, `jobs`, `defaults.jobs`, `variables` and `variablesFrom`; see [Validation](). The endpoint, drift and apply policy rules come from the CRD schema. - `spec.template.spec.identityRef` is not required. A template without one is admitted with the warning “spec.template.spec sets no identityRef; TerraformClusters created from this template are rejected unless a patch sets one”. - `spec.template.spec` is immutable. Any update that changes it is rejected with “TerraformClusterTemplate spec.template.spec is immutable; create a new TerraformClusterTemplate instead”. Only `spec.template.metadata` can change. - A ClusterClass topology dry-run is exempt from the immutability check, so Cluster API can validate patches that touch `spec.template.spec`. - A jobs policy is checked on an update only when it changed, and not at all while the template is deleting. - Deleting a template is always admitted. > [!WARNING] > > **Change a template by creating a new one** > > To change a field of `spec.template.spec`, create a template under a new name and point `ClusterClass.spec.infrastructure.templateRef` at it. See [Template immutability and rolling out a change](). > [!NOTE] > > **See also** > > - [TerraformCluster]() > - [Common fields]() > - [TerraformMachineTemplate]() > - [Templates and ClusterClass]() > - [The kinds]() # TerraformMachine A `TerraformMachine` is one node instance, provisioned by running a machine-role module image as a Kubernetes Job. It is the infrastructure machine object of Cluster API: a `Machine` points at it through `Machine.spec.infrastructureRef`, and Cluster API mirrors its `Ready` condition into the Machine’s `InfrastructureReady` condition. You rarely write one by hand. Cluster API creates it by cloning `spec.template` of a [`TerraformMachineTemplate`]() when a MachineDeployment, a MachineSet or a KubeadmControlPlane adds a Machine. The manager then runs the module image of the `machine` role (see the [machine role contract]()), keeps its Terraform state in a Secret, and reads the instance’s provider ID, addresses, failure domain and health from the module’s outputs. | | | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Kind | `TerraformMachine` | | Scope | Namespaced | | Module role | `machine` | | Created by | Cluster API, from a [`TerraformMachineTemplate`]() | | Referenced by | `Machine.spec.infrastructureRef` | | Finalizer | `terraformmachine.infrastructure.cluster.x-k8s.io` | | Short names | none | | Categories | `cluster-api` | | Status subresource | yes | ## Example The smallest valid object sets only `spec.source.image`. The manager needs an identity too, but a machine falls back to the one its `TerraformCluster` names, so a machine in a working cluster does not set it. terraformmachine.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachine metadata: name: demo-md-0-x7k2p namespace: default labels: cluster.x-k8s.io/cluster-name: demo spec: source: image: ghcr.io/captf-io/noop-machine:v0.1.0-opentofu ``` The `cluster.x-k8s.io/cluster-name` label and the owner reference to the `Machine` come from Cluster API. Without them the manager has no cluster to resolve and reports `DependenciesReady=False` or `Unknown` (see [Conditions](<#conditions>)). ## Full example Every spec field set. The `spec.providerID` line is shown for completeness: the controller writes it, you do not. terraformmachine-full.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachine metadata: name: demo-md-0-x7k2p namespace: default labels: cluster.x-k8s.io/cluster-name: demo spec: providerID: aws:///us-east-1a/i-0abc123def4567890 # (1)! source: image: ghcr.io/captf-io/aws-machine:v0.1.0-opentofu # (2)! identityRef: name: aws-prod # (3)! jobs: activeDeadlineSeconds: 3600 lockTimeoutSeconds: 300 variables: instance_type: m6i.large # (4)! variablesFrom: - secretRef: name: demo-machine-secrets # (5)! drift: intervalSeconds: 900 # (6)! remediation: annotateMachine: true # (7)! unhealthyThreshold: 5 healthCheckIntervalSeconds: 120 ``` 1. Written once by the controller from the module’s `provider_id` output. You never set it. See [Provider ID](<#provider-id>). 2. Required, and immutable: a different image means a new machine. 3. Optional when the owning `TerraformCluster` sets `spec.defaults.identityRef` or its own `spec.identityRef`. Immutable. 4. Inline module variables, a JSON object. Immutable, like every field that defines the machine. 5. A Secret in this namespace labeled `captf.io/variables=true`. Inline `variables` win over it on the same key. 6. Check for drift every 15 minutes. Without it, the cluster’s `spec.defaults.drift.intervalSeconds`, then the manager’s `--drift-default-interval`, applies. 7. Ask Cluster API to replace the Machine after 5 unhealthy samples. Needs a `MachineHealthCheck`; see [Remediation](<#remediation>). ## Spec `spec` must set at least one property, and `spec.source` is always required. `spec.source`, `spec.identityRef`, `spec.variables` and `spec.variablesFrom` define the machine and are immutable after creation. `spec.providerID` is set once. `spec.jobs`, `spec.drift` and `spec.remediation` are operational policy and stay mutable, so you can change a stuck machine’s deadline or its drift checks without replacing it. The workspace fields below are shared with other kinds and documented on [Common Fields](). | Field | Type | Description | | --- | --- | --- | | `spec.providerID` | string | The instance’s provider ID, set by the controller. See [Provider ID](<#provider-id>). **Optional.** **Immutable** once set. **Range:** 1 to 512 characters. | | `spec.source` | object | The machine-role module image and its pull policy ([Source]()). **Required.** **Immutable.** Never inherited from the cluster. | | `spec.identityRef` | object | The `TerraformClusterIdentity` whose credentials the Jobs use ([Identity reference]()). **Optional.** **Immutable.** **Default:** the owning `TerraformCluster`’s `spec.defaults.identityRef`, else its `spec.identityRef`. | | `spec.jobs` | object | Tuning of the Jobs that run the module ([Jobs]()). **Optional.** **Mutable.** **Default:** merged field by field over the `TerraformCluster`’s `spec.defaults.jobs`; fields neither sets take the built-in defaults. | | `spec.variables` | object | Inline module variables, a JSON object ([Variables]()). **Optional.** **Immutable.** | | `spec.variablesFrom` | list | ConfigMaps and Secrets that supply module variables ([Variable sources]()). **Optional.** **Immutable.** | | `spec.drift` | object | How often drift is checked. See [Drift](<#drift>). **Optional.** **Mutable.** **Default:** merged over the `TerraformCluster`’s `spec.defaults.drift`. | | `spec.remediation` | object | How an unhealthy instance is signaled to Cluster API. See [Remediation](<#remediation>). **Optional.** **Mutable.** | ### Provider ID `spec.providerID` binds the object to one instance. After the first apply succeeds, the controller copies the module’s `provider_id` output into it and emits a `ProviderIDSet` event. Cluster API then copies it to `Machine.spec.providerID`, where it must equal the Node’s `spec.providerID` (see [Node providerID matching]()). - The controller writes it only while it is empty and never clears it. If the instance disappears, `provider_id` turning `null` leaves the value in place and sets `InfrastructureHealthy` to `Unknown` (`ProviderIDMissing`), then to `False` (`InstanceTerminated`) if the next sample is `null` too. - If a later apply returns a different value, the controller keeps the old one and sets `OutputsValid=False` with reason `ProviderIDChanged`. - A state with no inputs hash, from an apply that has not succeeded, does not set it, because the retry may replace a tainted instance. > [!WARNING] > > **Only the manager can set providerID on an existing object** > > The webhook rejects any update that sets `spec.providerID` from a user other than the manager’s ServiceAccount, and any change after it is set. Create accepts a value because `clusterctl move` recreates objects with theirs. ### Drift `spec.drift` configures the periodic drift check, a refresh and plan that compares the real instance with the machine’s inputs. A machine’s drift is always reported (`DriftDetected`) and never remediated: the instance is immutable infrastructure, replaced by a rollout rather than patched, so there is no `action` field. See [Drift]() and [Drift and health](). | Field | Type | Description | | --- | --- | --- | | `spec.drift.intervalSeconds` | integer (int32) | Seconds between drift checks. **Optional.** **Mutable.** **Range:** 0 or more; 0 disables drift checks. **Default:** the `TerraformCluster`’s `spec.defaults.drift.intervalSeconds`, else the manager’s `--drift-default-interval` (30 minutes), applied at reconcile. | > [!NOTE] > > **Disabling drift also stops health sampling** > > Health is read from the same refresh Jobs. With `intervalSeconds: 0` and `spec.remediation.annotateMachine` not `true`, nothing refreshes the machine after provisioning, so `status.unhealthySamples` and `InfrastructureHealthy` stop updating. With `annotateMachine: true` the machine still refreshes at `healthCheckIntervalSeconds`. ### Remediation `spec.remediation` lets CAPTF ask Cluster API to replace a bad instance. With `annotateMachine: true`, CAPTF sets the `cluster.x-k8s.io/remediate-machine` annotation on the owner Machine when the instance has been unhealthy for `unhealthyThreshold` consecutive samples, or at once when it is terminated. A `MachineHealthCheck` that selects the Machine then acts on it. CAPTF also sets `captf.io/remediation-requested` (value: the reason) so it removes only its own annotation, which it does when the instance reads `Healthy` again and the Machine is not being deleted. An annotation someone else set is left alone. A single-replica control plane refuses the remediation. See [Machine Remediation](). | Field | Type | Description | | --- | --- | --- | | `spec.remediation.annotateMachine` | boolean | Annotate the owner Machine for remediation. **Optional.** **Mutable.** **Default:** `false`: CAPTF never touches the Machine. | | `spec.remediation.unhealthyThreshold` | integer (int32) | Consecutive unhealthy samples before the Machine is annotated. A sample is one completed refresh or drift Job; a terminated instance counts on the first sample. **Optional.** **Mutable.** **Range:** 1 to 100. **Default:** 3, applied at reconcile. | | `spec.remediation.healthCheckIntervalSeconds` | integer (int32) | How often a provisioned machine is refreshed to sample health while `annotateMachine` is `true`, independent of `spec.drift.intervalSeconds`. Ignored when `annotateMachine` is `false`. **Optional.** **Mutable.** **Range:** 60 to 86400. **Default:** 300, applied at reconcile. | An unhealthy sample is a reading of `InstanceUnhealthy`, `InstanceDegraded` or `InstanceStopped`. `Healthy` resets the count; `InstancePending`, `HealthUnknown` and `InstanceTerminated` leave it unchanged (see [Unhealthy samples]()). ## Status Nothing in `status` is load-bearing: the controller rebuilds every value from spec, the state Secret and the Job list, so it survives `clusterctl move`. `status.unhealthySamples` restarts at 0 after a move. Shared workspace status fields are documented on [Common Fields](). | Field | Type | Description | | --- | --- | --- | | `status.conditions` | list | The object’s conditions, keyed by `type`. **Range:** at most 32. See [Conditions](<#conditions>). | | `status.conditions[].type` | string | The condition type, such as `Ready` or `InfrastructureHealthy`. | | `status.conditions[].status` | string | `True`, `False` or `Unknown`. | | `status.conditions[].reason` | string | A CamelCase reason for the last change. Every reason is on [Conditions](). | | `status.conditions[].message` | string | A human-readable detail. | | `status.conditions[].lastTransitionTime` | time | When `status` last changed. | | `status.conditions[].observedGeneration` | integer | The `metadata.generation` the condition was computed for. | | `status.observedGeneration` | integer | The `metadata.generation` the controller last reconciled ([Observed generation]()). | | `status.initialization` | object | `provisioned` is true once the infrastructure is provisioned, then latched for the object’s life ([Initialization]()). The `Provisioned` printer column shows it. | | `status.activeJob` | object | The Job running now, if any ([Active Job]()). | | `status.lastRun` | object | The newest finished Job and its result ([Last run]()). | | `status.lastDriftCheck` | time | When the last drift check completed ([Drift checks and refreshes]()). | | `status.lastRefresh` | time | When the last refresh or drift Job completed, or the apply that stood in for the refresh ([Drift checks and refreshes]()). | | `status.pendingRefreshes` | integer | Consecutive health samples that read pending; spaces the refreshes while the instance is pending ([Drift checks and refreshes]()). | | `status.stateSecretSuffix` | string | The suffix of the state Secret ([State]()). | | `status.observedStateSerial` | integer | The Terraform state serial the outputs were read from ([State]()). | | `status.lastRestoredSerial` | integer | The serial of the last state restore ([State]()). | | `status.stateBackups` | list | Backups of the state available to restore ([State]()). | | `status.source` | object | What the last Job ran: the image reference, the digest it resolved to and the runtime version ([Image in use]()). | | `status.addresses` | list | The instance’s addresses, from the module’s `addresses` output, in the controller’s canonical order. **Range:** 1 to 256 items. | | `status.addresses[].type` | string | The Cluster API address type: `Hostname`, `ExternalIP`, `InternalIP`, `ExternalDNS` or `InternalDNS`. | | `status.addresses[].address` | string | The address itself. | | `status.failureDomain` | string | The failure domain the instance actually runs in, from the module’s `failure_domain` output. Must equal `Machine.spec.failureDomain` when that is set, else `OutputsValid=False` (`FailureDomainMismatch`). **Range:** 1 to 256 characters. | | `status.interruptible` | boolean | True for a spot or preemptible instance; Cluster API then labels the Node `cluster.x-k8s.io/interruptible`. Written as `false` when the module does not say. | | `status.unhealthySamples` | integer (int32) | Consecutive unhealthy health samples. Unset while the instance is healthy. **Range:** 1 or more when set. Drives [remediation](<#remediation>). | If a module output breaks the contract (a bad `addresses`, `failure_domain` or `interruptible`), the controller keeps the previous status value for that field and sets `OutputsValid=False`. > [!NOTE] > > **Example status** > > ```yaml > status: > observedGeneration: 2 > initialization: > provisioned: true > addresses: > - type: InternalIP > address: 10.0.12.34 > - type: InternalDNS > address: ip-10-0-12-34.ec2.internal > failureDomain: us-east-1a > interruptible: false > lastRefresh: "2026-10-02T09:41:07Z" > lastDriftCheck: "2026-10-02T09:41:07Z" > conditions: > - type: Ready > status: "True" > reason: Ready > lastTransitionTime: "2026-10-02T08:12:55Z" > observedGeneration: 2 > - type: InfrastructureHealthy > status: "True" > reason: Healthy > lastTransitionTime: "2026-10-02T08:12:55Z" > observedGeneration: 2 > - type: DriftDetected > status: "False" > reason: NoDrift > lastTransitionTime: "2026-10-02T09:41:07Z" > observedGeneration: 2 > ``` ## Conditions `Ready` is the only condition Cluster API reads. The others say why `Ready` is what it is. Every reason is listed on [Conditions](); this page covers what each type means for a machine. | Type | Meaning | | --- | --- | | [`Ready`]() | `True` when the instance is provisioned and healthy. Before `status.initialization.provisioned` first holds, it summarizes `DependenciesReady`, `IdentityAllowed`, `CredentialsMirrored`, `RunnerRBACReady`, `ApplyJobSucceeded`, `StateReadable`, `OutputsValid`, `InfrastructureHealthy` and `Deleting`. After that it follows `InfrastructureHealthy` and `Deleting` only, so a failed re-apply or drift Job does not flip the Machine’s `InfrastructureReady`. `False` when an input is `False`, `Unknown` when an input is `Unknown`. | | [`DependenciesReady`]() | `True` when the owner Machine and Cluster, the `TerraformCluster` exports, the bootstrap data Secret and the variable sources are all available. `Unknown` while it waits for one of them; `False` when the owner is missing, mismatched or not yet set, the Cluster is not a `TerraformCluster`, or a variable source is missing or invalid. | | [`IdentityAllowed`]() | `True` when the resolved identity exists and allows this namespace. `False` when no identity is set, it is missing, not allowed or has no Secret. `Unknown` when the check could not be completed. | | [`CredentialsMirrored`]() | `True` when the identity’s credentials are copied into the namespace for the Jobs. `False` when the copy failed; `Unknown` while the mirror is pending. | | [`RunnerRBACReady`]() | `True` when the Job’s ServiceAccount and RBAC are in place. `False` when they cannot be created or an override ServiceAccount lacks the `captf.io/runner=true` label. Never `Unknown`. | | [`ApplyJobSucceeded`]() | The outcome of the newest apply or destroy. `Unknown` before the first apply completes or while the run waits for a lease or the cluster; a running apply keeps the last result. | | [`StateReadable`]() | `True` when the Terraform state Secret reads cleanly. `False` when it is encrypted, corrupt, inconsistent, lost or locked. `Unknown` when no state exists yet. | | [`RestoreJobSucceeded`]() | The outcome of a state restore requested with `captf.io/restore-state`. Set only once a restore is requested; `Unknown` while the restore waits for a run lease. Never feeds `Ready`. | | [`OutputsValid`]() | `True` when the module’s outputs satisfy the contract. `False` for missing or invalid outputs, a failure domain mismatch or a changed provider ID. `Unknown` while required outputs are `null`. | | [`InfrastructureHealthy`]() | The instance’s health from the module’s `health` output. `False` while provisioning, pending, unhealthy, degraded, stopped or terminated; `Unknown` before the first apply, when the module reports unknown, or when `provider_id` first turns `null`. | | [`DriftJobSucceeded`]() | The outcome of the newest refresh or drift Job. `Unknown` before the first check, while a Job runs or waits for a lease, or when the durable inputs Secret is missing. Never feeds `Ready`. | | [`DriftDetected`]() | `True` when a drift check found changes (reported, pending remediation or being remediated), `False` when it found none, `Unknown` before the first check. Never feeds `Ready`. | | [`Paused`]() | `True` when the object or its Cluster is paused; no Job starts. `False` otherwise; never `Unknown`. Never feeds `Ready`. | | [`Deleting`]() | `True` once deletion has started, `False` before; never `Unknown`. Negative polarity: `True` makes `Ready` `False`. | ## Printer columns `kubectl get terraformmachines` shows these columns. `-o wide` adds `Image`. | Column | Source | Description | | --- | --- | --- | | `Cluster` | `.metadata.labels['cluster.x-k8s.io/cluster-name']` | The owning Cluster. | | `Machine` | the `Machine` entry in `.metadata.ownerReferences` | The Machine that owns this object. | | `Ready` | the `Ready` condition’s `status` | `True`, `False` or `Unknown`. | | `Provisioned` | `.status.initialization.provisioned` | Whether the infrastructure is provisioned. | | `ProviderID` | `.spec.providerID` | The instance’s provider ID. | | `Image` (wide) | `.spec.source.image` | The module image. | | `Age` | `.metadata.creationTimestamp` | Time since creation. | ## Validation The admission webhook enforces these rules on create, update and delete. A rejected request lists every violation it found. The CRD schema declares no defaults; defaults come from the controller, the cluster and the manager flags as noted in the field tables. - `spec.source.image` is required and must parse as an image reference. The registry and content are not checked. Checked on create and update. - `spec.source`, `spec.identityRef`, `spec.variables` and `spec.variablesFrom` are immutable. `variables` compares by value, so reordered keys or whitespace are no change. Checked on update. - `spec.providerID` can change only from empty to non-empty, and only by the manager’s ServiceAccount. Create accepts any value. Checked on update. - `spec.jobs`: neither the container nor the pod security context may weaken the hardened defaults (privileged, privilege escalation, added capabilities, a writable root filesystem, `procMount: Unmasked`, running as root, an `Unconfined` seccomp profile, a Windows host process). Checked on create, and update when `spec.jobs` changed. - `spec.jobs.lockTimeoutSeconds` must be less than `activeDeadlineSeconds`. A lone value is compared with the other’s built-in default. A merge with the cluster’s defaults is checked at reconcile, not here. Checked on create, and update when `spec.jobs` changed. - `spec.variables` is a JSON object of at most 256 keys. Each key is a Terraform identifier that is not a `captf_` name, a `machine` role contract input or a module meta-argument. Checked on create and update. - Each `spec.variablesFrom` entry sets exactly one of `configMapRef` and `secretRef`. Checked on create and update. - `spec` must set at least one property. `spec.providerID` is 1 to 512 characters. `spec.drift.intervalSeconds` is 0 or more. `spec.remediation.unhealthyThreshold` is 1 to 100. `spec.remediation.healthCheckIntervalSeconds` is 60 to 86400. Checked on create and update. - A delete is refused while a Machine that is not itself being deleted references the object through `spec.infrastructureRef`. Delete the Machine instead. Checked on delete. Metadata changes are always allowed: KubeadmControlPlane syncs labels and annotations on every reconcile, and `clusterctl move` annotates an object before it deletes it. A `jobs` policy stored before a rule existed does not block later updates, because it is checked only when it changes, and never on an object that is being deleted. The delete rule has two exceptions. An object that no Machine references (for example, its Machine is gone) can always be deleted. An object with the `clusterctl move` delete annotation can be deleted while its Cluster is paused. The annotation alone is not enough. ## Lifecycle **Create.** Cluster API clones the template and creates the object with the `cluster.x-k8s.io/cluster-name` label and an owner reference to the Machine. The manager adds the finalizer, resolves the Cluster and `TerraformCluster`, waits for the cluster’s infrastructure, its exports and the bootstrap data Secret, then starts an apply Job. Once the state carries a successful apply, the module’s `provider_id` is non-null and its health is not `pending`, the machine is provisioned. See [The reconcile lifecycle]() and [Job inputs](). **Change.** The machine definition is immutable, and the first apply’s inputs are pinned in a durable Secret: changing the cluster later does not re-apply the machine. To change an instance, change the template and roll the Machine through its MachineDeployment or KubeadmControlPlane. Policy fields (`jobs`, `drift`, `remediation`) can change at any time. Health and drift are sampled on the timers above. **Delete.** Delete the Machine, not the `TerraformMachine`. Cluster API drains the node and runs the lifecycle hooks, then deletes the infrastructure object. The manager runs a destroy Job from the durable inputs, so it works even when the Machine, the bootstrap Secret or the Cluster is already gone, then removes the state and the finalizer. A `TerraformCluster` waits for its machines to go before its own destroy. See [Deletion order](). If state is missing or unreadable, deletion is held until you set `captf.io/abandon-infrastructure` (see [Annotations and labels]()). > [!NOTE] > > **See also** > > - [Common Fields]() for the shared workspace fields. > - [TerraformMachineTemplate]() for how Cluster API creates machines. > - [TerraformCluster]() for `spec.defaults`, which a machine inherits. > - [TerraformClusterIdentity]() for the credentials an identity holds. > - [The Kinds]() > - [Drift and Health]() > - [Machine Remediation]() and [Drift]() > - [Machine role contract]() > - [Conditions]() and [Annotations and labels]() # TerraformMachineTemplate A `TerraformMachineTemplate` is the blueprint Cluster API clones into one [`TerraformMachine`]() per `Machine`. `spec.template` holds the metadata and spec each clone is created with. Three kinds of owner reference it: - a MachineDeployment or MachineSet, through `spec.template.spec.infrastructureRef` of the Machine template; - a KubeadmControlPlane (or another control-plane provider), through `spec.machineTemplate.infrastructureRef`; - a ClusterClass, as the infrastructure template of a control plane or a worker class. Unlike the other template kinds, a `TerraformMachineTemplate` has a status. The manager reads the module image’s labels and records the node size and platform the image declares in `status.capacity` and `status.nodeInfo`, so the Cluster Autoscaler can scale a MachineDeployment up from zero replicas without a running node to measure. | | | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Kind | `TerraformMachineTemplate` | | Scope | Namespaced | | Module role | `machine`, through the machines it creates | | Created by | You (or a cluster template or ClusterClass you apply) | | Referenced by | MachineDeployment, MachineSet, KubeadmControlPlane and ClusterClass `infrastructureRef` | | Finalizer | none | | Short names | none | | Categories | `cluster-api` | | Status subresource | yes | ## Example terraformmachinetemplate.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachineTemplate metadata: name: demo-md-0 namespace: default spec: template: metadata: labels: team: platform spec: source: image: ghcr.io/captf-io/aws-machine:v0.1.0-opentofu identityRef: name: aws-prod remediation: annotateMachine: true ``` `spec.template.spec` is the only required part, and inside it only `source.image`. Everything under it is a [`TerraformMachine` spec]() and means the same there. ## Spec `spec` holds only `spec.template`. | Field | Type | Description | | --- | --- | --- | | `spec.template` | object | The `TerraformMachine` created from this template. **Required.** Must set at least one property. | | `spec.template.metadata` | object | Metadata copied onto each created `TerraformMachine`. **Optional.** **Mutable.** | | `spec.template.metadata.labels` | map of string | Labels copied onto each machine. Keys and values must be valid Kubernetes labels. **Optional.** | | `spec.template.metadata.annotations` | map of string | Annotations copied onto each machine. Keys must be valid annotation keys. **Optional.** | | `spec.template.spec` | object | The spec of each created `TerraformMachine`. **Required.** **Immutable** (see [Validation](<#validation>)). | Each field directly under `spec.template.spec` is the `TerraformMachine` field of the same name: | Field | Type | Description | | --- | --- | --- | | `spec.template.spec.source` | object | The machine-role module image. **Required.** See [`spec.source`](). | | `spec.template.spec.identityRef` | object | The identity the machines’ Jobs use. Machines fall back to their `TerraformCluster`’s defaults when unset. See [`spec.identityRef`](). | | `spec.template.spec.jobs` | object | Job tuning for each machine. See [`spec.jobs`](). Only its `imagePullSecrets` take part in resolving this template’s status. | | `spec.template.spec.variables` | object | Inline module variables for each machine. See [`spec.variables`](). | | `spec.template.spec.variablesFrom` | list | Variable sources for each machine. See [`spec.variablesFrom`](). | | `spec.template.spec.drift` | object | Drift check policy of each machine. See [Drift](). | | `spec.template.spec.remediation` | object | Remediation policy of each machine. See [Remediation](). | | `spec.template.spec.providerID` | string | Must be empty. The controller assigns a provider ID to each machine. See [Provider ID](). | The fields under `spec.template.spec` are documented once, on [TerraformMachine](); the shared ones in depth on [Common Fields](). A change to a field in a template reaches machines only through a rollout: see [Validation](<#validation>). ## Status The controller resolves the status once for each image reference, from the image’s config labels. Tags are not polled again: a new image is a new template. When `spec.template.spec.source.image` differs from `status.capacitySource.image`, the controller resolves it again. | Field | Type | Description | | --- | --- | --- | | `status.capacity` | map of resource name to quantity | The resources of the node the template creates, from the image label `io.captf.capacity`, for example `cpu: "4"`. Unset when the image declares none or the label is invalid. | | `status.nodeInfo` | object | The platform of the node, from the image label `io.captf.node-info`. Unset when the image declares none or the label is invalid. Must set at least one property. | | `status.nodeInfo.architecture` | string | The node’s CPU architecture. **Allowed values:** `amd64`, `arm64`, `s390x`, `ppc64le`. | | `status.nodeInfo.operatingSystem` | string | The node’s operating system, for example `linux`. **Range:** 1 to 64 characters. | | `status.capacitySource` | object | Records which image `capacity` and `nodeInfo` came from. | | `status.capacitySource.image` | string | The `spec.template.spec.source.image` last resolved. **Range:** 1 to 512 characters. | | `status.conditions` | list | The `CapacityResolved` condition. **Range:** at most 32. | | `status.conditions[].type` | string | The condition type: `CapacityResolved`. | | `status.conditions[].status` | string | `True`, `False` or `Unknown`. | | `status.conditions[].reason` | string | A CamelCase reason. Every reason is on [Conditions](). | | `status.conditions[].message` | string | A human-readable detail. | | `status.conditions[].lastTransitionTime` | time | When `status` last changed. | | `status.conditions[].observedGeneration` | integer | The `metadata.generation` the condition was computed for. | ### Where the values come from The manager reads the image’s config over the registry API, not by pulling it. For an image index it takes the `linux` image of the manager’s own architecture, else the first `linux` image. It authenticates with the Secrets in `spec.template.spec.jobs.imagePullSecrets`, in the template’s namespace. A template has no `TerraformCluster`, so the cluster’s `spec.defaults.jobs.imagePullSecrets` do not apply here. The two labels are defined on the [image contract](). | Label | Value | Sets | | --- | --- | --- | | `io.captf.capacity` | A JSON object of resource name to Kubernetes quantity, such as `{"cpu":"4","memory":"16Gi"}`. Each key must be a valid resource name and each value a quantity. An empty object is invalid. | `status.capacity` | | `io.captf.node-info` | A JSON object `{"architecture":"amd64","operatingSystem":"linux"}` with at least one key and no others. | `status.nodeInfo` | status of a resolved template ```yaml status: capacity: cpu: "4" memory: 16Gi nodeInfo: architecture: amd64 operatingSystem: linux capacitySource: image: ghcr.io/captf-io/aws-machine:v0.1.0-opentofu conditions: - type: CapacityResolved status: "True" reason: CapacityResolved lastTransitionTime: "2026-10-02T08:10:31Z" ``` The Cluster API provider of the Cluster Autoscaler reads `status.capacity` and `status.nodeInfo` from the infrastructure template of a MachineDeployment or MachineSet that is at zero replicas, to decide whether a pending Pod would fit a new node. A module fixes the instance type, so an image describes one node size: an image that varies its size by variable needs one image per size to use the labels. The labels are optional, and pool images ignore them. ### Conditions `CapacityResolved` is the only condition. A template has no `Ready` condition. | Status | Reasons | Meaning | | --- | --- | --- | | `True` | [`CapacityResolved`]() | Every label the image carries parsed; `capacity` and `nodeInfo` are set from them. | | `True` | [`CapacityNotDeclared`]() | The image carries neither label. Both fields stay unset. This is not an error. | | `False` | [`ImageInspectFailed`]() | The registry could not be read: authentication failed, the image was not found, or the registry was unreachable. The previous values stay. The message names the class of failure, never the registry’s own text. The manager retries, doubling the delay from 30 seconds to at most 10 minutes. | | `False` | [`CapacityLabelInvalid`]() | A label is present but invalid. That field is unset; a valid other label is still applied. The manager does not retry the same image, because the tag is not polled again. | The manager emits a `CapacityResolved` Normal event when the values change and an `ImageInspectFailed` Warning event on the first failure. See [Events](). ## Printer columns `kubectl get terraformmachinetemplates` shows: | Column | Source | Description | | --- | --- | --- | | `Image` | `.spec.template.spec.source.image` | The module image. | | `Age` | `.metadata.creationTimestamp` | Time since creation. | ## Validation The admission webhook checks templates on create and update. Deletes are always allowed. - `spec.template.spec` is immutable. Create a new template and point the owner at it; Cluster API rolls the machines. This holds even for fields that are mutable on a `TerraformMachine` (`jobs`, `drift`, `remediation`). Checked on update. - The immutability rule does not apply to a ClusterClass topology dry-run, so a topology patch can be validated. Checked on update. - `spec.template.spec.providerID` must be empty: a value would stamp every clone with one instance’s identity. Checked on create and update. - `spec.template.metadata` labels and annotations must be valid keys and values. `spec.template.metadata` itself is mutable. Checked on create and update. - `spec.template.spec` follows the `TerraformMachine` spec rules: `source.image` is required and must parse, `jobs` may not weaken the security contexts, `lockTimeoutSeconds` must be below `activeDeadlineSeconds`, and `variables` and `variablesFrom` are well-formed. Checked on create and update. - `spec.template` must set at least one property, and `spec.template.spec` is required. Checked on create and update. A template’s rules are the same as the machine’s, so a bad image reference or a root-running security context is rejected when you create the template, not when Cluster API clones it. See [TerraformMachine validation]() for each one. > [!NOTE] > > **See also** > > - [TerraformMachine]() for every field under `spec.template.spec`. > - [Common Fields]() for the shared workspace fields. > - [The Kinds]() > - [Image contract: OCI labels]() > - [Conditions]() and [Events]() # TerraformMachinePool A `TerraformMachinePool` is the infrastructure object behind one Cluster API `MachinePool`: a group of nodes that CAPTF provisions and scales as a single native cloud scaling group, by running a machinepool-role module as a Kubernetes Job. The `MachinePool` references it through `MachinePool.spec.template.spec.infrastructureRef`, and Cluster API sets the `MachinePool` as its owner. You create it yourself, or Cluster API creates it from a [`TerraformMachinePoolTemplate`]() when a ClusterClass topology manages the `MachinePool`. Unlike a `TerraformMachine`, every field of a pool is mutable: a change re-applies the module. The controller writes the group’s membership back into `spec.providerID` and `spec.providerIDList`. See [The Kinds]() for how the three workload kinds compare and [Machine Pools]() for the task guide. | Property | Value | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Kind | `TerraformMachinePool` | | Scope | Namespaced | | Module role | `machinepool` (see the [machinepool contract]()) | | Created by | You, or Cluster API from a [`TerraformMachinePoolTemplate`]() through a ClusterClass | | Referenced by | `MachinePool.spec.template.spec.infrastructureRef` | | Finalizer | `terraformmachinepool.infrastructure.cluster.x-k8s.io` | | Short names | None | | Categories | `cluster-api` | | Status subresource | Yes | ## Example The smallest useful pool names a module image and carries the cluster’s name label. `spec.source.image` is the only required field, and `spec.identityRef` can be omitted when the `TerraformCluster` provides a default. pool.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: my-cluster-workers namespace: team-a labels: cluster.x-k8s.io/cluster-name: my-cluster # (1)! spec: source: image: ghcr.io/captf-io/aws-machinepool:v0.1.0-opentofu identityRef: name: aws-prod ``` 1. CAPTF finds the owning `Cluster` by this label, not by the `MachinePool`’s `spec.clusterName`. Cluster API adds it on its own next reconcile if it is missing, so setting it yourself avoids a wait. The `MachinePool` that owns the pool points back at it, and also needs a bootstrap provider object (see [Create a MachinePool]()). ## Full example pool-full.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePool metadata: name: my-cluster-workers namespace: team-a labels: cluster.x-k8s.io/cluster-name: my-cluster spec: source: image: ghcr.io/captf-io/aws-machinepool:v0.1.0-opentofu imagePullPolicy: IfNotPresent identityRef: name: aws-prod # (1)! jobs: activeDeadlineSeconds: 3600 # (2)! lockTimeoutSeconds: 300 variables: # (3)! instance_type: m6i.large root_volume_gib: 100 variablesFrom: - configMapRef: name: pool-common-vars drift: intervalSeconds: 900 # (4)! action: Remediate membershipRefreshIntervalSeconds: 30 # (5)! providerID: aws:///us-east-1/my-cluster-workers # (6)! providerIDList: # (7)! - aws:///us-east-1a/i-0a1b2c3d4e5f60001 - aws:///us-east-1b/i-0a1b2c3d4e5f60002 ``` 1. Falls back to the `TerraformCluster`’s `spec.defaults.identityRef`, then its `spec.identityRef`, when unset. 2. Merged field by field over the cluster’s `spec.defaults.jobs`. 3. The inline value wins over `variablesFrom` on the same key. The names `captf_*` and the role’s contract inputs are reserved. 4. Both sub-fields fall back to the cluster’s `spec.defaults.drift`. A pool’s drift cannot be turned off. 5. How often CAPTF refreshes the group’s membership between applies. 6. Set by the controller from the module’s `provider_id` output. You do not write it. 7. Set by the controller from the module’s `provider_id_list` output. You do not write it. ## Spec The `spec` block must not be empty, and `spec.source` is required. No spec field is immutable. Because the CRD declares no defaults, every default below is applied by the controller at reconcile, never stored in the object. The fields embedded from the shared workspace spec are documented in depth on [Common Fields](). The table links to each. | Field | Type | Description | | --- | --- | --- | | `spec.source` | object | The machinepool module image and pull policy. See [Source](). **Required.** **Mutable.** Never inherited from the cluster. | | `spec.identityRef` | object | The `TerraformClusterIdentity` whose credentials the pool’s Jobs use. See [Identity reference](). **Mutable.** **Default:** the owning `TerraformCluster`’s `spec.defaults.identityRef`, else its `spec.identityRef`. | | `spec.jobs` | object | Tuning of the Jobs that run the module. See [Jobs](). **Mutable.** **Default:** merged field by field over the cluster’s `spec.defaults.jobs`, then the built-in per-field defaults. | | `spec.variables` | object | Inline module variables, a JSON object. See [Variables](). **Mutable.** A change re-applies the pool. Wins over `spec.variablesFrom`. | | `spec.variablesFrom` | list | ConfigMaps and Secrets that supply module variables. See [Variable sources](). **Mutable.** | | `spec.providerID` | string | The scaling group’s provider ID. See [Provider IDs](<#provider-ids>). | | `spec.providerIDList` | list of strings | The provider IDs of every non-terminated member. See [Provider IDs](<#provider-ids>). | | `spec.drift` | object | How often to check for drift and what to do about it. See [Drift](<#drift>). **Mutable.** | | `spec.membershipRefreshIntervalSeconds` | integer | How often to refresh group membership between applies. See [Membership refresh](<#membership-refresh>). | ### Provider IDs The controller writes both fields from the module’s outputs. You do not set them, and a template accepts but does not use them. | Field | Type | Description | | --- | --- | --- | | `spec.providerID` | string | The scaling group’s provider ID, from the module’s `provider_id` output. Optional in the Cluster API `InfraMachinePool` contract and may stay unset for modules that have no group object. **Optional.** **Mutable.** **Range:** 1 to 512 characters. | | `spec.providerIDList` | list of strings | The provider IDs of every non-terminated member, from the module’s `provider_id_list` output, sorted and deduplicated. Each entry must equal the corresponding Node’s `spec.providerID`. Cluster API copies the list to `MachinePool.spec.providerIDList`, and a Node stays unschedulable until its ID is in it. **Optional.** **Mutable.** **Range:** at most 10000 entries, each 1 to 512 characters. | The list changes whenever the group’s membership does, which is why neither field is immutable. ### Drift `spec.drift` is merged field by field over the cluster’s `spec.defaults.drift`. A field you leave unset inherits the cluster’s value. See [Drift]() for the task guide and [Drift and Health]() for how a check runs. | Field | Type | Description | | --- | --- | --- | | `spec.drift.intervalSeconds` | integer | Seconds between drift checks. **Mutable.** **Default:** the cluster’s `spec.defaults.drift.intervalSeconds` when positive, else the manager’s `--drift-default-interval` (30 minutes; see [Manager Flags]()). **Range:** 1 or more. | | `spec.drift.action` | string | What CAPTF does when a check finds changes. **Mutable.** **Default:** `Report`. **Allowed values:** `Report` records the difference in the `DriftDetected` condition only; `Remediate` applies the pool’s current inputs to remove it. | A pool’s drift cannot be disabled, so `0` is rejected for `spec.drift.intervalSeconds`, and an inherited `0` from the cluster’s defaults is ignored in favor of the manager default. The drift Job’s refresh feeds the group’s current size into its plan; with no checks, a cloud-side scaling change would never register. > [!WARNING] > > **Remediation can fight a module’s own autoscaling** > > With `spec.drift.action: Remediate` and autoscaling enabled, the module must exclude its desired-count attribute from the plan. Otherwise every cloud-side scale reports as drift and is reverted. ### Membership refresh | Field | Type | Description | | --- | --- | --- | | `spec.membershipRefreshIntervalSeconds` | integer | Seconds between `apply -refresh-only` runs that pick up members joining or leaving the group between applies. **Optional.** **Mutable.** **Default:** 60 when unset or 0. **Range:** 15 to 86400. | The schema minimum of 15 means `0` is never a valid explicit value, so it always reads as unset. The controller also refreshes right after every apply. While the group has not converged (the length of `spec.providerIDList` differs from `status.replicas`), it refreshes every 30 seconds, or at the configured interval when that is shorter. A shorter interval gets new nodes ready sooner at the cost of more Jobs. See [Set the membership refresh interval](). ### Relation to the MachinePool The owning `MachinePool` supplies the pool’s size, not the `TerraformMachinePool`: the desired replica count reaches the module as its `replicas` input, from `MachinePool.spec.replicas`. - **Fixed size.** With no autoscaler annotations, `MachinePool.spec.replicas` is the only source of desired capacity. Change it to resize the group. - **Autoscaling.** With both `cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size` and `...-max-size` annotations on the `MachinePool`, the module owns the desired count. The controller then writes the observed capacity back to `MachinePool.spec.replicas` on every reconcile. The `AutoscalingActive` condition reports the mode; see [Conditions](<#conditions>). - **Membership.** Cluster API reads `spec.providerIDList` to match Nodes to the group, and reads `status.ready` to decide the pool is provisioned. ## Status The `status` block is rebuilt from the state Secret, the durable inputs and the Job list, so nothing in it is lost by `clusterctl move`. Fields shared with the other workload kinds are documented on [Common Fields](). | Field | Type | Description | | --- | --- | --- | | `status.conditions` | list | The pool’s conditions, keyed by `type`, at most 32. See [Conditions](<#conditions>). | | `status.conditions[].type` | string | The condition type, such as `Ready` or `ApplyJobSucceeded`. | | `status.conditions[].status` | string | `True`, `False` or `Unknown`. | | `status.conditions[].observedGeneration` | integer | The `metadata.generation` the condition was computed for. | | `status.conditions[].lastTransitionTime` | time | When the condition last changed status. | | `status.conditions[].reason` | string | A machine-readable reason. See [Conditions](). | | `status.conditions[].message` | string | A human-readable detail. It never contains a variable value. | | `status.ready` | boolean | The v1beta1 compatibility flag Cluster API reads to decide the pool is provisioned. Latched together with `status.initialization.provisioned`: once `true`, it stays `true`. | | `status.replicas` | integer | The group’s desired capacity at the last refresh, from the module’s `replicas` output. Outside a scaling transition it equals the length of `spec.providerIDList`. **Range:** 0 or more. | | `status.instances` | list | The group’s members, from the module’s `instances` output, at most 1000. See [Instances](<#instances>). | | `status.initialization` | object | The Cluster API contract’s provisioned flag. See [Initialization](). | | `status.observedGeneration` | integer | The generation this status was computed for. See [Observed generation](). | | `status.activeJob` | object | The Job running for the pool now, if any. See [Active job](). | | `status.lastRun` | object | The result of the last completed Job. See [Last run](). | | `status.lastDriftCheck` | time | When the last drift check completed. See [Drift checks and refreshes](). | | `status.lastRefresh` | time | When the last refresh or drift check completed. See [Drift checks and refreshes](). | | `status.pendingRefreshes` | integer | Consecutive health samples that read `pending`, which spaces refreshes while the group starts up. See [Drift checks and refreshes](). | | `status.observedStateSerial` | integer | The state serial the outputs were read from. See [State](). | | `status.lastRestoredSerial` | integer | The backup serial of the last consumed restore Job. See [State](). | | `status.stateSecretSuffix` | string | The state Secret’s backend suffix, derived by the controller. See [State](). | | `status.stateBackups` | list | The state backups the controller keeps, newest first. See [State](). | | `status.source` | object | What the last Job actually ran. See [Image in use](). | The whole status block is reported under [Workspace status](). ### Instances `status.instances` mirrors the module’s `instances` output. Its shape is module-defined and core Cluster API does not use it. If the module reports more than the controller keeps, the list is shortened and `OutputsValid` carries the `InstancesTruncated` reason (see [Conditions]()). | Field | Type | Description | | --- | --- | --- | | `status.instances[].providerID` | string | The instance’s provider ID. **Required.** **Range:** 1 to 512 characters. | | `status.instances[].instanceID` | string | A provider-defined identifier, distinct from `providerID` when the module has one to give. **Range:** 1 to 256 characters. | | `status.instances[].addresses` | list | The instance’s addresses, 1 to 256 entries. | | `status.instances[].addresses[].type` | string | The Cluster API machine address type. **Allowed values:** `Hostname`, `ExternalIP`, `InternalIP`, `ExternalDNS`, `InternalDNS`. | | `status.instances[].addresses[].address` | string | The address itself. | | `status.instances[].failureDomain` | string | The failure domain the instance actually runs in. **Range:** 1 to 256 characters. | | `status.instances[].state` | string | The instance’s health, the contract’s `health.state`. **Allowed values:** `pending`, `running`, `degraded`, `stopped`, `terminated`, `unknown`. | > [!NOTE] > > **Example status** > > ```yaml > status: > ready: true > replicas: 2 > initialization: > provisioned: true > observedGeneration: 4 > observedStateSerial: 17 > stateSecretSuffix: <16 hex digits>-mp # of sha256(namespace/kind/name) > lastRefresh: "2026-10-02T09:41:12Z" > lastDriftCheck: "2026-10-02T09:30:07Z" > instances: > - providerID: aws:///us-east-1a/i-0a1b2c3d4e5f60001 > instanceID: i-0a1b2c3d4e5f60001 > failureDomain: us-east-1a > state: running > addresses: > - type: InternalIP > address: 10.0.1.23 > - providerID: aws:///us-east-1b/i-0a1b2c3d4e5f60002 > instanceID: i-0a1b2c3d4e5f60002 > failureDomain: us-east-1b > state: pending > conditions: > - type: Ready > status: "True" > reason: Ready > observedGeneration: 4 > lastTransitionTime: "2026-10-02T09:12:40Z" > - type: AutoscalingActive > status: "False" > reason: AutoscalingDisabled > observedGeneration: 4 > lastTransitionTime: "2026-10-02T09:12:40Z" > ``` ## Conditions `Ready` is the summary: Cluster API mirrors it into the `MachinePool`’s `InfrastructureReady` condition. Every reason of every condition is listed in [Conditions](); this page names only which conditions a pool carries. | Condition | `True` | `False` | `Unknown` | | --- | --- | --- | --- | | `Ready` | Every input is healthy. | An input is `False`. | An input is `Unknown`. | | `DependenciesReady` | The owner, the cluster, its exports, the bootstrap data and any variable sources are in place. | The owner is gone, mismatched or has no `MachinePool` ownerRef yet, the `Cluster` is not a `TerraformCluster`, or a variable source is missing or invalid. | The pool waits for its owner, the cluster, its exports or bootstrap data. | | `IdentityAllowed` | The identity exists and permits this namespace. | No identity is set, it is missing, does not permit this namespace or has no Secret. | The identity check could not be completed. | | `CredentialsMirrored` | The Job’s credentials are in place. | They could not be copied. | The mirror is pending. | | `RunnerRBACReady` | The Job’s ServiceAccount and RBAC exist. | They could not be created, or an override ServiceAccount lacks the `captf.io/runner=true` label. | Never `Unknown`. | | `ApplyJobSucceeded` | The last apply or destroy succeeded. | It failed, hit its deadline or could not start, or a destructive change of the cluster’s exports waits for approval. | No apply has completed yet, or the run waits for a lease or for the cluster. A running apply keeps the last result. | | `StateReadable` | The state Secret reads cleanly. | It is encrypted, corrupt, inconsistent, lost or locked. | No state exists yet. | | `RestoreJobSucceeded` | The last state restore succeeded. | It failed, or the requested backup does not exist. | The restore waits for a run lease or for the cluster. Set only once a restore is requested. | | `OutputsValid` | The module’s outputs satisfy the contract. | An output is missing or invalid, or a provider ID changed. | Required outputs are still `null`. | | `InfrastructureHealthy` | The module reports the group healthy. | The pool is provisioning, or the module reports pending, unhealthy, degraded, stopped or terminated. | Before the first apply, or health is unknown. | | `DriftJobSucceeded` | The last drift check succeeded. | It failed or hit its deadline. | No check has completed, a check is running or waits for a lease, or the durable inputs Secret is missing. | | `DriftDetected` | The last check found drift: reported, pending remediation or being remediated. | It found none. | No check has completed. | | `AutoscalingActive` | The module owns the group’s desired count and the controller writes it back to `MachinePool.spec.replicas`. | Autoscaling is off, its annotations are invalid, or another controller owns `spec.replicas`. | Never `Unknown`. | | `Paused` | The `Cluster` or object is paused. | It is not. | Never `Unknown`. | | `Deleting` | The pool is being deleted. It makes `Ready` `False`. | It is not. | Never `Unknown`. | After provisioning, `Ready` is computed from `InfrastructureHealthy`, `ApplyJobSucceeded` and `Deleting`. Unlike for the cluster and the machine, a failed re-apply shows in a pool’s `Ready`, because a pool is re-applied regularly. `DriftDetected`, `DriftJobSucceeded`, `RestoreJobSucceeded`, `AutoscalingActive` and `Paused` never feed `Ready`. A pool becomes provisioned once its state carries a successful apply and the module reports a health state other than `pending`. That does not require `spec.providerIDList` to be non-empty, since a pool can scale to zero. ## Printer columns `kubectl get terraformmachinepools` shows: | Column | Source | | --- | --- | | `CLUSTER` | The `cluster.x-k8s.io/cluster-name` label. | | `MACHINEPOOL` | The name of the owner reference of kind `MachinePool`. | | `REPLICAS` | `status.replicas`. | | `READY` | The `Ready` condition’s status. | | `AGE` | Time since creation. | ## Validation The CRD schema and the validating webhook (`validation.terraformmachinepool.infrastructure.cluster.x-k8s.io`, on create and update, `failurePolicy: Fail`) enforce these rules. The webhook returns one `Invalid` error that lists every violation. It sets no defaults: all defaults come from the controller at reconcile. - **Required.** `spec` must have at least one field, and `spec.source.image` must be set. - **Image reference.** `spec.source.image` must parse as a container image reference. The webhook checks the syntax only, not the registry or the contents. `spec.source.imagePullPolicy` must be `IfNotPresent`, `Always` or `Never`. - **Immutability.** None. Every field of the pool can change on a live object, and a change re-applies the module. - **Ranges.** - `spec.providerID`: 1 to 512 characters. - `spec.providerIDList`: at most 10000 entries, each 1 to 512 characters. - `spec.drift.intervalSeconds`: 1 or more. - `spec.membershipRefreshIntervalSeconds`: 15 to 86400. - `status.instances`: at most 1000 entries. `status.conditions`: at most 32. - **Enums.** `spec.drift.action` is `Report` or `Remediate`. `status.instances[].state` is one of the six health states. - **Variables.** - `spec.variables` is a JSON object with 1 to 256 keys. - Each key is a Terraform identifier that does not start with `captf_`, is not a machinepool contract input and is not a module meta-argument. - Each `spec.variablesFrom` entry sets exactly one of `configMapRef` and `secretRef`. Keys inside a referenced ConfigMap or Secret are checked at reconcile and report `VariablesInvalid`. - Errors never include a value. - **Job policy.** `spec.jobs` may not weaken the hardened security context: no privileged container, privilege escalation, added capabilities, writable root filesystem, unmasked `/proc`, `runAsNonRoot: false`, UID 0, `Unconfined` seccomp profile or Windows host process, in the container or pod context. `lockTimeoutSeconds` must be less than `activeDeadlineSeconds`; an unset one is compared with the other’s built-in default. The check runs on create, and on update only when `spec.jobs` changed, never on a deleting object, so a stored policy cannot block a finalizer removal. The merge with the cluster’s defaults is not checked. See [Tuning Jobs](). - **Delete.** Always allowed. ## Lifecycle - **Create.** The controller waits for the owning `MachinePool`, the cluster to be provisioned and its exports and the bootstrap data to be ready, then applies the module. It refreshes right after the apply to read the group’s members. - **Scale.** A change to `MachinePool.spec.replicas` re-applies the pool. In autoscaling mode the controller instead writes the observed capacity back. - **Change.** Any spec change, and a rotation of the bootstrap data Secret, re-applies the module. An apply that renders a change of the cluster’s exports is guarded: if its plan deletes or replaces anything, it waits for the `captf.io/approve-destructive-plan` annotation (see [The destructive-plan guard]()). - **Drift and refresh.** Drift checks run on `spec.drift.intervalSeconds` and membership refreshes on `spec.membershipRefreshIntervalSeconds`. - **Delete.** Nothing blocks the deletion. The controller runs a destroy Job from the stored inputs and removes the finalizer once it succeeds, even if the `MachinePool` or `Cluster` no longer exists. See [The reconcile lifecycle](). > [!NOTE] > > **See also** > > - [TerraformMachinePoolTemplate]() and [Common Fields](). > - [Machine Pools](), [Drift](), [Identities and Credentials]() and [Tuning Jobs](). > - [The machinepool contract](). > - [The Kinds](), [The Reconcile Lifecycle]() and [Drift and Health](). > - [Conditions](), [Annotations, Labels and Finalizers]() and [Manager Flags](). # TerraformMachinePoolTemplate A `TerraformMachinePoolTemplate` is the blueprint a ClusterClass uses for the infrastructure of a `MachinePool`. A ClusterClass `machinePools` class references it in `infrastructure.templateRef`, and when a topology creates a `MachinePool` from that class, Cluster API clones `spec.template` into a [`TerraformMachinePool`]() that the `MachinePool` references. You do not need a template for a pool you create by hand; see [Templates and ClusterClass](). The template has no status. A pool has no scale-from-zero, so there is no capacity or node information to resolve. | Property | Value | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Kind | `TerraformMachinePoolTemplate` | | Scope | Namespaced | | Module role | `machinepool`, through the `TerraformMachinePool` it creates | | Created by | You | | Referenced by | A ClusterClass `machinePools[].infrastructure.templateRef` | | Finalizer | None | | Short names | None | | Categories | `cluster-api` | | Status subresource | No | ## Example machinepool-template.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformMachinePoolTemplate metadata: name: aws-workers-v1 namespace: team-a spec: template: metadata: labels: node-role.kubernetes.io/worker: "" # (1)! spec: source: image: ghcr.io/captf-io/aws-machinepool:v0.1.0-opentofu identityRef: name: aws-prod variables: instance_type: m6i.large drift: action: Remediate ``` 1. Metadata on `spec.template.metadata` is copied to the created `TerraformMachinePool`. The module also receives the labels as its `node_labels` input. ## Spec The `spec` block holds one field, `spec.template`, and `spec.template.spec` is required inside it. | Field | Type | Description | | --- | --- | --- | | `spec.template` | object | The `TerraformMachinePool` to create. **Required.** | | `spec.template.metadata` | object | Metadata copied onto the created object. **Optional.** | | `spec.template.metadata.labels` | map of strings | Labels for the created object. **Optional.** **Allowed values:** valid Kubernetes label keys and values. | | `spec.template.metadata.annotations` | map of strings | Annotations for the created object. **Optional.** **Allowed values:** valid Kubernetes annotation keys. | | `spec.template.spec` | object | The `TerraformMachinePool` spec to create. **Required.** **Immutable.** Takes the fields below. | | `spec.template.spec.source` | object | The module image. See [`spec.source`](). **Required.** | | `spec.template.spec.identityRef` | object | The identity for the pool’s Jobs. See [`spec.identityRef`](). | | `spec.template.spec.jobs` | object | Job tuning. See [`spec.jobs`](). | | `spec.template.spec.variables` | object | Inline module variables. See [`spec.variables`](). | | `spec.template.spec.variablesFrom` | list | Variable sources. See [`spec.variablesFrom`](). | | `spec.template.spec.drift` | object | Drift interval and action. See [Drift](). | | `spec.template.spec.membershipRefreshIntervalSeconds` | integer | Membership refresh interval. See [Membership refresh](). | | `spec.template.spec.providerID` | string | Accepted but not meaningful: the controller writes it on the created pool. See [Provider IDs](). | | `spec.template.spec.providerIDList` | list of strings | Accepted but not meaningful: the controller writes it on the created pool. See [Provider IDs](). | The fields take the same types, ranges and defaults as on the pool, and the same CRD schema limits apply. Defaults and inheritance from the cluster’s `spec.defaults` apply to the created pool, not to the template. ## Printer columns `kubectl get terraformmachinepooltemplates` shows: | Column | Source | | --- | --- | | `IMAGE` | `spec.template.spec.source.image`. | | `AGE` | Time since creation. | ## Validation The validating webhook (`validation.terraformmachinepooltemplate.infrastructure.cluster.x-k8s.io`, on create and update) applies these rules. - **Template spec rules.** `spec.template.spec` is checked like a pool’s spec: a valid `spec.template.spec.source.image` reference, the Job policy rules and the variable rules for the `machinepool` role. See [Validation]() on the pool. The Job policy is checked on update only when it changed. - **Template metadata.** `spec.template.metadata.labels` and `spec.template.metadata.annotations` must be valid label and annotation keys and values. - **Immutability.** After creation, `spec.template.spec` cannot change: the webhook rejects it with “TerraformMachinePoolTemplate spec.template.spec is immutable; create a new TerraformMachinePoolTemplate instead”. The template’s metadata can change. To roll out a change, create a new template and point the ClusterClass at it, which rolls a new pool through the `MachinePool`’s own mechanism. The check is skipped for the server-side apply dry-runs a ClusterClass topology performs. The pool a template creates is itself mutable. See [Template immutability and rolling out a change](). - **Delete.** Always allowed. > [!NOTE] > > **See also** > > - [TerraformMachinePool]() and [Common Fields](). > - [Templates and ClusterClass]() and [Machine Pools](). > - [The Kinds](). # TerraformClusterIdentity A `TerraformClusterIdentity` names a Secret of cloud credentials and the namespaces allowed to use it. It is cluster-scoped. A platform admin creates it, along with the Secret. `TerraformCluster.spec.identityRef`, `TerraformCluster.spec.defaults.identityRef` and the `spec.identityRef` of a `TerraformMachine` or `TerraformMachinePool` reference it by name. The manager copies the Secret into each allowed namespace that uses the identity and mounts that copy into the Jobs it runs there. | Property | Value | | --- | --- | | API version | `infrastructure.cluster.x-k8s.io/v1alpha1` | | Scope | Cluster | | Created by | A platform admin, with `kubectl apply` | | Referenced by | `TerraformCluster` (`spec.identityRef`, `spec.defaults.identityRef`), `TerraformMachine` and `TerraformMachinePool` (`spec.identityRef`) | | Finalizers | None. The delete webhook protects an identity in use instead | | Short names | None | | Categories | `cluster-api` | | Status subresource | Yes | ## Example A minimal identity for one namespace, and the Secret it points at: identity.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: aws spec: secretRef: name: aws namespace: captf-system allowedNamespaces: list: - team-a --- apiVersion: v1 kind: Secret metadata: name: aws namespace: captf-system type: Opaque stringData: AWS_ACCESS_KEY_ID: REPLACE_WITH_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY: REPLACE_WITH_SECRET_ACCESS_KEY ``` ## Full example Every spec field set. `list` and `selector` are combined: a namespace is allowed when either matches. identity.yaml ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformClusterIdentity metadata: name: aws spec: secretRef: name: aws # (1)! namespace: captf-system # (2)! allowedNamespaces: list: # (3)! - team-a - team-b selector: # (4)! matchLabels: captf.io/identity-aws: "true" ``` 1. The Secret’s name. A DNS subdomain, up to 253 characters. 2. The Secret’s namespace. Required, and any namespace works. It does not have to be an allowed namespace. 3. Up to 100 namespace names. At least one if `list` is set. 4. A Kubernetes label selector over namespace labels. Namespaces it matches are allowed in addition to those in `list`. ## Spec `spec` must have at least one property, and `spec.secretRef` is required. Every field is mutable: the CRD and webhook mark none immutable. | Field | Type | Description | | --- | --- | --- | | `spec` | object | The desired state: the credentials Secret and who may use it. **Required.** | | `spec.secretRef` | object | The credentials Secret. **Required.** | | `spec.allowedNamespaces` | object | Which namespaces may reference the identity. **Default:** unset, which allows no namespace. **Mutable.** | ### Secret reference | Field | Type | Description | | --- | --- | --- | | `spec.secretRef.name` | string | Name of the Secret. **Required.** **Mutable.** **Range:** 1 to 253 characters. **Allowed values:** a DNS subdomain (lowercase alphanumerics, `-` and `.`, starting and ending with an alphanumeric). | | `spec.secretRef.namespace` | string | Namespace of the Secret. **Required.** **Mutable.** **Default:** none; unlike some Kubernetes references it is never defaulted to the manager’s namespace. **Range:** 1 to 63 characters. **Allowed values:** a DNS label. | The Secret may live in any namespace. The admission webhook requires that the user who creates or changes the identity may `get` it, which keeps an identity from exposing a Secret its author cannot read. See [Validation](<#validation>). ### Allowed namespaces `spec.allowedNamespaces` decides which namespaces may use the identity. The manager checks it each time an object resolves its credentials, so changing it takes effect on the next reconcile of every object that uses the identity. | Field | Type | Description | | --- | --- | --- | | `spec.allowedNamespaces.list` | array of string | Names of allowed namespaces. **Range:** 1 to 100 items, each 1 to 63 characters. **Allowed values:** DNS labels. Items are unique (a set). | | `spec.allowedNamespaces.selector` | LabelSelector | Allows every namespace whose labels match. An empty selector (`{}`) matches every namespace, including ones that do not exist yet. | How the values combine: | `allowedNamespaces` | Allowed namespaces | | --- | --- | | Unset | None | | `list` only | Exactly the listed namespaces | | `selector` only | Namespaces whose labels match | | `list` and `selector` | The union: a namespace in `list`, or one the selector matches | | `selector: {}` | Every namespace | | `{}` (neither field) | Rejected at admission | The empty object is rejected rather than read as “every namespace” or “no namespace”, because the most permissive value should not look like the least. Write `selector: {}` for every namespace. A `list` of zero items also fails validation (`MinItems=1`). `spec.allowedNamespaces.selector` is a standard Kubernetes [`LabelSelector`]() applied to the labels of the `Namespace` object, so `matchLabels` and `matchExpressions` both work. CAPTF reads the namespace of the referencing object at reconcile time. It does not list namespaces ahead of time. A selector that is invalid, or a namespace that does not exist, denies access; a failed read of the namespace leaves the check `Unknown` (`IdentityCheckFailed`) and retries. identity.yaml ```yaml spec: allowedNamespaces: selector: matchExpressions: - key: environment operator: In values: ["staging", "production"] ``` > [!WARNING] > > **Anyone who can label a namespace can opt it in** > > A selector trusts namespace labels. Anyone who may edit the labels on a namespace can bring it under the selector, so restrict who may label namespaces, or prefer `list`. ## The credentials Secret The Secret is a plain `Opaque` Secret. Which keys it holds is defined by the module that uses the identity: name the keys for the variables its providers read, not for CAPTF. For the keys each cloud expects, see the [AWS](), [Azure](), [GCP](), [OCI]() and [OpenStack]() module pages. For creating it, see [Identities and Credentials](). A Job never mounts the source Secret. The manager copies it into the namespace of each object that uses the identity, as `captf-creds-` (labeled `captf.io/mirrored: "true"`, annotated with `captf.io/source-hash`; see [Annotations and labels]()). The Job gets the copy twice: - As environment variables, one per key, through `envFrom`. A key starting with `TF_` or `KUBE_` is not passed to the environment, except the variables the Job sets itself. - As read-only files, one per key, under `/var/run/captf/credentials/` with file mode `0440`. Every key appears as a file, including `TF_` and `KUBE_` keys. The identity never owns the source Secret. The mirror is owned by the objects that use it, and the manager deletes it when none remain. > [!CAUTION] > > **Every Job in an allowed namespace carries these credentials** > > Anything that runs in the Job, including the module’s Terraform code and its providers, can read them. Grant an identity only the cloud permissions the modules need, and allow only the namespaces that should have them. See the [security model](). ## Status The manager sets `status`; you do not write it. `status` has at least one property when present. | Field | Type | Description | | --- | --- | --- | | `status` | object | The observed state: whether the credentials Secret exists, and where it is mirrored. | | `status.namespaces` | array of string | Namespaces where a mirror of the credentials Secret currently exists and an object uses the identity. Sorted. **Range:** up to 1000 items, each 1 to 63 characters. | | `status.conditions` | array of Condition | The identity’s conditions. Only `Ready` is set. **Range:** up to 32 items. Keyed by `type`. | | `status.conditions[].type` | string | The condition type: `Ready`. | | `status.conditions[].status` | string | `True`, `False` or `Unknown`. | | `status.conditions[].reason` | string | A machine-readable reason: `SecretFound` or `SecretNotFound`. See [Conditions](). | | `status.conditions[].message` | string | A human-readable detail. For `SecretNotFound` it names the Secret as `namespace/name`. | | `status.conditions[].lastTransitionTime` | time | When `status` last changed. | | `status.conditions[].observedGeneration` | integer | The `metadata.generation` the condition was computed from. | A namespace appears in `status.namespaces` only when it holds the mirror Secret, correctly labeled and annotated, and at least one object there uses the identity. A Secret with the mirror’s name that someone else created does not count. > [!NOTE] > > **Example status** > > ```yaml > status: > conditions: > - type: Ready > status: "True" > reason: SecretFound > observedGeneration: 2 > lastTransitionTime: "2026-10-01T14:03:11Z" > namespaces: > - team-a > - team-b > ``` ## Conditions An identity sets one condition, `Ready`. | Status | Meaning | | --- | --- | | `True` | The credentials Secret named by `spec.secretRef` exists (`SecretFound`). | | `False` | The Secret does not exist (`SecretNotFound`). Objects that use this identity start no Job until it does. | | `Unknown` | Not set by the identity controller. | `Ready` says nothing about which namespaces are allowed or whether a mirror exists; the objects that reference the identity report that through their `IdentityAllowed` and `CredentialsMirrored` conditions. The controller does not watch the source Secret, so it re-reads it every five minutes and notices a Secret created or deleted out of band within that time. It also emits `IdentitySecretNotFound` and `IdentitySecretFound` events when `Ready` changes. Every reason and its meaning is in [Conditions](). ## Printer columns `kubectl get terraformclusteridentities` shows: | Column | Source | | --- | --- | | `Secret` | `.spec.secretRef.name` | | `Namespace` | `.spec.secretRef.namespace` | | `Ready` | The `status` of the `Ready` condition | | `Age` | `.metadata.creationTimestamp` | ## Validation The CRD schema and the validating admission webhook enforce these rules. The webhook runs on create, update and delete, and a request fails if the webhook is unreachable. Schema rules: - `spec` must have at least one property, and `spec.secretRef.name` and `spec.secretRef.namespace` are required, with the length and character limits in [Secret reference](<#secret-reference>). - `spec.allowedNamespaces` must set `list`, `selector` or both. The empty object is rejected with a message telling you to write `selector: {}`. - `spec.allowedNamespaces.list` has 1 to 100 unique items, each a DNS label. - `status.conditions` has at most 32 items and `status.namespaces` at most 1000. Webhook rules, on create and update: - `spec.secretRef.name` and `spec.secretRef.namespace` must be non-empty. - Each `spec.allowedNamespaces.list` item must be a valid DNS label. - `spec.allowedNamespaces.selector` must parse as a label selector. - The requesting user must be allowed to `get` the Secret. The webhook sends a `SubjectAccessReview` for verb `get` on `secrets` in `spec.secretRef.namespace`, named `spec.secretRef.name`, as the requester (user, groups and extra attributes). If it is denied, the request fails with `user "" may not get Secret /`. On create the check always runs. On update it runs only when `spec.secretRef` or `spec.allowedNamespaces` changed, in either direction, because the first chooses a Secret and the second chooses where it is copied. Changing only labels or annotations skips it. Webhook rules, on delete: - The delete is refused while any `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`, in any namespace, resolves to the identity through its own `identityRef` or a fallback. An object being deleted still counts. The message names one such object. - The delete is refused while `status.namespaces` is not empty. The message lists the namespaces. The `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` webhooks do not look up the identity they reference. A cluster must set a non-empty `spec.identityRef`, but the name need not exist when you create it. The controllers check it at reconcile time and report `IdentityNotFound`, `NamespaceNotAllowed` or `IdentityNotAllowed` through conditions, and start no Job. See [`identityRef`]() and the pages for [`TerraformCluster`](), [`TerraformMachine`]() and [`TerraformMachinePool`](). ## Lifecycle - **Create.** Create the Secret, then the identity. The creator must be able to `get` the Secret. `Ready` becomes `True` once the Secret exists. See [Create the identity](). - **Reference.** An object in an allowed namespace sets `identityRef`; the manager creates the mirror there. See [Reference it](). - **Rotate.** Edit the Secret’s data in place. The mirror in each allowed namespace is rewritten the next time an object that uses the identity reconciles, and within one `--sync-period` at the latest. A Job created after that gets the new values. A running Job keeps its old environment for its whole run, and its mounted files follow the mirror. To switch to another Secret, edit `spec.secretRef`, which reconciles every user at once and runs the `get` check again. See [Rotate credentials](). - **Revoke.** Remove a namespace from `spec.allowedNamespaces`. The manager starts no new Job there and deletes the mirror in that namespace. An object being deleted in a revoked namespace waits to destroy until the namespace is allowed again. See [Revoke access](). - **Delete.** Refused while the identity is in use or still mirrored, as described in [Validation](<#validation>). Deleting it never deletes the source Secret. See [Delete an identity](). > [!NOTE] > > **See also** > > - [Identities and Credentials](): create, reference, rotate and revoke an identity. > - [Credentials](): the mirror’s lifecycle. > - [Security model](): the trust boundary around mounted credentials. > - [Conditions](): every reason this kind and its users set. > - [Annotations and labels](): the mirror’s labels and annotations. > - [Common fields](): `identityRef` on the other kinds. > - [`TerraformCluster`]() # Common Fields `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool` embed one shared set of fields. In Go they are `WorkspaceSpec` and `WorkspaceStatus`; in YAML they sit directly under `spec` and `status` of each kind. This page documents them once, in full. Each kind’s own page ([TerraformCluster](), [TerraformMachine](), [TerraformMachinePool]()) lists them, notes what differs for that kind, and links here for the details. `TerraformCluster` also takes the same Job policy as `spec.defaults.jobs`, which its machines and pools inherit; see [Jobs](<#jobs>). Kind-specific behavior shows up in three places on this page: whether a field can change after creation, whether `spec.identityRef` is required, and how `spec.jobs` merges with the cluster’s defaults. Each is called out per field. Defaults on this page come from the admission webhook, the controllers or the manager’s flags. The CRDs declare no schema defaults, so an unset field stays unset in the stored object and the default is applied when the Job is built or at reconcile. ## Source `spec.source` names the module image. It is **required** on every kind. The image bundles the role module’s Terraform or OpenTofu code and the runtime binary at fixed paths; there is no separate module source and no separate runtime image. See the [image contract]() for the layout. | Field | Type | Description | | --- | --- | --- | | `spec.source` | `object` | The role image. **Required.** | | `spec.source.image` | `string` | The OCI image reference, `registry/repo:tag` or `registry/repo@sha256:`. The tag or digest is the module version. **Required.** **Range:** 1 to 512 characters. The webhook parses it as a normalized image reference and rejects a syntax error; it does not check the registry or the image content. | | `spec.source.imagePullPolicy` | `string` | Pull policy of the Job’s main container. **Default:** `IfNotPresent`, applied when the Job is built. **Allowed values:** `IfNotPresent`, `Always`, `Never`. Use `Always` with a mutable tag. | **Mutability.** `spec.source` is mutable on a `TerraformCluster` and a `TerraformMachinePool`: a new image is a new module version, applied against the existing state. It is **immutable** on a `TerraformMachine`; the webhook rejects any change, so you roll the machine by creating a new one (usually through a new template). Pin the module by digest ```yaml spec: source: image: ghcr.io/example/cluster-module@sha256:3f1c0e9d5a7b2c64e8f1a0b9d7c3e5f2a4b6c8d0e1f3a5b7c9d1e3f5a7b9c1d3 imagePullPolicy: IfNotPresent ``` > [!CAUTION] > > **Referencing an image grants its publisher access to this namespace** > > The module runs in a Job that can read Secrets and holds cloud credentials in this namespace. Whoever can publish to the image reference can run code with that access. Pin by digest, or use a registry and tag policy you control. See [Security model](). ## Identity reference `spec.identityRef` names the cluster-scoped `TerraformClusterIdentity` whose credentials the object’s Jobs use. See [Identities and Credentials]() for creating an identity and how its credentials reach a Job. | Field | Type | Description | | --- | --- | --- | | `spec.identityRef` | `object` | The identity to use. Whether it is required depends on the kind (below). | | `spec.identityRef.name` | `string` | The `TerraformClusterIdentity` name. **Required** when `spec.identityRef` is set. **Range:** 1 to 253 characters. | How the identity resolves depends on the kind: - A `TerraformCluster` uses only its own `spec.identityRef`. It is required: the webhook rejects a cluster without it, and `spec.defaults.identityRef` does not apply to the cluster itself. A `TerraformClusterTemplate` may leave it to a ClusterClass patch. - A `TerraformMachine` or `TerraformMachinePool` uses its own `spec.identityRef`, else the cluster’s `spec.defaults.identityRef`, else the cluster’s `spec.identityRef`. If none is set, `IdentityAllowed` is `False` with reason `IdentityNotFound`. **Mutability.** Mutable on a `TerraformCluster` and a `TerraformMachinePool`. **Immutable** on a `TerraformMachine`. ```yaml spec: identityRef: name: aws-prod ``` ## Jobs `spec.jobs` tunes the Kubernetes Jobs that run the module: deadlines, lock waits, history, ServiceAccount, resources, environment and security contexts. Every field is optional. It is mutable on all three kinds, including a provisioned `TerraformMachine`, so you can change a stuck machine’s deadline without replacing it. `TerraformCluster` also has `spec.defaults.jobs`, the same type. A `TerraformCluster` ignores its own `spec.defaults.jobs` and uses only `spec.jobs`; machines and pools merge it under their own `spec.jobs`. See [Inheritance](<#inheritance>). | Field | Type | Description | | --- | --- | --- | | `spec.jobs` | `object` | The Job policy. | | `spec.jobs.successfulJobsHistoryLimit` | `integer` | How many succeeded Jobs to keep per object and operation. The newest succeeded Job of each operation is kept even at 0. **Default:** 3, applied at reconcile. **Range:** 0 to 100. | | `spec.jobs.failedJobsHistoryLimit` | `integer` | How many failed Jobs to keep per object and operation. The newest failed Job of an operation is kept even at 0 while no newer Job of that operation succeeded, because retry backoff counts it. **Default:** 3, applied at reconcile. **Range:** 0 to 100. | | `spec.jobs.activeDeadlineSeconds` | `integer` | Bounds a Job’s run time, in seconds. A Job that reaches it is interrupted like an eviction and counts as a failure. **Default:** 3600, applied at reconcile when unset. **Range:** 1 to 86400 (one day). Must be greater than `lockTimeoutSeconds`. | | `spec.jobs.lockTimeoutSeconds` | `integer` | How long the runtime waits for the state lock, passed as `-lock-timeout`. **Default:** 300, applied at reconcile. **Range:** 0 to 3600. Must be less than `activeDeadlineSeconds`. | | `spec.jobs.serviceAccountName` | `string` | Overrides the runner ServiceAccount. **Default:** the controller creates `captf-runner`, bound to the static `captf-runner` ClusterRole. An override must exist and carry the label `captf.io/runner=true`, or no Job is created. **Range:** 1 to 253 characters. | | `spec.jobs.imagePullSecrets` | `[]object` | Pull secrets for the Job pod. They cover both the source image and the runner’s init image. Each item is a Kubernetes `LocalObjectReference` (`name`). **Range:** 1 to 10 items. See [Kubernetes fields](<#kubernetes-fields>). | | `spec.jobs.resources` | `object` | Resource requests and limits of the main container. **Default:** requests of 250m CPU and 512Mi memory, a 2Gi memory limit and no CPU limit. See [Kubernetes fields](<#kubernetes-fields>). | | `spec.jobs.env` | `[]object` | Environment variables added to the main container. Keyed by `name`. **Range:** 1 to 64 items. See [Kubernetes fields](<#kubernetes-fields>). | | `spec.jobs.securityContext` | `object` | Security context of the main container. Hardened defaults apply and the webhook rejects weakening it. See [Kubernetes fields](<#kubernetes-fields>). | | `spec.jobs.podSecurityContext` | `object` | Security context of the Job pod. See [Kubernetes fields](<#kubernetes-fields>). | Jobs never retry pods (`backoffLimit` is 0) and never get a TTL. The controller owns retries and prunes finished Jobs itself, which is what the history limits control. See [Retries]() and [Deadlines](). A tuned job policy ```yaml spec: jobs: activeDeadlineSeconds: 7200 lockTimeoutSeconds: 600 successfulJobsHistoryLimit: 1 failedJobsHistoryLimit: 5 serviceAccountName: team-runner imagePullSecrets: - name: ghcr-pull resources: requests: cpu: "1" memory: 1Gi limits: memory: 4Gi ``` ### Deadline and lock rules The webhook checks one policy at a time, on create and whenever the jobs policy changes: - When both are set, `lockTimeoutSeconds` must be less than `activeDeadlineSeconds`. - When only `activeDeadlineSeconds` is set, it must be greater than the built-in lock timeout, 300. - When only `lockTimeoutSeconds` is set, it must be less than the built-in deadline, 3600. `lockTimeoutSeconds: 4000` alone is rejected. The webhook cannot see a machine’s or pool’s merge with the cluster’s defaults. Reconcile checks the merged policy, and an inconsistent result sets `ApplyJobSucceeded` to `False` with reason `JobPolicyInvalid`; no Job runs until you fix one of the two values. See [Tuning Jobs](). ### Kubernetes fields These fields use Kubernetes’ own types. This page does not list their keys; see the linked Kubernetes API reference for each. **`spec.jobs.env`** A list of [`EnvVar`]() added to the main container (the one that runs your module). An entry whose name starts with `TF_` or `KUBE_` is accepted but silently left out of the Job, with no event: the runner and the Job own that namespace. See [Job Environment](). **`spec.jobs.resources`** A [`ResourceRequirements`]() applied to the main container as a whole. The init container that copies the runner binary is not configurable. See [Default resources](). **`spec.jobs.securityContext`** A [`SecurityContext`]() for the main container. The controller applies these defaults when it builds the Job: `seccompProfile` `RuntimeDefault`, `capabilities` drop `ALL`, `allowPrivilegeEscalation` `false` and `readOnlyRootFilesystem` `true`. `runAsNonRoot` is not defaulted. The container holds cloud credentials, so the webhook rejects `privileged: true`, `allowPrivilegeEscalation: true`, any `capabilities.add`, `readOnlyRootFilesystem: false`, an `Unconfined` seccomp profile, `procMount: Unmasked`, `windowsOptions.hostProcess: true`, and an explicit `runAsUser: 0` or `runAsNonRoot: false`. **`spec.jobs.podSecurityContext`** A [`PodSecurityContext`]() for the Job pod. The controller defaults `seccompProfile` to `RuntimeDefault` and `fsGroup` to the runner’s UID (65532), so a non-root image user can read the 0440 credential files through the group. The webhook applies the same seccomp, root and host-process rules as for the container. **`spec.jobs.imagePullSecrets`** A list of [`LocalObjectReference`]() naming Secrets in the object’s namespace. > [!WARNING] > > **The webhook does not stop an image that runs as root by default** > > It rejects only an explicit root setting. The Job sets neither `runAsUser` nor `runAsNonRoot`, so an image whose own `USER` is root is admitted and runs as root. Set `runAsNonRoot: true` or a non-zero `runAsUser` to require otherwise. See [Tuning Jobs](). ### Inheritance A `TerraformMachine` or `TerraformMachinePool` merges its `spec.jobs` field by field over its cluster’s `spec.defaults.jobs`. A field the object sets wins; an unset one comes from the defaults; a field neither sets gets the built-in default. Defaults are resolved at reconcile time and never persisted, so raising a cluster’s `spec.defaults.jobs` reaches existing machines and pools on their next reconcile, as does a provider upgrade that changes a built-in default. | Field | Merge | | --- | --- | | `spec.jobs.env` | By `name`: the object’s entries win on the same name, then the defaults’ remaining entries. | | `spec.jobs.imagePullSecrets` | Union of both lists, the object’s first, without duplicates. | | `spec.jobs.resources`, `spec.jobs.securityContext`, `spec.jobs.podSecurityContext` | Replaced as a whole: setting one on the object drops the default’s value entirely. | | Every other `spec.jobs` field | The object’s value, else the default’s, else the built-in. | One policy for every machine and pool ```yaml apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1 kind: TerraformCluster metadata: name: prod spec: defaults: jobs: activeDeadlineSeconds: 5400 imagePullSecrets: - name: ghcr-pull ``` See [What a cluster passes to its machines and pools]() and [Tuning Jobs](). ## Variables `spec.variables` passes module variables as an inline JSON object. Each key becomes a named argument of the role module, converted by the module’s declared type. A key the module does not declare fails the apply with “Unsupported argument”. | Field | Type | Description | | --- | --- | --- | | `spec.variables` | `map` | The variables, a JSON object of any values. Inline variables win over every `spec.variablesFrom` source. **Range:** when set, 1 to 256 keys. | Each key must be a Terraform identifier matching `^[a-zA-Z_][a-zA-Z0-9_-]*$`. The webhook also rejects a key that: - starts with `captf_`; - is a contract input of the object’s role (a machine cannot set `machine_name`, but a cluster can); - is a module meta-argument: `source`, `version`, `providers`, `count`, `for_each`, `depends_on`, `lifecycle` or `locals`. A value that is not a JSON object is rejected. Error messages name the key, never the value. **Mutability.** Mutable on a `TerraformCluster` and a `TerraformMachinePool`: a change is part of the inputs hash and re-applies the module (guarded on a cluster by its [apply policy]()). **Immutable** on a `TerraformMachine`. Variables are never inherited from `spec.defaults`. Inline variables ```yaml spec: variables: region: eu-west-1 node_count: 3 tags: team: platform env: prod ``` Inline variables are not sensitive, and every variable is stored in the inputs Secrets and in state. For merge order and the full story, see [Module Variables](). ## Variable sources `spec.variablesFrom` reads module variables from ConfigMaps and Secrets in the object’s own namespace. Every data key of a source becomes a variable of the same name. Sources apply in list order: a later source wins on a shared key, and inline `spec.variables` wins over all of them. | Field | Type | Description | | --- | --- | --- | | `spec.variablesFrom` | `[]object` | The sources. **Range:** 1 to 16 items. The list is replaced as a whole on update. | | `spec.variablesFrom[].configMapRef` | `object` | Names a ConfigMap. Set exactly one of `configMapRef` and `secretRef`; the webhook rejects both or neither. | | `spec.variablesFrom[].configMapRef.name` | `string` | The ConfigMap name. **Required** when `configMapRef` is set. **Range:** 1 to 253 characters. **Allowed values:** a DNS subdomain name, `^[a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*$`. | | `spec.variablesFrom[].secretRef` | `object` | Names a Secret. Set exactly one of `configMapRef` and `secretRef`. | | `spec.variablesFrom[].secretRef.name` | `string` | The Secret name. **Required** when `secretRef` is set. **Range** and **Allowed values:** as for `configMapRef.name`. | | `spec.variablesFrom[].optional` | `boolean` | When true, a missing or unlabeled source contributes nothing. When false, it holds the object at `DependenciesReady` `False` with reason `VariablesSourceNotFound`. **Default:** `false`. | | `spec.variablesFrom[].format` | `string` | How data values are passed. `String` passes each value as a string and lets the module’s declared type convert it (`"3"` to a number, `"true"` to a bool). `JSON` parses each value as JSON, for lists, maps and objects; a value that is not valid JSON is `VariablesInvalid`. **Default:** `String`, applied at reconcile. **Allowed values:** `String`, `JSON`. | The source must carry the label `captf.io/variables=true`. The manager reads and watches only labeled sources, and an unlabeled one counts as missing. Names in a source are checked with the same rules as inline keys, but at reconcile, so a bad key is `VariablesInvalid` naming the key and not a webhook rejection. A variable whose winning value came from a Secret is declared `sensitive = true` in the generated root, so Terraform redacts it in plan and apply output. It is still stored in the inputs Secrets and in state. **Mutability and effect of a change.** | Kind | Field | Source data change | | --- | --- | --- | | `TerraformCluster` | Mutable. | Re-read on every reconcile; a change to a labeled, referenced source wakes the controller and re-applies the module. | | `TerraformMachinePool` | Mutable. | The same as a cluster, but the re-apply is never guarded. | | `TerraformMachine` | **Immutable.** | Read only until the machine is provisioned, then pinned in its inputs Secret; later edits affect only machines created afterward. | ```yaml spec: variablesFrom: - configMapRef: name: network-vars # (1)! format: JSON ``` 1. The ConfigMap must carry `captf.io/variables: "true"`. ```yaml spec: variablesFrom: - secretRef: name: db-credentials optional: true # (1)! ``` 1. A missing Secret contributes nothing and does not block the object. > [!WARNING] > > **A source CAPTF does not own is not carried by `clusterctl move`** > > CAPTF puts no owner reference on a ConfigMap or Secret it reads. After a move, a `TerraformCluster` or `TerraformMachinePool` waits at `VariablesSourceNotFound` until you recreate the source in the target cluster. ## Workspace status Every shared status field is rebuilt from the spec, the state Secret, the durable inputs Secret or the Job list, so none is load-bearing: `clusterctl move` does not restore status, and the controller recomputes it. Each kind adds its own status fields (conditions, outputs and so on) on its own page. | Field | Type | Description | | --- | --- | --- | | `status.observedGeneration` | `integer` | The `metadata.generation` this status was computed for. | | `status.initialization` | `object` | The Cluster API v1beta2 initialization contract. | | `status.activeJob` | `object` | The Job running now, if any. | | `status.lastRun` | `object` | The result of the most recent completed Job. | | `status.lastDriftCheck` | `time` | When the last drift check completed. | | `status.lastRefresh` | `time` | When the last refresh or drift check completed. | | `status.pendingRefreshes` | `integer` | Consecutive pending health samples. | | `status.observedStateSerial` | `integer` | The state serial the outputs were read from. | | `status.lastRestoredSerial` | `integer` | The serial of the last consumed restore Job. | | `status.stateSecretSuffix` | `string` | The state backend’s `secret_suffix`. | | `status.stateBackups` | `[]object` | The kept state backups, newest first. | | `status.source` | `object` | What the last Job ran. | ### Observed generation `status.observedGeneration` is the `metadata.generation` the controller last computed status for. **Range:** 1 or more; unset until the first reconcile. When it is lower than `metadata.generation`, the status describes an older spec. ### Initialization | Field | Type | Description | | --- | --- | --- | | `status.initialization` | `object` | Cluster API’s initialization status. Omitted until it has a field. | | `status.initialization.provisioned` | `boolean` | True once the infrastructure is provisioned. | `status.initialization.provisioned` is derived from state until it first holds, then latched true for the object’s life. It first holds when an apply has completed, the module’s outputs are valid, and the module’s `health` output is present and not `pending`. Cluster API reads it to know when the infrastructure is ready. ### Active Job | Field | Type | Description | | --- | --- | --- | | `status.activeJob` | `object` | The Job currently running for this object. Unset when none is. | | `status.activeJob.name` | `string` | The Job name. **Range:** 1 to 63 characters. | | `status.activeJob.operation` | `string` | The operation the Job runs. **Allowed values:** `apply`, `destroy`, `drift`, `refresh`, `restore`, `plan`. | | `status.activeJob.attempt` | `integer` | The operation’s Job sequence number, the `a` in the Job name, starting at 1. It counts every Job of the operation still retained, not retries: the 40th refresh is attempt 40. **Range:** 1 or more. | | `status.activeJob.startTime` | `time` | When the Job started. | ```yaml status: activeJob: name: prod-apply-a3 operation: apply attempt: 3 startTime: "2026-10-02T09:14:07Z" ``` ### Last run `status.lastRun` is copied from the runner’s termination message when the newest Job finishes, so it holds the outcome of one Job whether it succeeded or failed. | Field | Type | Description | | --- | --- | --- | | `status.lastRun` | `object` | The most recent completed Job. | | `status.lastRun.job` | `string` | The Job name. **Range:** 1 to 63 characters. | | `status.lastRun.operation` | `string` | The operation the Job ran. **Allowed values:** `apply`, `destroy`, `drift`, `refresh`, `restore`, `plan`. | | `status.lastRun.steps` | `[]object` | The runtime commands the runner executed, in order. **Range:** 1 to 16 items. | | `status.lastRun.steps[].name` | `string` | The step name, such as `init`, `validate`, `plan`, `apply` or `apply-refresh-only`. **Range:** 1 to 64 characters. | | `status.lastRun.steps[].exitCode` | `integer` | The step’s exit code. | | `status.lastRun.steps[].durationMilliseconds` | `integer` | The step’s wall time in milliseconds. **Range:** 0 or more. | | `status.lastRun.error` | `object` | Set when the run failed. | | `status.lastRun.error.kind` | `string` | The failure class. **Allowed values:** `image-layout` (the image breaks the image contract), `step` (a runtime step failed), `interrupted` (the step was stopped from outside: a drain, an eviction, a Job deletion or its deadline), `blocked` (an apply stopped before a destructive plan that is not approved; nothing changed), `plan-changed` (an approved plan planned other changes; nothing changed). | | `status.lastRun.error.step` | `string` | The step that failed, for kind `step`. **Range:** 1 to 64 characters. | | `status.lastRun.error.summary` | `string` | The runner’s short description of the failure. **Range:** 1 to 512 bytes. It is not raw stderr: anyone who can get the object can read status, so the full output stays in the Job’s logs. | | `status.lastRun.drift` | `object` | Set when a drift run found changes. | | `status.lastRun.drift.add` | `integer` | Resources the plan would create. **Range:** 0 or more. | | `status.lastRun.drift.change` | `integer` | Resources the plan would update in place. **Range:** 0 or more. | | `status.lastRun.drift.destroy` | `integer` | Resources the plan would destroy, counting replacements. **Range:** 0 or more. | | `status.lastRun.drift.resources` | `[]string` | Addresses of the drifted resources. **Range:** 1 to 20 items of 1 to 512 characters. | ```yaml status: lastRun: job: prod-drift-a12 operation: drift steps: - name: init exitCode: 0 durationMilliseconds: 4120 - name: plan exitCode: 2 durationMilliseconds: 31877 drift: add: 0 change: 1 destroy: 0 resources: - aws_security_group.nodes ``` `error.kind` `blocked` and `plan-changed` are covered in [Destructive plan guard]() and [Manual approval](). ### Drift checks and refreshes | Field | Type | Description | | --- | --- | --- | | `status.lastDriftCheck` | `time` | When the last drift check completed. | | `status.lastRefresh` | `time` | When the last refresh or drift check completed. For a kind whose apply can itself give a definite health reading, it is when that apply finished, because the reading stands in for the refresh after the apply. | | `status.pendingRefreshes` | `integer` | The number of consecutive health samples (completed refresh or drift Jobs) that read `pending` since the last other reading or the last apply. Unset otherwise. **Range:** 1 or more. | While health is `pending`, `status.pendingRefreshes` spaces the refreshes of a `TerraformCluster` or `TerraformMachine`: 30 seconds, then 1, 2 and 4 minutes, and at most 5 minutes. A `TerraformMachinePool` instead refreshes every 30 seconds flat while its health is pending or its membership is converging, whatever the count. The count lives in status only, so after `clusterctl move` it restarts at 0 and the spacing restarts at 30 seconds. A disabled drift interval (`spec.drift.intervalSeconds: 0`) means no drift checks, so `status.lastDriftCheck` does not advance. See [Drift and Health](). ### State | Field | Type | Description | | --- | --- | --- | | `status.observedStateSerial` | `integer` | The Terraform state serial the outputs were read from. **Range:** 1 or more. | | `status.lastRestoredSerial` | `integer` | The backup serial of the last restore Job the controller consumed, whether it succeeded or failed, so a `captf.io/restore-state` annotation naming it does not run again. **Range:** 1 or more. | | `status.stateSecretSuffix` | `string` | The `kubernetes` backend `secret_suffix` of this object’s state. Informational: the controller derives it deterministically and never reads it back. **Range:** 1 to 63 characters. | | `status.stateBackups` | `[]object` | The state backups the controller keeps, newest first, as of the last backup, prune or restore request. **Range:** 1 to 16 items. | | `status.stateBackups[].serial` | `integer` | The state serial the backup holds. Set it as the `captf.io/restore-state` annotation to restore it. **Range:** 1 or more. | | `status.stateBackups[].takenAt` | `time` | When the controller copied the state. | | `status.stateBackups[].bytes` | `integer` | The compressed size of the backup, summed over its Secrets. **Range:** 1 or more. | The number of backups kept comes from the manager’s `--state-backups` flag, which defaults to 5; the field itself holds at most 16. See [Terraform State]() for how and when backups are taken, and the [state restore runbook]() for restoring one. ```yaml status: observedStateSerial: 42 stateSecretSuffix: <16 hex digits>-c # of sha256(namespace/kind/name) stateBackups: - serial: 42 takenAt: "2026-10-02T09:20:51Z" bytes: 18342 - serial: 41 takenAt: "2026-10-01T16:02:13Z" bytes: 18310 ``` ### Image in use `status.source` records what the last Job actually ran, as opposed to what `spec.source` asks for. It is omitted until a Job has run. | Field | Type | Description | | --- | --- | --- | | `status.source` | `object` | What the last Job ran. | | `status.source.image` | `string` | The image reference that ran last, as written in the spec. **Range:** 1 to 512 characters. | | `status.source.imageDigest` | `string` | The digest the container runtime resolved the image to (the pod’s `imageID`). Informational: the pinned copy lives on the durable inputs Secret as `captf.io/image-digest`. **Range:** 1 to 512 characters. | | `status.source.runtimeVersion` | `string` | The version reported by ` version -json`. **Range:** 1 to 64 characters. | ```yaml status: source: image: ghcr.io/example/cluster-module:v1.4.0 imageDigest: ghcr.io/example/cluster-module@sha256:3f1c0e9d5a7b2c64e8f1a0b9d7c3e5f2a4b6c8d0e1f3a5b7c9d1e3f5a7b9c1d3 runtimeVersion: 1.9.8 ``` > [!NOTE] > > **See also** > > - [TerraformCluster](), [TerraformMachine]() and [TerraformMachinePool]() for the kind-specific fields. > - [Module Variables]() for merge order, formats and what a change does. > - [Tuning Jobs]() for choosing job settings. > - [Identities and Credentials]() for creating the identity `spec.identityRef` names. > - [Terraform State]() for the state backend, backups and restore. > - [Drift and Health]() for the checks behind `status.lastDriftCheck` and `status.lastRefresh`. > - [Job Environment]() for the Job’s fixed environment, resources and security contexts. # Annotations, Labels and Finalizers CAPTF reads a few keys that you set to approve a plan, restore a state or release a stuck deletion. It writes many more for its own bookkeeping, owns three finalizers, and honors several Cluster API and clusterctl keys. This page lists all of them, grouped by who uses them. ## The keys you are most likely to use | Key | Put it on | Value | Effect | | --- | --- | --- | --- | | `captf.io/approve-plan` | `TerraformCluster` with `applyPolicy: Manual` | Plan hash from `status.plan.planHash` | Applies that one plan | | `captf.io/approve-destructive-plan` | `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` | Hash from the `DestructivePlanBlocked` message | Lets one blocked apply delete or replace resources | | `captf.io/restore-state` | `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` | Serial from `status.stateBackups` | Pushes that backup back as the state | | `captf.io/abandon-infrastructure` | A deleting `Terraform*` object | The object’s `metadata.uid` | Removes the finalizer without a destroy | | `captf.io/variables` | A `ConfigMap` or `Secret` | `true` | Allows it as a `variablesFrom` source | | `captf.io/runner` | A `ServiceAccount` | `true` | Allows it as a custom runner account | ## Keys you set You set these on your own objects. Approvals and restores are one-shot: CAPTF removes the annotation after it acts, so a later change needs a new one. ### `captf.io/approve-plan` Approves one plan of a `TerraformCluster` whose `spec.applyPolicy` is `Manual`. Only `TerraformCluster` has an `applyPolicy`, so only it accepts this key. | | | | --- | --- | | Kind | `TerraformCluster` | | Value | The plan hash in `status.plan.planHash` | | Removed by CAPTF | Yes, once the apply of that plan succeeds | The apply runs only if it plans exactly the same changes again. If the plan changed, the apply stops, `status.plan` shows the new plan, and the old approval no longer matches. Approving a plan also approves the deletes and replacements it lists, so you do not also need `captf.io/approve-destructive-plan`. ```sh kubectl annotate terraformcluster -n \ captf.io/approve-plan= --overwrite ``` `` is `status.plan.planHash`; the `PlanAwaitingApproval` condition message prints the full command. See [Plan Approval]() and [Manual Plan Approval](). ### `captf.io/approve-destructive-plan` Approves one apply whose plan deletes or replaces resources, or one drift remediation that would. | | | | --- | --- | | Kinds | `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` | | Value | The inputs hash from the `DestructivePlanBlocked` message. On a pool, the approval hash | | Removed by CAPTF | Yes, after an apply of the approved hash succeeds | The value names the inputs the apply renders, not the object. Any later change to the inputs produces a different hash, so the approval stops matching and the next destructive plan is blocked again. A value that does not equal the hash CAPTF rendered approves nothing and is not an error. CAPTF consumes the approval after any successful apply of that hash, even one whose plan was not destructive, so a later destructive remediation of the same inputs needs a new approval. On a pool, CAPTF also removes the annotation when the cluster exports it approved a change of return to the exports the pool already applied. ```yaml metadata: annotations: captf.io/approve-destructive-plan: ``` Approving needs only `patch` on the object. The `ApplyJobSucceeded` condition message prints the exact `kubectl annotate` command. See [The Destructive-Plan Guard](). ### `captf.io/restore-state` Restores a state backup. A restore Job pushes the backup into the backend with `state push -force`, replacing the current state. | | | | --- | --- | | Kinds | `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` | | Value | The serial of a backup listed in `status.stateBackups` | | Removed by CAPTF | Yes, once the restore Job succeeds | A value that is not a serial, or that names no complete backup, sets `RestoreJobSucceeded=False` with reason `RestoreBackupNotFound` and starts nothing. A restore takes precedence over apply, drift and refresh, and no approval gates it. A failed restore is not retried for the same serial: set another serial, or remove the annotation, wait until `status.lastRestoredSerial` clears, and set it again. On an object that is being deleted, the annotation is ignored while the state reads. It runs only while the deletion is held because the state is missing or unreadable; the restore then runs first and the destroy follows. ```sh kubectl annotate terraformcluster -n \ captf.io/restore-state= --overwrite ``` `` is a number from `kubectl get terraformcluster -n -o jsonpath='{.status.stateBackups}'`. See the [state restore runbook]() and [Backups and Restore](). ### `captf.io/abandon-infrastructure` Releases a deletion that cannot finish, without running a destroy. | | | | --- | --- | | Kinds | `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool`, while deleting | | Value | The object’s `metadata.uid`. Any other value is ignored | | Removed by CAPTF | No. The object is deleted with it | The annotation applies when the deletion is held because the state is missing or unreadable, the last destroy failed, or the destroy cannot start: the durable inputs are gone, the identity no longer allows the namespace or no longer exists, or the runner credentials cannot be prepared. CAPTF then runs its normal cleanup, removes the finalizer without a destroy, and records an `InfrastructureAbandoned` `Warning` event naming the cause. An object whose state reads and whose destroy can start is destroyed as usual, even with the annotation set. > [!CAUTION] > > **Abandoning leaves the infrastructure running and untracked** > > CAPTF does not destroy anything. Whatever the module created keeps running and costing money, and the state backups are deleted with the object. Back up what you need and clean up the resources yourself. ```sh kubectl annotate terraformcluster -n \ captf.io/abandon-infrastructure="$(kubectl get terraformcluster \ -n -o jsonpath='{.metadata.uid}')" ``` See [Held Deletions]() and the [stuck destroy runbook](). Prefer this key to removing the finalizer by hand, which skips cleanup; see [Stripping a Finalizer by Hand](). ### `captf.io/variables` A label that opts a `ConfigMap` or `Secret` in as a `variablesFrom` source. | | | | --- | --- | | Put it on | The `ConfigMap` or `Secret` that `spec.variablesFrom` names | | Value | `true` | | Removed by CAPTF | Never. It is a standing opt-in | The manager reads and watches only labeled sources, so an unlabeled source counts as missing. Create it in each namespace that uses it. ```sh kubectl label configmap -n captf.io/variables=true ``` See [Module Variables](). ### `captf.io/runner` A label that opts a custom runner `ServiceAccount` in. | | | | --- | --- | | Put it on | The `ServiceAccount` named by `spec.jobs.serviceAccountName` | | Value | `true` | | Removed by CAPTF | Never | The default `captf-runner` account needs no label. Without the label on a custom account, `RunnerRBACReady` is `False` with reason `ServiceAccountNotOptedIn` and CAPTF creates no Job. The label is consent, not authorization: anyone who can label the account opts it in. ```sh kubectl label serviceaccount -n captf.io/runner=true ``` See [RBAC](). ## Keys CAPTF sets CAPTF writes these for its own bookkeeping and reads them back on the next reconcile. Don’t edit or remove them: a wrong value can cause a repeated apply, a skipped approval or an orphaned Secret. They are listed so you can recognize them in `kubectl get -o yaml` output. ### On your objects and their Machines | Key | Kind | On | Records | | --- | --- | --- | --- | | `captf.io/endpoint-source` | annotation | `TerraformCluster` | Who owns the control-plane endpoint: `user` if one existed before the first apply, `module` if the module created it. Written once, never changed | | `captf.io/remediation-requested` | annotation | `Machine` | That CAPTF set `cluster.x-k8s.io/remediate-machine`; the value is the reason. CAPTF removes only a request it made | | `clusterctl.cluster.x-k8s.io/block-move` | annotation | `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` | That a Job is starting or running, so `clusterctl move` waits. Set before a Job is created and cleared when none is active | ### On state and backup Secrets | Key | Kind | On | Records | | --- | --- | --- | --- | | `captf.io/managed` | label | Everything CAPTF owns: state Secrets, Leases, mirrors, run Jobs, the default runner `ServiceAccount` | Selects the objects the manager caches and sweeps | | `captf.infrastructure.cluster.x-k8s.io/owner-kind` | label | State Secrets, Leases | The owning object’s kind | | `captf.infrastructure.cluster.x-k8s.io/owner-name` | label | State Secrets, Leases | The owning object’s name | | `captf.io/inputs-hash` | annotation | The base state Secret, chunk 0 of a backup | The inputs hash of the state currently adopted | | `captf.io/state-backup` | label | Backup chunk Secrets | Marks a Secret as a state backup chunk | | `captf.io/state-backup-suffix` | label | Backup chunk Secrets | The backend Secret suffix the backup came from | | `captf.io/state-backup-serial` | annotation | Backup chunk Secrets | The backup’s state serial; the value `captf.io/restore-state` names | | `captf.io/state-backup-lineage` | annotation | Backup chunk Secrets | The state’s lineage ID at backup time | | `captf.io/state-backup-taken-at` | annotation | Backup chunk Secrets | When the backup was taken | | `captf.io/state-backup-source-job` | annotation | Backup chunk Secrets | The Job whose apply produced the backed-up state | | `captf.io/state-backup-digest` | annotation | Backup chunk Secrets | A digest of the state, used to detect a conflicting concurrent backup | | `captf.io/state-backup-resources` | annotation | Backup chunk Secrets | The backup’s managed resource count | | `captf.io/state-backup-set` | annotation | Backup chunk Secrets | Groups the chunks of one backup | | `captf.io/state-backup-chunk` | annotation | Backup chunk Secrets | The chunk’s index within its set | | `captf.io/state-backup-chunks` | annotation | Backup chunk Secrets | The set’s total chunk count | The `captf.io/inputs-hash` key is an annotation, not a label. ### On leases and the identity mirror | Key | Kind | On | Records | | --- | --- | --- | --- | | `captf.io/lease` | label | `coordination.k8s.io` Leases | Tells a CAPTF run or cluster lease from the backend’s own lock Lease | | `captf.io/lease-op` | annotation | Leases | The operation the holder runs, for diagnostics | | `captf.io/lease-acquired-at` | annotation | Leases | When the lease was acquired | | `captf.io/mirrored` | label | The identity credential mirror Secret | Marks a Secret as a credential mirror | | `captf.io/source-hash` | annotation | The identity credential mirror Secret | A hash of the source identity’s credentials, so a change rewrites the mirror | ### On the durable inputs Secret The durable inputs Secret holds what CAPTF needs to re-render and destroy an object. Unlike `status`, these annotations move with the Secret. | Key | Kind | On | Records | | --- | --- | --- | --- | | `captf.io/image` | annotation | Durable inputs Secret | `spec.source.image` as last written | | `captf.io/image-digest` | annotation | Durable inputs Secret | The resolved image digest, so a floating tag stays pinned across reconciles | | `captf.io/identity` | annotation | Durable inputs Secret, identity mirror | The `TerraformClusterIdentity` the credentials came from | | `captf.io/applied` | annotation | Durable inputs Secret | `true` once an apply succeeded or a state with an inputs hash was read, so a later missing state reads as lost, not never written | | `captf.io/interrupted-apply` | annotation | Durable inputs Secret of a cluster or pool | An apply Job that was deleted while it ran. It may have applied part of its change, so an apply stays due, and is guarded, until one started after it succeeds | | `captf.io/pending-cluster-outputs` | annotation | A pool’s durable inputs Secret | A change of the cluster’s exports whose pool apply was blocked before a destructive plan, as JSON. The pool keeps applying the exports of its last successful apply until you approve | | `captf.io/partial-cluster-outputs` | annotation | A pool’s durable inputs Secret | A change of the cluster’s exports whose pool apply failed, so the state may hold part of it, as JSON. Every apply is guarded until one succeeds | | `captf.io/applied-cluster-outputs-hash` | annotation | A pool’s durable inputs Secret | The hash of the cluster exports the last successful pool apply rendered. A hash, never an exported value | ### On run Jobs | Key | Kind | On | Records | | --- | --- | --- | --- | | `captf.infrastructure.cluster.x-k8s.io/op` | label | Job and pod | The operation: `apply`, `destroy`, `refresh`, `drift`, `restore` or `plan` | | `captf.infrastructure.cluster.x-k8s.io/attempt` | label | Job and pod | The Job’s attempt number | | `captf.io/bookkept` | annotation | Job | That a finished Job is already counted in `status`, so it is not counted twice | | `captf.io/interrupted` | annotation | Job | That the Job stopped without a clean result, for example it was evicted | | `captf.io/drift-remediation` | annotation | Job | That the apply is a drift remediation, not an ordinary apply | | `captf.io/destructive-plan-blocked` | annotation | Job | That the apply stopped because its plan was destructive and unapproved | | `captf.io/plan-changed` | annotation | Job | That the apply stopped because a re-plan under `Manual` no longer matched the approved plan | | `captf.io/plan-unreadable` | annotation | Job | That the Job’s plan result could not be parsed | | `captf.io/approved-plan` | annotation | Job | The plan hash an apply Job was created to satisfy under `Manual` | | `captf.io/after-failed-apply` | annotation | Job | That the apply started while the newest apply had failed, so an earlier failure still counts and the apply waits for approval instead of being dropped | | `captf.io/after-interrupted-apply` | annotation | Job | The vanished apply Job that this apply started after. Its success removes the `captf.io/interrupted-apply` record | | `captf.io/restore-serial` | annotation | Restore Job | The backup serial the Job pushes | | `captf.io/approval-hash` | annotation | Pool apply Job | The approval hash of a pool apply guarded for a change of the cluster’s exports: what `captf.io/approve-destructive-plan` must name | | `captf.io/cluster-outputs-hash` | annotation | Pool apply Job | The hash of the cluster exports the apply renders | | `captf.io/held-cluster-outputs` | annotation | Pool apply Job | That the apply rendered the exports of the last successful apply while a change waited for approval, so its success does not count as applying that change | ## Finalizers CAPTF puts one finalizer on each workload kind. It blocks deletion until the controller has destroyed the object’s Terraform-managed resources and cleaned up its Secrets. | Finalizer | Kind | Guards | | --- | --- | --- | | `terraformcluster.infrastructure.cluster.x-k8s.io` | `TerraformCluster` | The cluster’s Terraform-managed resources | | `terraformmachine.infrastructure.cluster.x-k8s.io` | `TerraformMachine` | The machine’s Terraform-managed resources | | `terraformmachinepool.infrastructure.cluster.x-k8s.io` | `TerraformMachinePool` | The pool’s Terraform-managed resources | `TerraformClusterIdentity` and the `*Template` kinds have no finalizer. Nothing external depends on an identity directly, so a delete webhook protects an identity that is still in use instead. When a deletion is stuck, see [Held Deletions](); the finalizer is removed by the controller, or by [`captf.io/abandon-infrastructure`](<#captfioabandon-infrastructure>). Each kind’s page under [Resources]() names its finalizer. ## Cluster API and clusterctl keys These keys belong to Cluster API or clusterctl. CAPTF reads or writes them as noted. | Key | Kind | On | CAPTF | | --- | --- | --- | --- | | `cluster.x-k8s.io/cluster-name` | label | `Terraform*` objects, run Jobs, state Secrets, Leases | Reads it to find the owning `Cluster`, and scopes everything it creates with it. Cluster API sets it on the objects it creates | | `cluster.x-k8s.io/remediate-machine` | annotation | `Machine` | Sets it with `captf.io/remediation-requested` when health sampling finds an unhealthy instance, and removes it once the instance reads `Healthy` and only if CAPTF set it. Cluster API then remediates the `Machine` | | `cluster.x-k8s.io/replicas-managed-by` | annotation | `MachinePool` | Sets it to `captf` on an autoscaled pool, so Cluster API stops treating `spec.replicas` as authoritative. Removes only a `captf` value. A foreign value is left alone and CAPTF stops writing replicas back | | `cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size` | annotation | `MachinePool` | You set it. The pool’s minimum replica count | | `cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size` | annotation | `MachinePool` | You set it. The pool’s maximum replica count | | `cluster.x-k8s.io/cloned-from-name` | annotation | `Terraform*` objects | Reads it as the `captf.io/template` tag. Cluster API sets it when it clones from a template | | `clusterctl.cluster.x-k8s.io/move` | label | State Secrets, Leases | Sets it so `clusterctl move` copies them | | `clusterctl.cluster.x-k8s.io/move-hierarchy` | label | The `TerraformClusterIdentity` CRD | Ships it on the CRD so `clusterctl move` brings objects that reference an identity | | `clusterctl.cluster.x-k8s.io/block-move` | annotation | `Terraform*` objects | Sets it while a Job runs, so `clusterctl move` waits | | `clusterctl.cluster.x-k8s.io/delete-for-move` | annotation | `TerraformMachine` | Reads it in the delete webhook. clusterctl sets it when it deletes the source of a move | The two autoscaler annotations go on the Cluster API `MachinePool`, not the `TerraformMachinePool`. You must set both, as non-negative integers with the minimum not above the maximum; otherwise the pool reports `AutoscalingActive=False` and applies without autoscaling. ```yaml apiVersion: cluster.x-k8s.io/v1beta2 kind: MachinePool metadata: name: namespace: annotations: cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "3" cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "10" ``` See [Choose fixed replicas or autoscaling](). CAPTF honors `delete-for-move` only while the owning `Cluster` is paused, which clusterctl sets before it deletes anything. On an unpaused cluster the annotation changes nothing, so setting it by hand does not skip the ordinary delete checks. See [clusterctl move]() and [Machine Remediation](). ## `captf_tags` keys These are not Kubernetes metadata. CAPTF renders them into the `captf_tags` input of every module, as a Terraform map of strings that the module should put on the resources it creates. All six keys are always present. See [Job Inputs](). | Key | Value | | --- | --- | | `captf.io/cluster` | The owning `Cluster`’s name | | `captf.io/namespace` | The owning object’s namespace | | `captf.io/kind` | The owning object’s kind: `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool` | | `captf.io/name` | The owning object’s name | | `captf.io/managed-by` | Always `captf` | | `captf.io/template` | The object’s `cluster.x-k8s.io/cloned-from-name` annotation, or empty when absent | # clusterctl Variables clusterctl replaces shell-style `${NAME}` and `${NAME:=default}` references in the CAPTF templates with environment variables. A variable with no default and no value makes `clusterctl generate` fail, so a missing name never deploys something you did not ask for. The provider ships three templates: | Template | Flavor | What it creates | | --- | --- | --- | | `cluster-template.yaml` | Default (no `--flavor`) | A `Cluster`, `TerraformCluster`, control plane, `MachineDeployment` and `MachineHealthCheck`s | | `cluster-template-clusterclass.yaml` | `clusterclass` | A `Cluster` that references the `noop` `ClusterClass` through `spec.topology` | | `identity.yaml` | None. Applied by an admin | A `TerraformClusterIdentity` and its credentials `Secret` | Both cluster flavors read the same variables, so the first table covers them. The identity template reads two. See [Templates and ClusterClass]() for what each flavor creates and [Quick Start]() for a full walkthrough. ## Cluster flavors: default and `clusterclass` Both templates read every variable below. Variables marked as set by a `clusterctl generate cluster` option need no `export`. ### Identity and module images | Variable | Required | Default | Sets | | --- | --- | --- | --- | | `TERRAFORM_IDENTITY_NAME` | Yes | None | The `TerraformClusterIdentity` the cluster uses, and by default its machines. Default flavor: `TerraformCluster.spec.identityRef.name` and `spec.defaults.identityRef.name`. `clusterclass` flavor: the `identityName` topology variable | | `TERRAFORM_CLUSTER_IMAGE` | Yes | None | The cluster module image: `TerraformCluster.spec.source.image`, or the `clusterImage` topology variable | | `TERRAFORM_MACHINE_IMAGE` | Yes | None | The machine module image for control-plane and worker machines: `spec.template.spec.source.image` of both `TerraformMachineTemplate`s, or the `machineImage` topology variable | ### Cluster shape | Variable | Required | Default | Sets | | --- | --- | --- | --- | | `CLUSTER_NAME` | Yes | None | The name of every object. Set by the name you pass to `clusterctl generate cluster` | | `KUBERNETES_VERSION` | Yes | None | The version of the `KubeadmControlPlane` and `MachineDeployment`, or of the topology. Set by `--kubernetes-version` | | `CONTROL_PLANE_MACHINE_COUNT` | Yes | None | Control-plane replicas. Set by `--control-plane-machine-count`. An odd number of at least 3 keeps a rollout with `maxSurge: 0` available | | `WORKER_MACHINE_COUNT` | Yes | None | `MachineDeployment` replicas. Set by `--worker-machine-count` | | `POD_CIDR` | No | `192.168.0.0/16` | `Cluster.spec.clusterNetwork.pods` | | `SERVICE_CIDR` | No | `10.128.0.0/12` | `Cluster.spec.clusterNetwork.services` | ### Health check timeouts Each timeout is in seconds. They tune the two `MachineHealthCheck`s: one for the control plane and one for workers. In the `clusterclass` flavor they set the same fields on the `Cluster` topology. | Variable | Required | Default | Sets | | --- | --- | --- | --- | | `TERRAFORM_NODE_STARTUP_TIMEOUT` | No | `1200` | Worker `nodeStartupTimeoutSeconds` | | `TERRAFORM_UNHEALTHY_TIMEOUT` | No | `1800` | Worker: seconds of `InfrastructureReady=False` before remediation | | `TERRAFORM_CP_NODE_STARTUP_TIMEOUT` | No | `1800` | Control-plane `nodeStartupTimeoutSeconds`. Longer than the worker value because kubeadm init or join plus kube-vip does more work than a worker join | | `TERRAFORM_CP_UNHEALTHY_TIMEOUT` | No | `3600` | Control-plane unhealthy timeout. Longer because replacing a control-plane machine costs more | > [!WARNING] > > **The unhealthy timeout must exceed one apply plus one drift interval** > > The unhealthy window starts when an apply starts, so set `TERRAFORM_UNHEALTHY_TIMEOUT` above your longest apply time plus one drift interval (30 minutes by default). Below that, Cluster API can replace a machine whose apply is still running. ## Identity template `identity.yaml` is not a flavor. An admin applies it once per set of credentials, with `clusterctl generate yaml`: | Variable | Required | Default | Sets | | --- | --- | --- | --- | | `TERRAFORM_IDENTITY_NAME` | Yes | None | The name of the `TerraformClusterIdentity` and of its credentials `Secret` | | `NAMESPACE` | Yes | None | The namespace allowed to use the identity (`spec.allowedNamespaces.list`) | The generated `Secret` holds placeholder keys. Replace them with the environment variables your module’s providers read. See [Identities and Credentials](). ```sh export TERRAFORM_IDENTITY_NAME= NAMESPACE= clusterctl generate yaml --from templates/identity.yaml | kubectl apply -f - ``` Replace `` with the name clusters reference, and `` with the namespace they live in. ## Generate a cluster Export the required variables that no option sets, then generate. This example uses the default flavor and overrides two optional values: ```sh export TERRAFORM_IDENTITY_NAME= \ TERRAFORM_CLUSTER_IMAGE= \ TERRAFORM_MACHINE_IMAGE= \ POD_CIDR= \ TERRAFORM_UNHEALTHY_TIMEOUT= clusterctl generate cluster --infrastructure terraform \ --target-namespace \ --kubernetes-version \ --control-plane-machine-count 3 --worker-machine-count 2 \ | kubectl apply -f - ``` Placeholders: - ``: an existing `TerraformClusterIdentity` that allows ``. - `` and ``: the module images, for example a tag you built for the cluster and machine roles. - `` and ``: optional. Leave them out to use the defaults in the tables above. - ``, `` and ``: the name, the target namespace and the Kubernetes version, such as `v1.36.3`. `--infrastructure terraform` needs the provider registered with clusterctl. To render a template file directly instead, replace it with `--from templates/cluster-template.yaml`. To use the `clusterclass` flavor, apply the class to the namespace once, enable the `ClusterTopology` feature gate when you register the provider, and add `--flavor clusterclass`: ```sh kubectl apply -n -f templates/clusterclass-noop.yaml clusterctl generate cluster --infrastructure terraform \ --flavor clusterclass --target-namespace \ --kubernetes-version \ --control-plane-machine-count 3 --worker-machine-count 2 \ | kubectl apply -f - ``` `clusterclass-noop.yaml` has no clusterctl variables and no namespace, so apply it to every namespace that generates clusters from it. ## ClusterClass topology variables The `noop` `ClusterClass` declares three topology variables. The `clusterclass` flavor fills them from the clusterctl variables above, in `Cluster.spec.topology.variables`. To change one later, edit the value on the `Cluster`; you do not edit the class or its templates. | Variable | Required | Type | Filled from | Description | | --- | --- | --- | --- | --- | | `identityName` | Yes | `string` | `TERRAFORM_IDENTITY_NAME` | Name of the `TerraformClusterIdentity` the cluster and, by default, its machines use | | `clusterImage` | Yes | `string` | `TERRAFORM_CLUSTER_IMAGE` | Source image of the cluster module (`TerraformCluster.spec.source.image`) | | `machineImage` | Yes | `string` | `TERRAFORM_MACHINE_IMAGE` | Source image of the machine module, for control-plane and worker machines. Changing it rolls the machines out | ```yaml spec: topology: classRef: name: noop variables: - name: identityName value: - name: clusterImage value: - name: machineImage value: ``` A `*Template` kind is immutable once created, so the topology controller rolls a new `machineImage` out by creating a new `TerraformMachineTemplate`. See [Template immutability and rolling out a change]() and [The noop ClusterClass](). # Conditions Conditions are the typed status entries in `status.conditions` of a CAPTF object. Each has a `type`, a `status` (`True`, `False` or `Unknown`), a machine-readable `reason` and a human-readable `message`. The type says what is being tracked, the status says how it stands, and the reason says why. This page lists every type and every reason CAPTF sets. To go from a stuck object to the right fix, start with [Troubleshooting conditions](), which adds the likely cause of each reason. Five kinds carry conditions: `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool`, `TerraformMachineTemplate` and `TerraformClusterIdentity`. Below, “all three” means the first three. The per-kind pages list which types each kind sets: [`TerraformCluster`](), [`TerraformMachine`](), [`TerraformMachinePool`]() and [`TerraformClusterIdentity`](). ## Read conditions `Ready` is the only condition Cluster API reads. Cluster API mirrors it into `InfrastructureReady` on the owning `Cluster`, `Machine` or `MachinePool`. Every other type is informational: it explains why `Ready` is what it is, and nothing outside CAPTF acts on it. List every condition of an object with `jsonpath`: ```sh kubectl get terraformcluster -n \ -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\t"}{.message}{"\n"}{end}' ``` `kubectl describe` prints the same list under `Conditions:`, with the events below it. To see the whole Cluster API tree, with each object’s `Ready` and the reason it is not ready, run `clusterctl describe cluster `. See [Observability]() for how conditions surface in events and metrics. Most types follow normal polarity: `True` is healthy and `False` is a problem. Three types have negative polarity, where `True` is the problem or the unwanted state: `Deleting`, `DriftDetected` and `DeletionBlocked`. `Paused` is a state rather than a health reading: `True` means paused. A type that has no reason for `Unknown` is never `Unknown`. A type that is absent has not been set yet. In the `Ready` summary an absent input counts as `Unknown`, except `Deleting`. > [!NOTE] > > **Waiting is `Unknown`, not `False`** > > Waiting for a lease, for another object or for an approval sets `Unknown`. A wait never turns `Ready` `False`, and it never counts toward retry backoff. ## Summary | Type | Carried by | Polarity | Meaning | | --- | --- | --- | --- | | [`Ready`](<#ready>) | All three, and `TerraformClusterIdentity` | Normal | Summary of the object’s other conditions. The one Cluster API reads. | | [`DependenciesReady`](<#dependenciesready>) | All three | Normal | The owner, cluster, exports, bootstrap data and variable sources the object waits for exist. | | [`IdentityAllowed`](<#identityallowed>) | All three | Normal | The identity exists, allows the namespace and has its Secret. | | [`CredentialsMirrored`](<#credentialsmirrored>) | All three | Normal | The identity’s credentials are copied into the namespace for the Jobs. | | [`RunnerRBACReady`](<#runnerrbacready>) | All three | Normal | The runner ServiceAccount and its RoleBinding exist. | | [`ApplyJobSucceeded`](<#applyjobsucceeded>) | All three | Normal | Outcome of the newest apply or destroy, and why one waits or cannot start. | | [`StateReadable`](<#statereadable>) | All three | Normal | The Terraform state can be read. | | [`RestoreJobSucceeded`](<#restorejobsucceeded>) | All three | Normal | Outcome of a state restore you requested. | | [`OutputsValid`](<#outputsvalid>) | All three | Normal | The module’s outputs satisfy the contract. | | [`InfrastructureHealthy`](<#infrastructurehealthy>) | All three | Normal | The health the module reports. | | [`DriftJobSucceeded`](<#driftjobsucceeded>) | All three | Normal | Outcome of the newest refresh or drift Job. | | [`DriftDetected`](<#driftdetected>) | All three | Negative | The last drift check found differences. | | [`DeletionBlocked`](<#deletionblocked>) | `TerraformCluster` | Negative | A deleting cluster waits for its machines and pools. | | [`EndpointAvailable`](<#endpointavailable>) | `TerraformCluster` | Normal | A valid control-plane endpoint is known. | | [`AutoscalingActive`](<#autoscalingactive>) | `TerraformMachinePool` | Normal | The module owns the pool’s desired count. | | [`CapacityResolved`](<#capacityresolved>) | `TerraformMachineTemplate` | Normal | Capacity and node info were read from the image’s labels. | | [`Paused`](<#paused>) | All three | State | The object or its Cluster is paused. | | [`Deleting`](<#deleting>) | All three | Negative | The object is being deleted. | ## Ready Carried by `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool` and `TerraformClusterIdentity`. Polarity: normal. `Ready` summarizes other conditions; see [Ready summarization](<#ready-summarization>) for which. `True` means every input is `True`. `False` means at least one input is `False`, and `Unknown` means none is `False` but at least one is `Unknown`. The message names the inputs that decide it, and the input’s own reason, such as `Provisioning`, appears there, not in `Ready`’s reason. A `TerraformClusterIdentity` does not summarize: it sets `Ready` from whether its credentials Secret exists. | Status | Reason | Meaning | | --- | --- | --- | | `True` | `Ready` | Every input is healthy. Nothing to do. | | `True` | `SecretFound` | Identity only: the Secret named by `spec.secretRef` exists. | | `False` | `NotReady` | An input is `False`. Read the message, then the named condition below. | | `False` | `SecretNotFound` | Identity only: the Secret does not exist, and objects that use the identity start no Job until it does. See the [identity runbook](). | | `Unknown` | `ReadyUnknown` | No input is `False`, but one is `Unknown`: the object waits for its owner, a first apply, a lease or an approval. Usually transient. | ## DependenciesReady Carried by all three. Polarity: normal. The gates that must open before CAPTF renders inputs or starts a Job: the owner, the `Cluster`, the cluster’s infrastructure and exports, bootstrap data and variable sources. While a gate is closed, nothing is rendered or run. A deleting object skips the owner gates, so its destroy still runs. | Status | Reason | Meaning | | --- | --- | --- | | `True` | `DependenciesReady` | Every gate is open. | | `False` | `ClusterNotTerraform` | The owning `Cluster`’s `infrastructureRef` is not a `TerraformCluster`. Nothing is rendered or run. Point the `Cluster` at a `TerraformCluster`, or move the object to the right one. | | `False` | `OwnerMismatch` | An `ownerReference` of the expected kind resolves to an object that does not reference this one back: a wrong or missing `infrastructureRef`, a UID mismatch, or a `cluster.x-k8s.io/cluster-name` label that disagrees with the owner. CAPTF treats the reference as forged or stale, so the object has no valid owner, no Job runs and nothing is written to the named owner or its `Cluster`. The message says which check failed. | | `False` | `OwnerNotFound` | The owner object is gone. Delete this object if it is orphaned. See [deletion order](). | | `False` | `VariablesInvalid` | A `variablesFrom` source has a key that is not a Terraform identifier or is reserved, or a value that is not UTF-8 or, for format `json`, not valid JSON. The message names the key, never the value. No Job starts until you correct the source. See [module variables](). | | `False` | `VariablesSourceNotFound` | A ConfigMap or Secret named in `spec.variablesFrom`, not marked optional, is missing or lacks the `captf.io/variables=true` label. No Job starts. See [module variables](). | | `False` | `WaitingForOwnerMachine` | A fresh `TerraformMachine` has only a non-controller control-plane `ownerReference` and no `Machine` one yet. Wait for Cluster API to set it. | | `False` | `WaitingForOwnerMachinePool` | A `TerraformMachinePool` has `ownerReferences` but none to a `MachinePool` yet. Wait for Cluster API to set it. | | `Unknown` | `WaitingForOwner` | The owner is not set yet. Wait, and check the object’s `cluster.x-k8s.io/cluster-name` label. | | `Unknown` | `WaitingForClusterInfrastructure` | The `TerraformCluster` is not provisioned yet. Read its `ApplyJobSucceeded` and `Ready`. | | `Unknown` | `WaitingForClusterExports` | The cluster module’s `exports` output is not readable. Check the `TerraformCluster`’s [`StateReadable`](<#statereadable>) and [`OutputsValid`](<#outputsvalid>). | | `Unknown` | `WaitingForBootstrapData` | The `Machine`’s or `MachinePool`’s bootstrap data Secret is not set or not found. The bootstrap provider creates it, not CAPTF. See [control planes](). | ## IdentityAllowed Carried by all three. Polarity: normal. Whether the resolved `TerraformClusterIdentity` exists, allows this namespace and has its credentials Secret. While it is not `True`, no Job starts, except that a deletion that needs no Job still finishes. See [Identities]() and the [identity runbook](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `IdentityAllowed` | The identity exists, allows the namespace and its Secret exists. | | `False` | `IdentityNotFound` | No identity resolved: none is set on the object or the cluster defaults, or the named one does not exist. [Fix](). | | `False` | `NamespaceNotAllowed` | The identity’s `allowedNamespaces` excludes this namespace. CAPTF revokes the mirror. [Fix](). | | `False` | `SecretNotFound` | The identity allows the namespace but its credentials Secret does not exist. [Fix](). | | `Unknown` | `IdentityCheckFailed` | CAPTF could not finish the check, for example because an API read failed. It usually clears by itself. [Fix](). | ## CredentialsMirrored Carried by all three. Polarity: normal. Whether the identity’s credentials are copied into the object’s namespace as the mirror Secret that its Jobs mount. | Status | Reason | Meaning | | --- | --- | --- | | `True` | `Mirrored` | The mirror exists and is current. | | `False` | `MirrorFailed` | Creating or updating the mirror failed, for example because a Secret with the mirror’s name belongs to something else. [Fix](). | | `Unknown` | `MirrorPending` | No mirror exists yet, either before the first mirror or because [`IdentityAllowed`](<#identityallowed>) is not `True`. Fix that first. [Details](). | ## RunnerRBACReady Carried by all three. Polarity: normal. Never `Unknown`; until the controller first reaches it, the condition is absent and counts as `Unknown` in `Ready`. Whether the runner ServiceAccount exists and is bound to the runner role in the object’s namespace. See [RBAC](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `RBACReady` | The ServiceAccount and RoleBinding are in place. | | `False` | `RBACFailed` | Creating the ServiceAccount or the `captf-runner` RoleBinding failed, for example because a RoleBinding with that name is not CAPTF’s. [Fix](). | | `False` | `ServiceAccountNotOptedIn` | An override ServiceAccount lacks the `captf.io/runner=true` label. Label it, create it or drop the override; the controller retries every 30 seconds. [Fix](). | ## ApplyJobSucceeded Carried by all three. Polarity: normal. The outcome of the newest apply, plan or destroy Job, and the reason one waits or cannot start. A plan Job stands for the apply it plans. After one has completed the status is not `Unknown`, except while the next operation waits for a lease or an approval. A running apply keeps the last result. The three lease reasons are shared with [`DriftJobSucceeded`](<#driftjobsucceeded>) and [`RestoreJobSucceeded`](<#restorejobsucceeded>). To diagnose a failure, read `status.lastRun` and the Job’s pod logs; see [Failing Jobs](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `ApplySucceeded` | The newest apply succeeded. | | `True` | `DestroySucceeded` | The destroy succeeded and cleanup follows. | | `False` | `ApplyFailed` | An apply Job failed, or disappeared while it ran because it was deleted before it finished. The message names the Job and the failed step. The controller retries with backoff. After a Job vanished, an apply of the current inputs stays due until one succeeds; this message outranks an older blocked or plan-changed apply. See [retries](). | | `False` | `DestroyFailed` | A destroy Job failed, or a destroy cannot be rendered because the durable inputs Secret is missing. It retries until it succeeds. If it cannot, see [stuck destroy](). | | `False` | `DestructivePlanBlocked` | An apply stopped before a plan that deletes or replaces resources, and no approval names it. On a `TerraformCluster`, which includes a drift remediation, the `captf.io/approve-destructive-plan` annotation does not name the inputs hash it renders, and no apply of that hash runs until it does or the inputs change. On a `TerraformMachinePool`, it is an apply of a change of the cluster’s exports that stopped, and the annotation does not name its approval hash (the inputs hash without `bootstrap_data`). The reason stays while the change waits. The pool keeps applying everything else with the exports of its last successful apply, and the condition reports that apply, `True`, again once the change is approved, or the exports change or return to the applied ones. A pool that cannot fall back to those exports waits for approval as a cluster does, and the message says why. See the [destructive-plan guard](). | | `False` | `IdentityNotAllowed` | No Job could be created because the identity does not allow the namespace or is gone. Allow the namespace again, or [abandon]() a deleting object. | | `False` | `ImageInvalid` | The runner reported an image-layout error: no `/captf/module` or a non-executable command. Rebuild the image to the [image contract](). | | `False` | `ImagePullFailed` | The pod stayed in `ErrImagePull` or `ImagePullBackOff` past `activeDeadlineSeconds`. See [image pull failures](). | | `False` | `InputsTooLarge` | The rendered root module and variables exceed what a Secret can carry, so no Job starts. See [size limits](). | | `False` | `JobDeadlineExceeded` | The Job hit `activeDeadlineSeconds`. See [deadlines](). | | `False` | `JobPolicyInvalid` | The effective job policy, the object’s own merged over the cluster defaults and the built-in defaults, gives a `lockTimeoutSeconds` that is not below `activeDeadlineSeconds`, so no Job starts. See [the merged-policy check](). | | `Unknown` | `NoApplyYet` | No apply has completed yet. Check [`DependenciesReady`](<#dependenciesready>), [`IdentityAllowed`](<#identityallowed>) and the Jobs. | | `Unknown` | `PlanAwaitingApproval` | A `TerraformCluster` with `applyPolicy: Manual` waits for approval of the plan in `status.plan`, through the `captf.io/approve-plan` annotation naming its hash. See [manual approval](). | | `Unknown` | `PlanChanged` | An approved apply planned other changes than the approved plan and stopped before applying them. The new plan in `status.plan` waits for approval. See [manual approval](). | | `Unknown` | `WaitingForClusterOperation` | A machine’s or pool’s apply or destroy waits for its `TerraformCluster`’s to finish. | | `Unknown` | `WaitingForMachineOperations` | A `TerraformCluster`’s apply or destroy waits for the applies and destroys of its machines and pools in flight. New ones wait behind it. | | `Unknown` | `WaitingForRunLease` | Another live Job, such as one another manager instance started, holds the object’s run lease. No Job starts until it finishes. Never delete a Lease by hand. See [leases](). | ## StateReadable Carried by all three. Polarity: normal. Whether CAPTF can read the object’s Terraform state from its backend Secrets. While it is `False`, no Job runs, and a deleting object’s deletion is held, except for `StateLocked`. See [unreadable state]() and [State](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `StateRead` | The state was read. | | `False` | `StateCorrupt` | The state cannot be decoded. Restore a backup taken before the damage. [Fix](). | | `False` | `StateEncrypted` | The state carries OpenTofu client-side encryption, which CAPTF does not support in v1. [Fix](). | | `False` | `StateInconsistent` | The state chunks disagree and do not form one complete state. [Fix](). | | `False` | `StateLocked` | Something other than the object’s own runner, such as a workstation, holds the state lock. Every Job waits `lockTimeoutSeconds` for it and then fails. See the [stale-lock runbook](). | | `False` | `StateLost` | The state Secret of an object that applied before is missing, or carries no inputs hash for an immutable kind. An object counts as having applied if it is provisioned, marked applied, digest-pinned or backed up on its own Secrets, which survive a `clusterctl move`. No Job runs until you restore the state. A deleting object keeps its finalizer until you restore it with `captf.io/restore-state` or abandon the infrastructure with `captf.io/abandon-infrastructure`. [Fix](). | | `Unknown` | `StateNotFound` | No state exists yet. The first apply creates it. [Details](). | ## RestoreJobSucceeded Carried by all three. Polarity: normal. Set only once you request a restore with the `captf.io/restore-state` annotation. Never feeds `Ready`. The outcome of the newest state restore Job. A requested serial without a backup, or a restore that waits for a lease, is reported here too. See [state restore](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `StateRestored` | The restore Job pushed the backup into the backend. | | `False` | `RestoreBackupNotFound` | The annotation names no existing backup, or is not a serial. No Job starts. Pick a serial from `status.stateBackups`. | | `False` | `RestoreFailed` | The restore Job failed. CAPTF does not retry it for the same serial, in `status.lastRestoredSerial`. To retry, remove the annotation, wait for `lastRestoredSerial` to clear and set it again. | | `Unknown` | `WaitingForClusterOperation` | A machine’s or pool’s restore waits for its `TerraformCluster`’s operation to finish. | | `Unknown` | `WaitingForMachineOperations` | A `TerraformCluster`’s restore waits for the operations of its machines and pools in flight. | | `Unknown` | `WaitingForRunLease` | Another live Job holds the object’s run lease. No Job starts until it finishes. | ## OutputsValid Carried by all three. Polarity: normal. Whether the module’s outputs, read from state, satisfy the contract for the object’s role and the Cluster API field markers. See the [module contract](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `OutputsValid` | The outputs are valid. | | `True` | `InstancesTruncated` | A pool’s `instances` output had more entries than the controller keeps, and CAPTF shortened it. The pool still provisions and nothing else about its outputs is invalid. See the [`MachinePool` role](). | | `False` | `FailureDomainMismatch` | A machine’s `failure_domain` output differs from the failure domain it requested. | | `False` | `OutputsInvalid` | An output violates the contract or a Cluster API marker. The message names it. | | `False` | `OutputsMissing` | A required output is not declared. Check the module with [`tfcapi-lint`](). | | `False` | `ProviderIDChanged` | `provider_id` changed after it was first written. It is immutable. Delete the `Machine` to replace the instance. | | `Unknown` | `OutputsPending` | Required outputs are `null`, so the first apply has not produced them yet. | ## InfrastructureHealthy Carried by all three. Polarity: normal. The health the module reports, read from state at each refresh. After provisioning it is the input that moves `Ready`. See [Drift and health](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `Healthy` | The health state is `running` and healthy. | | `False` | `InstanceDegraded` | The module reports `degraded`. | | `False` | `InstancePending` | The module reports `pending`. Refresh runs on a doubling schedule up to five minutes. | | `False` | `InstanceStopped` | The module reports `stopped`. | | `False` | `InstanceTerminated` | The module reports `terminated`, or the instance vanished. See [terminated instances](). | | `False` | `InstanceUnhealthy` | The state is `running` but healthy is `false`. | | `False` | `Provisioning` | The first apply started and the object is not provisioned yet. This starts the clock for a MachineHealthCheck. | | `Unknown` | `HealthUnknown` | The module reports `unknown`, an unrecognized state or no health at all. | | `Unknown` | `ProviderIDMissing` | `provider_id` turned `null` after provisioning, for the first time. CAPTF reports the instance terminated only if the next sample is `null` too. | | `Unknown` | `WaitingForProvisioning` | Before the first apply. | To replace an unhealthy machine, see [machine remediation](). ## DriftJobSucceeded Carried by all three. Polarity: normal. Never feeds `Ready`. The outcome of the newest refresh or drift Job. See [Drift](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `DriftChecked` | The newest refresh or drift Job succeeded. | | `False` | `DriftJobDeadlineExceeded` | The drift Job hit `activeDeadlineSeconds`. Raise the deadline. | | `False` | `DriftJobFailed` | The drift Job failed. The check retries with backoff. Read the Job’s logs. | | `Unknown` | `DriftJobRunning` | A drift Job runs. | | `Unknown` | `DriftNotChecked` | No drift check has completed, or drift checks are disabled. | | `Unknown` | `DurableInputsMissing` | A refresh or drift Job is due but has nothing to run against: the durable inputs Secret `captf-inputs--` is gone, and the object renders no current inputs in its place (a `TerraformMachine`, which is immutable, never does). No refresh or drift check runs until you restore the Secret, so `InfrastructureHealthy` keeps its last reading and drift goes unchecked. | | `Unknown` | `WaitingForRunLease` | A refresh or drift Job waits for the object’s run lease, which another live Job holds. | ## DriftDetected Carried by all three. Polarity: negative. Never feeds `Ready`. Whether the last drift check found the infrastructure differing from the desired inputs. `True` is the problem. See [Drift](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `DriftPending` | Remediation is pending, or a cluster re-apply failed. The message names a failed, blocked or changed remediation Job. See [the remediation cap](). | | `True` | `DriftRemediating` | A remediation apply runs. | | `True` | `DriftReported` | The drift action is `Report`, so CAPTF reports the drift and changes nothing. Accept it or set the action to `Remediate`; see [Remediate drift](). | | `False` | `NoDrift` | The last check found no drift. | | `Unknown` | `DriftNotChecked` | No drift check has completed. CAPTF sets this on the first visit. | ## DeletionBlocked Carried by `TerraformCluster`. Polarity: negative. Never feeds `Ready`. Whether a deleting `TerraformCluster` waits for dependents. CAPTF sets it on the first visit as `False`. | Status | Reason | Meaning | | --- | --- | --- | | `True` | `DependentsExist` | `TerraformMachines` or `TerraformMachinePools` of the cluster still exist, so no destroy runs. The message names them. See [the cluster waits for its machines](). | | `False` | `NotBlocked` | Nothing blocks the deletion. | ## EndpointAvailable Carried by `TerraformCluster`. Polarity: normal. Never feeds `Ready`. Not set before the cluster is provisioned. Whether a valid control-plane endpoint is known. See the [cluster role](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `EndpointAvailable` | A valid endpoint exists. | | `False` | `WaitingForEndpoint` | The cluster is provisioned, and neither the module’s `control_plane_endpoint` output nor `Cluster.spec.controlPlaneEndpoint` has a valid endpoint. Output one from the module or set it on the `Cluster`. | ## AutoscalingActive Carried by `TerraformMachinePool`. Polarity: normal. Never `Unknown`. Never feeds `Ready`, and Cluster API does not mirror it to the `MachinePool`. Whether the pool’s autoscaler annotations are present and valid, so the module owns the desired count. See [Machine pools](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `ReplicasManagedByModule` | Both annotations are present and valid, and the controller writes the observed replicas back to `MachinePool.spec.replicas`. | | `False` | `AutoscalingAnnotationsInvalid` | An annotation is present but the pair is incomplete, unparsable or has min above max. The message names the problem. The pool still applies without autoscaling. | | `False` | `AutoscalingDisabled` | Neither annotation is set. `MachinePool.spec.replicas` is the only source of desired capacity. | | `False` | `ReplicasManagedExternally` | Both annotations are valid, but the `MachinePool`’s `cluster.x-k8s.io/replicas-managed-by` annotation names another controller. CAPTF does not write the observed replicas back, so the two controllers do not fight over `spec.replicas`. Choose one owner. | ## CapacityResolved Carried by `TerraformMachineTemplate`, which has no `Ready` condition. Polarity: normal. Never `Unknown`. Whether the template’s capacity and node info were read from the module image’s labels. See [Templates](). | Status | Reason | Meaning | | --- | --- | --- | | `True` | `CapacityNotDeclared` | The image carries neither label. Fine unless a ClusterClass autoscaler needs the capacity. | | `True` | `CapacityResolved` | Both labels parsed. | | `False` | `CapacityLabelInvalid` | A label is present but invalid. Fix the label in the image; see the [image contract](). | | `False` | `ImageInspectFailed` | The registry fetch or authentication failed. Fix the image reference or the credentials. The controller also emits a `Warning` event. | ## Paused Carried by all three. Polarity: a state, not a health reading. Never `Unknown`. Never feeds `Ready`. Whether reconciliation is paused. A paused object does bookkeeping only and starts no Job. CAPTF sets it on the first visit. | Status | Reason | Meaning | | --- | --- | --- | | `True` | `Paused` | The object or its `Cluster` is paused: `spec.paused` on the `Cluster`, or the `cluster.x-k8s.io/paused` annotation on the object. `clusterctl move` pauses on purpose. A deletion waits until you unpause; see [pause stops a deletion](). | | `False` | `NotPaused` | The object is not paused. | ## Deleting Carried by all three. Polarity: negative. Never `Unknown`. Whether the object is being deleted. `True` makes `Ready` `False`. The message can say what the destroy waits for. CAPTF sets it on the first visit. | Status | Reason | Meaning | | --- | --- | --- | | `True` | `Deleting` | The `deletionTimestamp` is set and no more specific reason applies. Wait for the destroy. If the object does not delete, follow [my object will not delete](). | | `False` | `NotDeleting` | The `deletionTimestamp` is not set. | ## Ready summarization `Ready` is built with the Cluster API `SetSummaryCondition` helper from the inputs below. Which inputs it uses depends on the kind and on whether `status.initialization.provisioned` has latched `True`; it latches once and stays. `Deleting` has negative polarity, so `True` there makes `Ready` `False`. Before provisioning, all three kinds summarize the same nine inputs: - `DependenciesReady` - `IdentityAllowed` - `CredentialsMirrored` - `RunnerRBACReady` - `ApplyJobSucceeded` - `StateReadable` - `OutputsValid` - `InfrastructureHealthy` - `Deleting` After provisioning, the inputs differ per kind: | Kind | Inputs after provisioning | | --- | --- | | `TerraformCluster` | `InfrastructureHealthy`, `Deleting` | | `TerraformMachine` | `InfrastructureHealthy`, `Deleting` | | `TerraformMachinePool` | `InfrastructureHealthy`, `ApplyJobSucceeded`, `Deleting` | > [!WARNING] > > **A failed re-apply does not flip a cluster’s or machine’s `Ready`** > > After provisioning, a `TerraformCluster`’s `Ready` ignores `ApplyJobSucceeded`. If it did not, a failed re-apply would flip the `Cluster`’s `InfrastructureReady` and suspend every MachineHealthCheck of the cluster. A `TerraformMachine` is immutable and behaves the same. A `TerraformMachinePool` is mutable and re-applied regularly, so a failed re-apply shows in its `Ready`. These types never feed `Ready`: [`Paused`](<#paused>), [`DriftDetected`](<#driftdetected>), [`DriftJobSucceeded`](<#driftjobsucceeded>), [`DeletionBlocked`](<#deletionblocked>), [`EndpointAvailable`](<#endpointavailable>), [`RestoreJobSucceeded`](<#restorejobsucceeded>) and [`AutoscalingActive`](<#autoscalingactive>). An input that has not been set yet counts as `Unknown`, so a new object is `Ready: Unknown`, never `True`. Only `Deleting` is ignored while missing. `Ready`’s reason is always `Ready`, `NotReady` or `ReadyUnknown`. A `TerraformClusterIdentity` does not summarize. It sets its own `Ready` from whether its credentials Secret exists (`SecretFound` or `SecretNotFound`). > [!NOTE] > > **See also** > > - [Troubleshooting conditions]() > - [Troubleshooting: start here]() > - [Events]() > - [Observability]() > - [Runbooks]() > - [Drift and health]() > - [Annotations and labels]() # Events CAPTF records a Kubernetes Event for each step in the life of an object: the manager once per transition, and a Job’s runner in real time while the Job runs. This page lists every reason, grouped by situation. Jump to: [lifecycle](<#lifecycle-and-readiness>), [Jobs](<#jobs>), [ordering and leases](<#ordering-and-leases>), [plans and approvals](<#plans-and-approvals>), [state](<#state>), [drift and health](<#drift-and-health>), [identity and credentials](<#identity-and-credentials>), [deletion](<#deletion>), [pools and templates](<#pools-and-templates>), [runner events](<#runner-events>). ## Reading events Events are `events.k8s.io/v1` objects recorded on the CAPTF object they are about. Read them in order with `kubectl events`: ```sh kubectl events --for terraformcluster/ -n kubectl events --for terraformmachine/ -n kubectl events --for terraformmachinepool/ -n kubectl events --for terraformmachinetemplate/ -n kubectl events --for terraformclusteridentity/ -n default ``` `kubectl describe` shows the same events at the end of its output. A `TerraformClusterIdentity` is cluster-scoped, so its events land in the `default` namespace. Add `-o wide` to see which events the manager recorded (`captf-manager`) and which a runner did (`captf.io/runner`). Three rules hold for every event: - The manager emits an event once per transition or occurrence, never on every reconcile. A reason that appears again means the thing happened again. - A condition that first appears emits an event only when it starts in its bad state, so a new object’s first reconcile is quiet. - Notes never carry credentials, tfvars, outputs, plan values or raw stderr. Where a table says “any provisioned kind”, the event is recorded on a `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`. The **Type** column is the Kubernetes event type: `Normal` reports progress, `Warning` reports something that needs a look. For a Warning, the **Action** column says what to check next; the condition that backs the event is in [Conditions](). > [!NOTE] > > **Events expire** > > Kubernetes keeps events for about an hour by default. For history, read the object’s conditions and `status.lastRun`, or ship events to your log pipeline. ## Lifecycle and readiness | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `Provisioned` | Normal | any provisioned kind | `Provisioned` latched true: the first apply finished and the infrastructure exists. | None. | | `InputsChanged` | Normal | any provisioned kind | The inputs hash differs from the state’s and an apply of the new inputs starts. | None. See [Inputs](). | | `ControlPlaneEndpointSet` | Normal | `TerraformCluster` | `spec.controlPlaneEndpoint` was written from the module output. | None. | | `FailureDomainsChanged` | Normal | `TerraformCluster` | `status.failureDomains` changed. | None. | | `ProviderIDSet` | Normal | `TerraformMachine`, `TerraformMachinePool` | `spec.providerID` was written. A pool’s value can change later. | None. | | `Paused` | Normal | any provisioned kind | The `Paused` condition became True. Reconciliation stops starting Jobs. | Resume when ready: clear the pause on the object or its `Cluster`. | | `Resumed` | Normal | any provisioned kind | `Paused` went from True to False. Reconciliation resumes. | None. | | `OutputsInvalid` | Warning | any provisioned kind | The module’s outputs broke the contract. | Fix the module output named in the note. See [`OutputsValid`](). | | `ConditionChanged` | Normal or Warning | any provisioned kind | Any owned condition without a more specific reason changed status or reason. Warning when it moved into its bad state, Normal otherwise. | For a Warning, read the condition named in the note. | | `DigestPinned` | Normal | any provisioned kind | An image digest was recorded on the durable inputs Secret, or pinned again after an apply of a mutable kind. | None. | | `DigestUnknown` | Warning | any provisioned kind | No digest could be pinned, or an operation runs the `spec` reference for lack of one. | Check that the registry is reachable and the reference resolves. | ## Jobs | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `JobCreated` | Normal | any provisioned kind | A Job started. The note gives the operation, attempt, image and why. | None. | | `JobSucceeded` | Normal | any provisioned kind | An apply, destroy, refresh or drift Job succeeded. | None. | | `JobFailed` | Warning | any provisioned kind | A Job failed, or an apply or destroy could not start. | Read the Job logs and `status.lastRun`. See the [job failures runbook](). | | `JobDeadlineExceeded` | Warning | any provisioned kind | A Job hit `activeDeadlineSeconds`. | Find the slow step in the runner events, then raise the deadline or fix the module. See the [slow jobs runbook](). | | `JobInterrupted` | Warning | any provisioned kind | Something outside CAPTF stopped the Job, such as a node drain, an eviction or a deletion. CAPTF retries without backoff. | None if it was planned. Otherwise check node pressure and preemption. | | `StuckJobDeleted` | Warning | any provisioned kind | A Job that could never start, because its per-run Secret is missing, was deleted so it can start again. | None. Repeats mean something removes the Secret. | ## Ordering and leases An object runs one operation at a time, and a machine waits for its cluster. These events say what an operation waits for. See [Leases](). | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `WaitingForRunLease` | Normal | any provisioned kind | Another live Job holds the object’s run lease. | Wait. If no Job is running, see the [stale lock runbook](). | | `WaitingForClusterOperation` | Normal | `TerraformMachine`, `TerraformMachinePool` | A machine’s apply or destroy waits for its `TerraformCluster`’s apply or destroy. | Wait. Check the cluster if it never ends. | | `WaitingForMachineOperations` | Normal | `TerraformCluster` | A cluster’s apply or destroy waits for its machines’ applies and destroys in flight. | Wait. Check the machines if it never ends. | ## Plans and approvals These events belong to the destructive-plan guard and to `applyPolicy: Manual`. See [Approvals](). | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `PlanReady` | Normal | `TerraformCluster` | A plan Job planned a change under `applyPolicy: Manual`. The note has the counts, the plan hash and the approve command. Once per plan Job. | Review the plan, then approve it. | | `PlanApproved` | Normal | `TerraformCluster` | The apply of an approved plan started. | None. | | `PlanApplied` | Normal | `TerraformCluster` | The approved plan was applied and its approval annotation removed. | None. | | `PlanChanged` | Warning | `TerraformCluster` | An approved apply planned other changes and stopped before applying them. Once per such Job. | Review the new plan and approve it again. | | `DestructivePlanBlocked` | Warning | `TerraformCluster` | An apply stopped before a plan that deletes or replaces resources. Once per blocked Job, in place of `JobFailed`. | Read the plan and approve it if the deletes are intended. See [Destructive guard](). | | `DestructivePlanApprovalConsumed` | Normal | `TerraformCluster` | The approved destructive apply succeeded and its approval annotation was removed. | None. | `DestructivePlanBlocked` is also recorded for a `TerraformMachinePool` apply that would change the cluster’s exports. The pool keeps applying with the exports of its last successful apply until you approve the change. If an earlier apply may have left such a change partly applied, the pool waits for the approval, as a cluster does. ## State See [State]() and [Backups](). | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `StateAdopted` | Normal | any provisioned kind | The state written by a successful apply was adopted with its new inputs hash. | None. | | `StateBackedUp` | Normal | any provisioned kind | A new state serial was copied into a backup. Once per backup. | None. | | `StateRestored` | Normal | any provisioned kind | A restore Job pushed a backup into the backend and the `captf.io/restore-state` annotation was removed. | None. | | `StateRestoreFailed` | Warning | any provisioned kind | A restore Job failed. CAPTF does not retry it for the same serial. | Read the restore Job logs. See the [state restore runbook](). | | `StateLocked` | Warning | any provisioned kind | Something else holds the state lock (`StateReadable` False, reason `StateLocked`). | Wait, or find the holder. See the [stale lock runbook](). | | `ForceUnlocked` | Warning | any provisioned kind | CAPTF force-unlocked a stale state lock. | Find out why the previous Job died. Alert: [`CAPTFForceUnlocks`](). | | `StateUnreadable` | Warning | any provisioned kind | The state could not be read. | See the [state unreadable runbook](). Alert: [`CAPTFStateUnreadable`](). | | `StateLost` | Warning | any provisioned kind | A provisioned object’s state is gone or carries no inputs hash (`StateReadable` False, reason `StateLost`). | Restore from a backup. See the [total state loss runbook](). | | `OwnerReferencesRepaired` | Normal | any provisioned kind | Secrets of the object (state, backups, durable inputs, plan key or its credential mirror entry) had no owner reference, or one to an earlier UID, as a management-cluster restore leaves them. They are owned by the object again. | None. See the [move runbook](). | ## Drift and health See [Drift and health](). | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `DriftDetected` | Warning | any provisioned kind | A drift check found a difference. Once per finding. | Read the drift plan summary. With `drift.action: Report` you decide; with `Remediate` CAPTF applies. Alert: [`CAPTFClusterDrift`](). | | `DriftRemediationStarted` | Normal | any provisioned kind | An apply that remediates drift started. | None. | | `DriftResolved` | Normal | any provisioned kind | `DriftDetected` went from True to False. | None. | | `InstanceHealthy` | Normal | any provisioned kind | `InfrastructureHealthy` became True. | None. | | `InstanceUnhealthy` | Warning | any provisioned kind | `InfrastructureHealthy` became False for an unhealthy, degraded, stopped or terminated instance. | Check the instance in the cloud console. See [`InfrastructureHealthy`](). | | `RemediationRequested` | Warning | `TerraformMachine` | CAPTF annotated the owner `Machine` with `cluster.x-k8s.io/remediate-machine`. | Cluster API replaces the machine if a `MachineHealthCheck` acts on it. | | `RemediationWithdrawn` | Normal | `TerraformMachine` | The instance read healthy again and CAPTF removed the annotation it set. | None. | ## Identity and credentials See [Identity and credentials](). | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `IdentityNotAllowed` | Warning | any provisioned kind | The identity does not allow the object’s namespace, or does not exist. | Add the namespace to the identity’s allowed namespaces. See [`IdentityAllowed`](). | | `IdentitySecretFound` | Normal | `TerraformClusterIdentity` | The identity’s credentials Secret appeared. | None. | | `IdentitySecretNotFound` | Warning | `TerraformClusterIdentity` | The identity’s credentials Secret went missing. | Recreate the Secret. | | `MirrorCreated` | Normal | any provisioned kind | The credential mirror of the namespace was created on behalf of the object. | None. | | `MirrorRemoved` | Normal | any provisioned kind | The credential mirror of the namespace was deleted on behalf of the object. | None. | ## Deletion See [Deletion](). | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `DeletionStarted` | Normal | any provisioned kind | The first reconcile with a `deletionTimestamp`. | None. | | `Destroyed` | Normal | any provisioned kind | The destroy succeeded and cleanup ran. | None. | | `FinalizerRemoved` | Normal | any provisioned kind | The finalizer was removed. The object goes away. | None. | | `InfrastructureAbandoned` | Warning | any provisioned kind | A deletion held on lost or unreadable state, or whose destroy failed or cannot start, was released by `captf.io/abandon-infrastructure` naming the object’s UID. The finalizer was removed without a destroy. | Delete the orphaned cloud resources by hand. See [Manual finalizer](). | ## Pools and templates | Reason | Type | On | Fires when | Action | | --- | --- | --- | --- | --- | | `ReplicasWrittenBack` | Normal | `TerraformMachinePool` | An autoscaled pool’s observed `replicas` output was written to `MachinePool.spec.replicas`. The note reads `X → Y`. | None. | | `ReplicasManagedExternally` | Warning | `TerraformMachinePool` | The autoscaler annotations are valid, but another owner holds `replicas-managed-by`. CAPTF does not write `spec.replicas` back. | Decide which controller owns the replica count. | | `CapacityResolved` | Normal | `TerraformMachineTemplate` | The template’s `capacity` or `nodeInfo` changed from its image labels. | None. | | `ImageInspectFailed` | Warning | `TerraformMachineTemplate` | The registry could not be read for capacity. | Check registry access and credentials. | ## Runner events A Job’s runner posts its own progress as events on the object the Job is for, related to the Job itself. They fire in real time, which makes them the quickest way to see which step a long Job is in. The manager enables them with `--runner-events`, which defaults to `true`; see [Manager flags](). The runner needs `create` on `events` in its ClusterRole, which the shipped manifests grant. | Reason | Type | Fires when | | --- | --- | --- | | `RunStarted` | Normal | The runtime is ready and the first step is about to run. The note gives the operation, image reference and runtime version. | | `StepStarted` | Normal | A runtime step started. | | `StepSucceeded` | Normal | A runtime step finished. | | `StepFailed` | Warning | A runtime step failed. The note is the runner’s curated failure summary, never raw stderr. | | `PlanSummary` | Normal | The runner parsed a plan (drift, or a guarded cluster apply). The note has counts only. | | `ResourcesChanged` | Normal | An apply or destroy step finished. The note has the counts from the runtime’s summary line. | | `RunFinished` | Normal or Warning | The run ended. The note gives the result and total duration. Warning when the run did not succeed. | > [!NOTE] > > **Runner events are best effort** > > Each request has a two-second timeout. After three consecutive failures the runner stops emitting for the rest of that run. Emission never fails or slows the run. Turn `--runner-events` off on a large fleet: a single scheduled drift check alone produces several events. For where to read events during an incident, see [Observability](). # Alerts CAPTF ships eleven alerts as one `PrometheusRule`, `captf-alerts`, in the `captf` rule group. They fire on the series in [Metrics](), and each links to a section of [Observability]() that says where to look first. | Alert | Severity | Means | | --- | --- | --- | | [`CAPTFJobFailing`](<#captfjobfailing>) | warning | More than two Jobs of a kind and op failed in 30 minutes. | | [`CAPTFDestroyStuck`](<#captfdestroystuck>) | critical | Destroy Jobs for a kind have not succeeded for 30 minutes. | | [`CAPTFClusterDrift`](<#captfclusterdrift>) | warning | A `TerraformCluster` has stayed drifted for an hour. | | [`CAPTFStateUnreadable`](<#captfstateunreadable>) | critical | A state Secret turned unreadable; nothing applies. | | [`CAPTFForceUnlocks`](<#captfforceunlocks>) | warning | A Job force-unlocked a state lock. | | [`CAPTFReconcileErrors`](<#captfreconcileerrors>) | warning | A controller keeps returning reconcile errors. | | [`CAPTFJobSlow`](<#captfjobslow>) | info | The p90 Job duration is above 30 minutes. | | [`CAPTFJobQueueSlow`](<#captfjobqueueslow>) | warning | The p90 wait before a Job starts is above 5 minutes. | | [`CAPTFStateNearSecretLimit`](<#captfstatenearsecretlimit>) | warning | An object’s state is above 900 KiB of the 1 MiB limit. | | [`CAPTFInputsNearLimit`](<#captfinputsnearlimit>) | warning | An object’s rendered inputs are near the size limit. | | [`CAPTFNoRecentSuccess`](<#captfnorecentsuccess>) | warning | No drift check or health refresh succeeded in six hours. | ## Install the rules `config/prometheus` in the provider repository is an opt-in kustomize component: a metrics `Service`, a `ServiceMonitor`, the `PrometheusRule` and the RBAC Prometheus needs to read the endpoint. It is not part of `infrastructure-components.yaml`, so `clusterctl init` does not install it and `clusterctl upgrade` does not update it. Build it from a checkout of the tag you run and apply it again after each upgrade to pick up rule changes. [Enabling the Prometheus component]() has the steps. The Prometheus Operator must select the rule. If your `Prometheus` object filters `PrometheusRule` objects by label, add that label to `captf-alerts` in an overlay. Promtool unit tests for the rules are in `config/prometheus/tests/rules_test.yaml`; `make promtool-check` and `make promtool-test` run them. ## Read the alerts Two properties apply to all eleven alerts: - The failure counters count transitions, not reconciles. They go up once when an object or Job newly enters the bad state, so `increase(...) > 0` means something newly broke, not that it is still broken. A counter alert resolves after its window even if the cause remains. - Only `CAPTFClusterDrift`, `CAPTFStateNearSecretLimit`, `CAPTFInputsNearLimit` and `CAPTFNoRecentSuccess` carry `namespace` and `name` labels. The others aggregate by `kind` with `op` or `reason` (`CAPTFReconcileErrors` by `controller`), so you find the object through its [conditions](), as each runbook says. Blocked and plan-changed Jobs, which wait for an approval, are not failures and no alert counts them. See [Approvals](). ## Tune thresholds Change a rule in an overlay instead of editing the shipped file, so an upgrade does not discard the change. A kustomize patch on `captf-alerts` replaces a rule’s `expr` or `for`. Rules are a list, so the patch targets the rule by index; confirm the index against the built output first: kustomization.yaml ```yaml patches: - target: kind: PrometheusRule name: captf-alerts patch: |- - op: replace path: /spec/groups/0/rules/6/expr # (1)! value: histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1h]))) > 3600 ``` 1. Rule 6 is `CAPTFJobSlow`: rules are numbered from 0 in the order of the table above. This example raises its threshold from 30 to 60 minutes. Two thresholds mirror real limits, so raise them with care. `CAPTFStateNearSecretLimit` warns at 900 KiB because one Secret holds at most 1 MiB, and `CAPTFInputsNearLimit` warns at 900000 bytes because no Job starts above 1000000. To silence an alert for one object, use an Alertmanager silence on its labels instead of loosening the rule for all. ## CAPTFJobFailing **Severity:** warning. **For:** none, so it fires on the first evaluation that matches. ```promql sum by (kind, op) (increase(captf_jobs_total{result=~"failed|deadline"}[30m])) > 2 ``` More than two Jobs of one kind and op failed or hit their deadline within 30 minutes. Blocked, plan-changed and interrupted Jobs do not count. Likely causes: - A module error that every object of the kind hits, such as a bad input, an API quota or an expired credential. - A cloud outage or throttling. - A deadline set too short for the module. Find the objects with `ApplyJobSucceeded=False` or `DriftJobSucceeded=False` and read `status.lastRun` and the Job logs. Runbook: [`CAPTFJobFailing`](), then [Failing Jobs](). ## CAPTFDestroyStuck **Severity:** critical. **For:** 30m. ```promql sum by (kind) (increase(captf_jobs_total{op="destroy",result!="succeeded"}[30m])) > 0 ``` Destroy Jobs for a kind have failed, hit their deadline or been interrupted, with none succeeding, for 30 minutes. The objects keep their finalizer and their state, so nothing is orphaned yet, but the deletion does not finish. Likely causes: - Cloud resources that cannot be deleted because something still depends on them, such as a load balancer or a network interface. - Credentials that expired or were removed before the destroy ran. - An identity that no longer allows the namespace. Runbook: [`CAPTFDestroyStuck`](), then [Stuck Destroy](). ## CAPTFClusterDrift **Severity:** warning. **For:** 1h. ```promql captf_drift_detected{kind="TerraformCluster"} == 1 ``` A `TerraformCluster`’s last drift check found changes, and it has stayed that way for an hour. The alert carries the cluster’s `namespace` and `name`. Likely causes: - `drift.action: Report`: CAPTF reports and waits for you to decide. - `drift.action: Remediate`: the remediation apply keeps failing. Read `ApplyJobSucceeded`. - Someone changed the infrastructure outside CAPTF. Runbook: [`CAPTFClusterDrift`](), then [Drift](). Machines and pools can drift too, but no shipped alert covers them; query `captf_drift_detected` for those kinds. ## CAPTFStateUnreadable **Severity:** critical. **For:** none. ```promql sum by (kind, reason) (increase(captf_state_read_errors_total[15m])) > 0 ``` A state Secret turned unreadable in the last 15 minutes. The `reason` label says how: `inconsistent`, `encrypted`, `corrupt`, `lost` or `locked`. Nothing applies to the object until its state reads again. Likely causes: - `lost`: the state Secret was deleted, or a management-cluster restore left an object without its state. - `corrupt` or `inconsistent`: a partial write or an edit by hand. - `encrypted`: the state is encrypted and CAPTF has no key for it. - `locked`: a holder other than this object’s runner holds the lock. Find the object with `StateReadable=False` and read its reason and message. Runbook: [`CAPTFStateUnreadable`](), then [Unreadable State](). ## CAPTFForceUnlocks **Severity:** warning. **For:** none. ```promql sum by (kind) (increase(captf_lock_force_unlocks_total[1h])) > 0 ``` A Job force-unlocked a state lock whose holder pod no longer existed. The unlock is safe by design, since the holder is gone, but the previous Job died without releasing the lock. Likely causes: - The runner pod was evicted, OOM-killed or lost with its node. - A node drain or spot preemption stopped a Job mid-apply. Look for `ForceUnlocked` and `JobInterrupted` events on the object, and find why the previous Job died. Runbook: [`CAPTFForceUnlocks`](), then [Stale State Lock](). ## CAPTFReconcileErrors **Severity:** warning. **For:** 10m. ```promql sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m])) > 0.1 ``` One of the CAPTF controllers returned more than 0.1 reconcile errors per second, sustained for 10 minutes. This is a controller-runtime metric, not a `captf_*` series; the `controller` label is the controller’s name, such as `terraformcluster`. Likely causes: - The API server is unavailable or throttling the manager. - A webhook or a CRD the controller depends on is missing. - A bug that fails every reconcile of one kind. Read the manager logs for the failing kind. Runbook: [`CAPTFReconcileErrors`](), then [Reconcile Errors](). ## CAPTFJobSlow **Severity:** info. **For:** 30m. ```promql histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1h]))) > 1800 ``` The 90th percentile of Job duration for a kind and op has been above 30 minutes, measured over the last hour, for 30 minutes. Compare it with the Jobs’ `activeDeadlineSeconds`: a Job near its deadline is about to fail. Likely causes: - A module that creates slow resources, such as a managed Kubernetes control plane or a database. - A slow step: find it with `captf_job_step_duration_seconds`. - Cloud API throttling. Runbook: [`CAPTFJobSlow`](), then [Slow Jobs](). ## CAPTFJobQueueSlow **Severity:** warning. **For:** 10m. ```promql histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_queue_seconds_bucket[10m]))) > 300 ``` The 90th percentile of the time from a Job’s creation to its source container starting has been above five minutes, for 10 minutes. That time is scheduling, image pulls and the runner’s init copy; the module has not started yet. Likely causes: - No node capacity or a namespace quota blocks the runner pods. - A slow or failing pull of the module image. - Node autoscaling that is slow to add nodes. Look for `Pending` runner pods and their events. Runbook: [`CAPTFJobQueueSlow`](), then [Slow Jobs](). ## CAPTFStateNearSecretLimit **Severity:** warning. **For:** 15m. ```promql captf_state_bytes > 900 * 1024 ``` An object’s compressed state is above 900 KiB for 15 minutes. The Kubernetes state backend keeps one Secret per state, and a Secret holds at most 1 MiB, so the next apply that crosses the limit fails to save its state. The alert carries the object’s `namespace` and `name`. Likely causes: - A module that manages too many resources in one state. Check the count with `captf_state_resources`. - Large attributes stored in state. Split the module across kinds or reduce what it stores. Runbook: [`CAPTFStateNearSecretLimit`](), then [Size Limits](). ## CAPTFInputsNearLimit **Severity:** warning. **For:** 15m. ```promql captf_inputs_bytes > 900000 ``` The rendered `main.tf.json` and `terraform.tfvars.json` for an object are above 900000 bytes for 15 minutes. Above 1000000 bytes no Job starts, and `ApplyJobSucceeded` reports `InputsTooLarge`. The alert carries the object’s `namespace` and `name`. Likely causes: - A very large `spec` field, such as an inline list or a long user-data string. - Many machines or node groups in one object. Runbook: [`CAPTFInputsNearLimit`](), then [Size Limits](). See also [Inputs](). ## CAPTFNoRecentSuccess **Severity:** warning. **For:** 30m. ```promql time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 6 * 3600 ``` An object’s scheduled drift check or health refresh has not succeeded in six hours, so drift and health go unobserved. The `op` label says which. The series exists only while the op is scheduled: not while the object is deleting or paused, has drift checks off, or (for `refresh`) samples no health. The rule assumes intervals well under six hours. If you set a drift or refresh interval near or above six hours, raise the threshold. Likely causes: - The scheduled Jobs fail. Read `DriftJobSucceeded` and `status.lastRun`. - The Jobs never start: see `CAPTFJobQueueSlow` and the run lease events. Runbook: [`CAPTFNoRecentSuccess`](), then [Failing Jobs](). # Metrics The manager exposes `captf_*` series that describe Jobs, reconcile decisions, state, drift and health. The [shipped alerts]() are written on them, and the examples below build on the same series. Jump to: [Jobs](<#jobs>), [reconcile decisions](<#reconcile-decisions>), [state and inputs](<#state-and-inputs>), [drift and health](<#drift-and-health>), [locks, leases and approvals](<#locks-leases-and-approvals>), [identity and images](<#identity-and-images>), [build info](<#build-info>), [controller-runtime series](<#controller-runtime-and-other-series>). ## Where they are served The metrics are on the manager’s diagnostics endpoint, HTTPS on `:8443` (`--diagnostics-address`), at `/metrics`. Requests are authenticated with a `TokenReview` and authorized with a `SubjectAccessReview` for `get` on the non-resource URL `/metrics`; a request without a token gets `401`. `--insecure-diagnostics` serves plain HTTP with no checks and is for local development only. See [Observability]() and [Manager flags](). The opt-in `config/prometheus` component ships a `Service` (`captf-controller-manager-metrics`, port `8443`, named `metrics`), a `ServiceMonitor` that scrapes it over HTTPS with the Prometheus ServiceAccount’s token, and the RBAC that token needs. It is not part of the release manifest; see [Enabling the Prometheus component](). ## Conventions - Metrics are declared with `k8s.io/component-base/metrics` at stability level `ALPHA`, so names and labels can change between releases. On the wire, each `HELP` string starts with `[ALPHA]`. - Durations are in seconds, sizes in bytes and timestamps in Unix seconds. - Labels are bounded enums: `kind`, `op`, `result`, `reason`, `step`, `action` and `error_kind`. Only the eight per-object gauges carry `namespace` and `name`, and CAPTF removes a gauge’s series when its object is deleted, so they do not accumulate. - `kind` is the CAPTF kind: `TerraformCluster`, `TerraformMachine` or `TerraformMachinePool`. `op` is the operation: `apply`, `destroy`, `refresh`, `drift`, `plan` or `restore`. - Failure counters count transitions, not reconciles. They go up once when an object or Job newly enters a bad state, so use `increase()` to ask whether something newly broke. ## Jobs Every runner Job is observed when it completes. A Job is one `op` for one object; see [Jobs](). | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_jobs_total` | counter | `kind`, `op`, `result` | Jobs | Jobs completed, by result. | | `captf_job_duration_seconds` | histogram | `kind`, `op`, `result` | seconds | Wall time of a completed Job, start to finish. | | `captf_job_step_duration_seconds` | histogram | `kind`, `op`, `step` | seconds | Wall time of each runner step of a completed Job. | | `captf_job_queue_seconds` | histogram | `kind`, `op` | seconds | Time from a Job’s creation to its source container’s start. | | `captf_job_errors_total` | counter | `kind`, `op`, `error_kind`, `step` | Jobs | Jobs that did not succeed, by error kind and failing step. | | `captf_jobs_active` | gauge | `kind`, `op` | Jobs | Jobs running now, counted from the Job cache at scrape time. | | `captf_job_attempts` | histogram | `kind`, `op` | attempts | Retry number of a Job that succeeded. | | `captf_resources_changed_total` | counter | `kind`, `op`, `action` | resources | Resources an apply or destroy Job changed. | Label values: | Label | Values | | --- | --- | | `result` | `succeeded`, `failed`, `deadline`, `interrupted` (stopped from outside: a drain, eviction or deletion), `blocked` (a guarded apply stopped before a plan that deletes or replaces resources, awaiting approval), `plan_changed` (an approved apply planned other changes and stopped). | | `step` | `init`, `force-unlock`, `validate`, `plan`, `show-json`, `apply`, `apply-refresh-only`, `destroy`, `state-push`, `state-list`, `prepare`; anything else is `other`. On `captf_job_errors_total`, `none` when no step failed. | | `error_kind` | `step`, `image-layout`, `interrupted`, `blocked`, `plan-changed`; `deadline` or `unknown` when the Job left no result. | | `action` | `add`, `change`, `destroy`, `import`. | Notes on how each series is built: - `op="plan"` is a plan Job under `applyPolicy: Manual`; it applies nothing. - `blocked` and `plan_changed` are not failures. Each is reported by its own condition and event, and `CAPTFJobFailing` does not count them. - `captf_job_duration_seconds` buckets run from 15 seconds to 2 hours, and `captf_job_step_duration_seconds` from 1 second to 2 hours. - `captf_job_queue_seconds` has buckets from 5 seconds to 30 minutes, with 300 seconds, the `CAPTFJobQueueSlow` threshold, as a boundary. CAPTF does not record it when the pod reports no start. - `captf_job_attempts` is 1 plus the failed Jobs of the op since it last succeeded. Interrupted Jobs do not count. - `captf_resources_changed_total` comes from the runtime’s final summary line. ```promql # p95 apply duration by kind over the last day, for Jobs that ran their course histogram_quantile(0.95, sum by (le, kind) ( rate(captf_job_duration_seconds_bucket{op="apply",result=~"succeeded|failed"}[1d]))) # Failures by error kind and failing step, last hour sum by (kind, op, error_kind, step) (increase(captf_job_errors_total[1h])) > 0 ``` ```promql # Slowest runner steps (p90) topk(5, histogram_quantile(0.9, sum by (le, kind, op, step) ( rate(captf_job_step_duration_seconds_bucket[6h])))) # Resources destroyed by applies in the last 24 hours sum by (kind) (increase(captf_resources_changed_total{op="apply",action="destroy"}[24h])) ``` ## Reconcile decisions | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_reconcile_op_decisions_total` | counter | `kind`, `op`, `reason` | decisions | What a reconcile decided to run, and why. | `op` is `none` when the reconcile decided to run nothing. `reason` says why: for example `NoState`, `InputsChanged`, `DriftDue`, `DriftRemediation` or `Deleting` start a Job, while `UpToDate`, `JobActive`, `LastApplyFailedBackoff` or `DeletingBackoff` start nothing. ```promql # Why applies started in the last day sum by (kind, reason) (increase(captf_reconcile_op_decisions_total{op="apply"}[1d])) ``` ## State and inputs The per-object gauges hold the last value the manager read. See [State](). | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_state_read_errors_total` | counter | `kind`, `reason` | reads | State reads that turned unreadable. | | `captf_state_resources` | gauge | `kind`, `namespace`, `name` | resources | Managed resources, not data sources, in the object’s state. | | `captf_state_bytes` | gauge | `kind`, `namespace`, `name` | bytes | Compressed state size summed over its Secrets. | | `captf_inputs_bytes` | gauge | `kind`, `namespace`, `name` | bytes | Size of the rendered `main.tf.json` and `terraform.tfvars.json`. | | `captf_state_backups_total` | counter | `kind`, `result` | backups | State backups, by result. | | `captf_state_restores_total` | counter | `kind`, `result` | restores | State restores requested with `captf.io/restore-state`. | | `captf_outputs_invalid_total` | counter | `kind`, `reason` | outputs | Outputs that turned invalid against the module contract. | | `captf_inputs_hash_changes_total` | counter | `kind` | applies | Applies started because the inputs of a mutable kind changed. | Label values: | Label | Values | | --- | --- | | `state_read_errors_total{reason}` | `inconsistent`, `encrypted`, `corrupt` (an unsupported state version counts as corrupt), `lost` (a provisioned object’s state is gone), `locked` (held by something other than this object’s runner). | | `state_backups_total{result}` | `taken` (a new serial copied into `captf-state-backup-*` Secrets), `pruned` (a backup beyond `--state-backups` deleted), `skipped` (a new serial not backed up: encrypted, unreadable or oversized state, or a failed copy). | | `state_restores_total{result}` | `succeeded`, `failed` (a restore Job finished), `not_found` (the annotation names no backup). | The Kubernetes backend holds at most 1 MiB per Secret, and no Job starts when the inputs exceed 1000000 bytes. [`CAPTFStateNearSecretLimit`]() and [`CAPTFInputsNearLimit`]() warn at 900 KiB and 900000 bytes. ```promql # The ten largest states, as a share of one Secret's 1 MiB topk(10, captf_state_bytes / (1024 * 1024)) # Objects with unreadable state in the last 15 minutes, by reason sum by (kind, reason) (increase(captf_state_read_errors_total[15m])) > 0 ``` ## Drift and health See [Drift and health](). The gauges use the condition value: `1` for True, `0` for False and `-1` for Unknown. | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_ready` | gauge | `kind`, `namespace`, `name` | 1, 0 or -1 | The `Ready` condition. | | `captf_infrastructure_healthy` | gauge | `kind`, `namespace`, `name` | 1, 0 or -1 | The `InfrastructureHealthy` condition. | | `captf_drift_detected` | gauge | `kind`, `namespace`, `name` | 1 or 0 | 1 while `DriftDetected` is True, else 0. | | `captf_drift_resources_total` | counter | `kind`, `action` | resources | Resources a drift Job that found drift would `add`, `change` or `destroy`. | | `captf_last_success_timestamp_seconds` | gauge | `kind`, `namespace`, `name`, `op` | Unix seconds | When the newest successful Job of an op finished. | | `captf_unhealthy_samples` | gauge | `namespace`, `name` | samples | A `TerraformMachine`’s consecutive unhealthy health samples (`status.unhealthySamples`). | | `captf_remediation_requests_total` | counter | `action` | requests | `cluster.x-k8s.io/remediate-machine` annotations set on (`requested`) or removed from (`withdrawn`) a `Machine`. | `captf_last_success_timestamp_seconds` exports `drift` and `refresh` only while that op is scheduled: not while the object is deleting or paused, and only with a drift interval or health checks set. A series that goes missing is therefore not a failure by itself. ```promql # Objects that are not Ready captf_ready == 0 # Objects whose drift check or health refresh is overdue, in hours (time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"}) / 3600 > 6 ``` ## Locks, leases and approvals | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_lock_force_unlocks_total` | counter | `kind` | unlocks | Stale state locks force-unlocked. | | `captf_lease_waits_total` | counter | `kind`, `reason` | waits | Operations that started waiting for a run lease, once per wait. | | `captf_plan_approvals_total` | counter | `kind`, `result` | approvals | Plans approved with `captf.io/approve-plan` under `applyPolicy: Manual`. | | `captf_destructive_plan_approvals_consumed_total` | counter | `kind` | approvals | Destructive-plan approvals removed after the approved apply succeeded. | Label values: | Label | Values | | --- | --- | | `lease_waits_total{reason}` | `run_lease` (another live Job of the object holds it), `cluster_operation` (a machine’s apply or destroy waits for its `TerraformCluster`’s), `machine_operations` (a cluster’s apply or destroy waits for its machines’). | | `plan_approvals_total{result}` | `approved` (the apply of the approved plan succeeded and the annotation was removed), `changed` (the approved apply planned other changes and stopped; the new plan waits for approval). | See [Leases]() and [Approvals](). ```promql # Where operations wait, by reason sum by (kind, reason) (increase(captf_lease_waits_total[1h])) ``` ## Identity and images | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_identity_denied_total` | counter | `reason` | refusals | Identity refusals: `notfound` or `namespace`. | | `captf_image_inspect_errors_total` | counter | `reason` | errors | Registry or image-label failures while resolving template capacity. | ```promql # Identity refusals in the last day sum by (reason) (increase(captf_identity_denied_total[1d])) ``` ## Build info | Metric | Type | Labels | Unit | Measures | | --- | --- | --- | --- | --- | | `captf_build_info` | gauge | `version`, `commit`, `contract` | constant 1 | The manager’s version, commit and module contract. | ```promql captf_build_info ``` ## Controller-runtime and other series The same endpoint serves series that CAPTF does not define: | Family | What it covers | | --- | --- | | `controller_runtime_*` | Reconcile counts, errors, time and active workers per controller. `controller_runtime_reconcile_errors_total` backs [`CAPTFReconcileErrors`](). | | `workqueue_*` | Queue depth, latency and retries per controller. | | `rest_client_*` | Client requests to the API server, by code and verb. | | `go_*`, `process_*` | Go runtime and process statistics. | | `kubernetes_feature_enabled` | Feature gates, from component-base. | CAPTF merges its registry with the controller-runtime defaults and the component-base registry, and drops component-base’s own `go_*` and `process_*` families because controller-runtime already serves them. ```promql # Reconcile errors per second by controller sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m])) ``` # Manager Flags The CAPTF manager is the `manager` binary in the `manager` container of the `captf-controller-manager` Deployment. It takes every setting as a command-line flag; it has no configuration file, and it reads only two settings from the environment (see [Environment variables](<#environment-variables>)). The flags follow the conventions of `kube-controller-manager` and the CAPI core providers. A flag accepts `--name=value` or `--name value`, and underscores in place of dashes (`--leader_elect` is `--leader-elect`). Durations are Go duration strings such as `30m`, `90s` or `1h30m`. For what each group of flags changes in practice, see [Configuration](). This page is the complete list. ## Set a flag The shipped Deployment sets its arguments in `config/manager/manager.yaml` in the provider repository. `clusterctl init` installs that manifest, and there are no clusterctl variables for the manager’s arguments, so you change a flag by editing the `args` of the `manager` container on the live Deployment, or in a kustomize overlay over the released manifest. Deployment (abridged) ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: captf-controller-manager namespace: captf-system spec: template: spec: containers: - name: manager args: - --leader-elect - --diagnostics-address=:8443 - --insecure-diagnostics=false - --webhook-port=9443 - --terraformcluster-concurrency=20 - --drift-default-interval=1h ``` The first four arguments are the shipped ones; the highlighted two are additions. The manager logs every parsed flag once at start, one `FLAG: --name="value"` line each, and refuses to start on a value it cannot run with: a concurrency below 1, a `--sync-period` or `--drift-default-interval` that is not positive, a negative `--state-backups`, or no valid runner image. > [!WARNING] > > **A strategic-merge patch that adds one flag drops the others** > > `args` is a plain list, so a strategic-merge patch replaces the whole list. Add a flag with a JSON patch, or patch the full list. Either way, a hand-patched flag does not survive `clusterctl upgrade apply` unless you reapply it. [Changing a flag]() has the commands. ## The flags you are most likely to change | Flag | Default | Change it to | | --- | --- | --- | | `--leader-elect` | `false` (set by the shipped manifest) | Keep it on whenever more than one replica can run. | | `--runner-image` | the manager’s own image | Pin or mirror the image that supplies the runner binary to every Job. | | `--watch-filter` | empty | Share one management cluster between provider instances. | | `--namespace` | empty (all namespaces) | Confine the manager to one namespace. | | `--drift-default-interval` | `30m` | Move the fleet-wide drift cadence. | | `--state-backups` | `5` | Keep fewer or more state backups, or none. | | `--runner-events` | `true` | Turn off runner progress events on a busy cluster. | | `--terraformcluster-concurrency` and its three siblings | `10` | Raise a kind’s parallelism when its objects queue. | | `--cluster-operation-gate` | `true` | Allow a cluster’s operation and its machines’ to overlap. | | `--tls-min-version` | `VersionTLS12` | Raise the TLS floor to `VersionTLS13`. | | `-v, --v` | `2` | Raise verbosity while diagnosing, then lower it again. | ## CAPTF These flags set how the manager runs Jobs and what it keeps around them. | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--cluster-operation-gate` | `bool` | `true` | Keeps a `TerraformCluster`’s apply, destroy or restore from running at the same time as its machines’ operations, through a per-Cluster write Lease. The per-object run Lease that stops two Jobs for one object is always on. Turn the gate off only if you accept those overlapping. | | `--drift-default-interval` | `duration` | `30m` | The drift check interval an object falls back to when neither it nor its cluster’s defaults set one. Must be positive. A machine pool whose own or inherited interval is `0` also uses it, since pool drift cannot be disabled. | | `--runner-events` | `bool` | `true` | Has each Job’s runner post progress events (`RunStarted`, `Step*`, `PlanSummary`, `ResourcesChanged`, `RunFinished`) on the owning `Terraform*` object. Emission is best effort and never fails a run. Needs `create` on `events` in the runner ClusterRole. Turn it off to cut event volume. | | `--runner-image` | `string` | `$CAPTF_MANAGER_IMAGE` | The image of the init container that copies the runner binary into every Job. It must contain `/runner`, which means a CAPTF manager or runner image, never a module image. Unset, it takes the manager’s own image. The manager refuses to start when this is empty or not a valid image reference. | | `--state-backups` | `int` | `5` | How many state backups to keep per object. Each new state serial is copied into `captf-state-backup-*` Secrets, and older copies are pruned. `0` takes no new backups; existing ones stay and can still be restored. Must not be negative. | ## Controllers and concurrency These flags choose what the manager watches and how many objects it reconciles at once. | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--namespace` | `string` | empty | Restricts the manager to one namespace. Empty watches all, which is what `clusterctl init` installs. A manager watches one namespace or all of them, never a chosen set. `TerraformClusterIdentity` is cluster-scoped and is watched everywhere regardless. | | `--sync-period` | `duration` | `10m` | The minimum interval at which the informers re-enqueue every cached object, on top of event-driven reconciles. Also the interval of the orphan sweep. It reads the local cache only and does not set the drift or health cadence. Must be positive. | | `--terraformcluster-concurrency` | `int` | `10` | How many `TerraformCluster` objects reconcile at once. Must be at least 1. | | `--terraformmachine-concurrency` | `int` | `10` | How many `TerraformMachine` objects reconcile at once. Must be at least 1. | | `--terraformmachinepool-concurrency` | `int` | `10` | How many `TerraformMachinePool` objects reconcile at once. Must be at least 1. | | `--terraformmachinetemplate-concurrency` | `int` | `10` | How many `TerraformMachineTemplate` objects reconcile at once. Must be at least 1. | | `--watch-filter` | `string` | empty | Reconcile only objects labeled `cluster.x-k8s.io/watch-filter` with this value. Empty reconciles every object the manager can see. The label key is fixed. | A reconcile is a handful of API reads and a status patch, and mostly waits on a Job, so a higher concurrency rarely costs much CPU. Lower it to soften bursts against a small API server. The admission webhooks serve every namespace whatever `--namespace` and `--watch-filter` say, so an object the manager does not watch is admitted and then never reconciled. [Namespace scoping and `--watch-filter`]() explains the trade-offs. ## Leader election Leader election makes sure one manager reconciles at a time. The Lease is named `controller-leader-election-captf` and is not parameterized. | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--leader-elect` | `bool` | `false` | Turns leader election on. Enable it whenever more than one replica can run, including during a rolling update. The shipped manifest sets it. | | `--leader-elect-lease-duration` | `duration` | `15s` | How long a non-leader candidate waits before it forces leadership. | | `--leader-elect-renew-deadline` | `duration` | `10s` | How long the leader keeps retrying to renew before it gives up leadership. | | `--leader-elect-retry-period` | `duration` | `2s` | How long a candidate waits between attempts. | The defaults match `kube-controller-manager` and rarely need changing. Lower them to replace a crashed leader faster, at the cost of more Lease traffic. ## Webhooks The webhook server hosts the validating webhooks. cert-manager issues its serving certificate, and the Deployment mounts the Secret into the manager pod. | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--webhook-cert-dir` | `string` | `/tmp/k8s-webhook-server/serving-certs/` | The directory that holds the serving certificate and key. | | `--webhook-cert-name` | `string` | `tls.crt` | The certificate’s file name in that directory. | | `--webhook-key-name` | `string` | `tls.key` | The key’s file name in that directory. | | `--webhook-port` | `int` | `9443` | The port the webhook server listens on. The shipped manifest sets `9443`; the webhook Service targets it, so change both together. | The shipped manifests wire these together. You change them only for a custom certificate layout. ## Diagnostics and TLS (CAPI) These flags come from Cluster API’s shared manager options and behave as they do in the CAPI core providers. The diagnostics endpoint serves Prometheus metrics over HTTPS, authenticated and authorized against the API server. See [Observability]() for what it serves. | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--diagnostics-address` | `string` | `:8443` | The address the diagnostics endpoint binds to. With authentication on, it also serves the pprof endpoints and an endpoint that changes the log level at run time. | | `--insecure-diagnostics` | `bool` | `false` | Serves diagnostics over plain HTTP with no authentication or authorization, and without the pprof and log-level endpoints. For local development only. | | `--tls-cipher-suites` | `stringSlice` | empty | A comma-separated list of cipher suites for the webhook server and the metrics server (the latter only when diagnostics are secure). Empty uses Go’s defaults. The flag’s help text lists the preferred and the insecure names. | | `--tls-curve-preferences` | `int32Slice` | empty | A comma-separated list of numeric Go `crypto/tls` `CurveID` values to allow as key exchange mechanisms. The order is ignored. Empty uses Go’s defaults. The supported values depend on the Go version. | | `--tls-min-version` | `string` | `VersionTLS12` | The minimum TLS version of the webhook and metrics servers: `VersionTLS10`, `VersionTLS11`, `VersionTLS12` or `VersionTLS13`. | > [!CAUTION] > > **`--insecure-diagnostics` removes authentication from the metrics endpoint** > > Anything that can reach the address then reads the metrics with no token. Use it on a workstation, never on a cluster others can reach. ## Logging The logging flags are the Kubernetes component-base set. The manager runs at verbosity 2, where the usual reconcile flow is visible. Credentials, bootstrap data, tfvars content and output values are never logged, whatever the level. | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--feature-gates` | `mapStringBool` | empty | A comma-separated `Key=value` list. Only the logging gates in [Feature gates](<#feature-gates>) exist. | | `--log-flush-frequency` | `duration` | `5s` | The longest interval between log flushes. | | `--log-json-info-buffer-size` | `quantity` | `0` | Alpha. Buffers info messages in the JSON format with split streams, to increase performance. `0` disables buffering. Takes `512`, `1K`, `2Ki` and the like. Needs `LoggingAlphaOptions`. | | `--log-json-split-stream` | `bool` | `false` | Alpha. In the JSON format, writes errors to stderr and info messages to stdout, instead of one stream to stdout. Needs `LoggingAlphaOptions`. | | `--log-text-info-buffer-size` | `quantity` | `0` | Alpha. The same buffering for the text format with split streams. Needs `LoggingAlphaOptions`. | | `--log-text-split-stream` | `bool` | `false` | Alpha. In the text format, writes errors to stderr and info messages to stdout. Needs `LoggingAlphaOptions`. | | `--logging-format` | `string` | `text` | The log format: `text` or `json`. `json` is gated by `LoggingBetaOptions`, which is on by default. | | `-v, --v` | `Level` | `2` | The log verbosity. Raise it while you diagnose a problem and lower it afterward, since higher levels log more of each reconcile. | | `--vmodule` | `pattern=N,...` | empty | Per-file verbosity overrides, such as `reconcile*=4`. Works only with the text format. | To change the level of a running manager without a restart, use the authenticated diagnostics endpoint; see [Logs and verbosity](). ## Other flags | Flag | Type | Default | What it does | | --- | --- | --- | --- | | `--health-addr` | `string` | `:9440` | The address of the health endpoint that serves `/healthz` and `/readyz`. The Deployment’s probes target port `9440`, so change both together. | | `--kubeconfig` | `string` | empty | The path to a kubeconfig. You need it only out of cluster. Without it, the manager uses `$KUBECONFIG`, then its in-cluster configuration. | | `--profiler-address` | `string` | empty | A bind address for a separate pprof profiler, such as `localhost:6060`. Empty leaves it off. Unrelated to the pprof endpoints on the diagnostics address. | | `-h, --help` | `bool` | `false` | Prints the flag help, grouped by section, and exits. | | `--version` | `version` | `false` | Prints version information and exits. `--version=raw` prints the raw build stamp. `--version=vX.Y.Z` sets the reported version instead. | `manager version` prints the same line as `manager --version` and exits. ## Feature gates `--feature-gates` takes a comma-separated `Key=value` list, for example `--feature-gates=LoggingAlphaOptions=true`. CAPTF registers no feature gate of its own. The gates that exist are the component-base logging gates. | Name | Default | Maturity | | --- | --- | --- | | `AllAlpha` | `false` | Alpha | | `AllBeta` | `false` | Beta | | `ContextualLogging` | `true` | Beta | | `LoggingAlphaOptions` | `false` | Alpha | | `LoggingBetaOptions` | `true` | Beta | ## Environment variables The manager reads these from its own environment. A custom Deployment or an overlay that replaces the container’s `env` must keep them. | Variable | What it does | | --- | --- | | `CAPTF_MANAGER_IMAGE` | The default of `--runner-image`. The shipped Deployment sets it to the manager’s own image, and `clusterctl init` replaces it with the release image. | | `KUBECONFIG` | Read by the `--kubeconfig` handling when the flag is not given. Only an out-of-cluster manager needs it. | The manager also reads `POD_NAMESPACE` and `SERVICE_ACCOUNT_NAME`, which the shipped Deployment sets through the downward API. Without them the webhook cannot tell which ServiceAccount is the manager and refuses its `providerID` writes. See [Manager environment](). ## Shipped arguments `config/manager/manager.yaml` sets these arguments on the `manager` container. Every other flag keeps its default. | Flag | Value | | --- | --- | | `--leader-elect` | set | | `--diagnostics-address` | `:8443` | | `--insecure-diagnostics` | `false` | | `--webhook-port` | `9443` | > [!NOTE] > > **See also** > > - [Configuration]() — when and why to change each group of flags. > - [Observability]() — the diagnostics endpoint, metrics and verbosity. > - [Installation]() — what `clusterctl init` installs. > - [Job Environment]() — what the runner Job a manager creates looks like. # Requeue Intervals and Schedules The controller is event-driven first: a change to the object, a Job finishing or an annotation triggers a reconcile. Requeues are the fallback for time and for waits that have no event. This page lists every interval, its default, and where it comes from. All of them are compiled in unless a field or flag is named. ## Schedules (when a Job is due) | Schedule | Default | Set by | Notes | | --- | --- | --- | --- | | Drift check | 30 min | `spec.drift.intervalSeconds`; else `--drift-default-interval` | `0` disables it for a cluster or machine; a pool’s drift is never disabled | | Health check | 300 s | `remediation.healthCheckIntervalSeconds` | Only while `remediation.annotateMachine` is true; otherwise health is sampled at the drift or refresh cadence | | Membership refresh | 60 s | A pool’s `membershipRefreshIntervalSeconds` | Pools only | | Pending-health refresh | 30 s, doubling to 5 min | Compiled in | Doubles per consecutive pending reading: 30 s, 1 min, 2 min, 4 min, 5 min | | Converging-membership refresh | 30 s | Compiled in | Fixed, not doubled: members join on the provider’s schedule | | Refresh after apply | Once, at once | Compiled in | Machines and pools; skipped when the apply’s own reading is definite | Each schedule’s deadline adds a **jitter** of up to a tenth of its interval. It is deterministic: derived from a hash of the object’s UID, so the same object always gets the same offset, and objects created together (a `MachineDeployment`’s machines, or everything after a `clusterctl move`) do not check in lockstep. An object with no UID gets none. ## Waits (when the controller looks again) | Interval | Value | Used while | | --- | --- | --- | | `GateRequeue` | 30 s | An owner reference, a dependency, credentials, deletion order or a lease is not ready | | `StateRequeue` | 1 min | The state is lost or unreadable, including a held deletion | | `LagRequeue` | 5 s | The Job cache trails the API server, or a live Job holds the run lease at cleanup, or an annotation write conflicted | | `ActiveJobRequeue` | 1 min | A Job runs; the Job watch normally wakes the reconcile first | | `RetryBase` and `RetryMax` | 1 min, 10 min | The backoff after a failed Job; see [Retries]() | | `RetryMax` | 10 min | An approval wait, `JobPolicyInvalid`, `InputsTooLarge`, or missing durable inputs: nothing to do until something changes, and the change wakes the reconcile | | Stuck-Job age | 1 min | A Job younger than this is not checked for a missing per-run Secret | | Lease grace | 1 min | A lease whose holder Job does not exist stays held this long | | Identity re-read | 5 min | A `TerraformClusterIdentity` re-reads its credentials Secret | ## The resync `--sync-period` (default 10 minutes) is the minimum interval at which the informers re-enqueue every cached object. It reads the local cache, not the API server, and sets no schedule of drift or health. It recovers a requeue that was missed, at the cost of more reconciles. The orphan sweep ticks at the same interval, with 10% jitter, on the leader. See [Configuration](). ## How they combine A reconcile with nothing to run requeues at the **soonest** of the due times: the next refresh, health, membership or drift deadline, whichever applies. A reconcile that waits on something requeues at that wait’s interval. A failed op’s backoff replaces its Job with a requeue for the remaining delay, and pauses nothing else: another op still runs on its own schedule. A waiting input change is the exception: it pauses the refresh and drift schedule, because they would render the unapplied inputs and report the change as drift. | Object state | Typical next reconcile | | --- | --- | | Idle, healthy | The sooner of the drift and health deadlines, and any event | | A Job running | The Job’s finish event; a minute at the latest | | Backing off after a failure | The end of the backoff | | Waiting for a lease | 30 s | | Waiting for an approval | The annotation, or 10 min | | Held on the state | 1 min | ## See also - [Choosing the operation](). - [Drift and Health](). - [Configuring drift](). # Job Environment Every Terraform or OpenTofu operation runs as one Kubernetes Job. The manager builds it, the runner binary inside it drives the runtime, and your module image supplies the module. This page is an anatomy of that Job: the containers, the environment each one gets, the volumes, the fixed fields, the default resources and security contexts, the runner’s command line, and what the runner does to the environment before it runs your module. Use it to answer what a module sees at run time. For the fields you can change, see [Tuning Jobs]() and [`spec.jobs`](). For the order of the steps the runner runs, see [Runtime Environment](). For the Job’s place in the lifecycle, see [Jobs, Retries and Concurrency](). ## Anatomy of a Job A Job has one init container and one main container, sharing five volumes (six for a plan or an apply): - The **init container**, named `runner`, runs the manager’s runner image (`--runner-image`). It copies one static binary, the runner, into a shared volume and exits. - The **main container**, named `source`, runs your module image (`spec.source.image`) with the copied runner as its command. The runner checks the image layout, builds the working directory and runs the runtime as a child process. The image’s own `ENTRYPOINT` and `CMD` never run. The example below is an apply Job for a `TerraformCluster`, abridged to the fields this page describes. Job pod spec (abridged) ```yaml spec: backoffLimit: 0 # the controller owns retries activeDeadlineSeconds: 3600 # spec.jobs.activeDeadlineSeconds template: spec: restartPolicy: Never terminationGracePeriodSeconds: 600 # SIGTERM, then SIGKILL after 10 min serviceAccountName: captf-runner # spec.jobs.serviceAccountName securityContext: seccompProfile: {type: RuntimeDefault} fsGroup: 65532 # lets a non-root image user read 0440 files initContainers: - name: runner # the manager's runner image image: ghcr.io/captf-io/cluster-api-provider-terraform:vX.Y.Z command: ["/runner", "copy", "/captf/bin/runner"] volumeMounts: - {name: runner, mountPath: /captf/bin} # fixed resources and a locked-down security context: see below containers: - name: source # your module image image: registry.example.com/modules/cluster:v1 command: ["/captf/bin/runner", "run"] args: - --op=apply - --bin=/captf/runtime - --lock-timeout=300s - --stop-timeout=570s # ... the rest of the runner flags: see "Runner command and args" env: - {name: TF_IN_AUTOMATION, value: "1"} - {name: TF_INPUT, value: "0"} - {name: HOME, value: /captf/work} - {name: TMPDIR, value: /tmp} - {name: KUBE_NAMESPACE, value: default} - {name: CHECKPOINT_DISABLE, value: "1"} # ... then the entries of spec.jobs.env that are not TF_* or KUBE_* envFrom: - secretRef: {name: captf-creds-} # the credentials mirror volumeMounts: - {name: runner, mountPath: /captf/bin, readOnly: true} - {name: work, mountPath: /captf/work} - {name: tmp, mountPath: /tmp} - {name: config, mountPath: /captf/config, readOnly: true} - {name: creds, mountPath: /var/run/captf/credentials, readOnly: true} - {name: plan-key, mountPath: /captf/plan-key, readOnly: true} securityContext: allowPrivilegeEscalation: false capabilities: {drop: [ALL]} readOnlyRootFilesystem: true resources: requests: {cpu: 250m, memory: 512Mi} limits: {memory: 2Gi} terminationMessagePolicy: ReadFile # the runner writes its result here ``` The Job also carries labels that name the owner and the operation, and, for an apply, a plan and a restore, an annotation with the inputs hash it renders. See [Job names, attempts and history]() for the naming scheme. ## Fixed Job fields | Field | Value | Meaning | | --- | --- | --- | | `backoffLimit` | `0` | The controller owns retries and the attempt number is in the Job name, so the pod never retries itself. | | `restartPolicy` | `Never` | A failed container is not restarted. | | `ttlSecondsAfterFinished` | unset | No Job expires on its own. The controller reads backoff and conditions from retained Jobs and prunes them by the history limits. | | `terminationGracePeriodSeconds` | `600` | SIGTERM lets the runner finish in-flight provider calls and write its result before SIGKILL. A shorter period would kill a run mid-call and leave resources created but never recorded. | | `activeDeadlineSeconds` | `3600` | The default. Set `spec.jobs.activeDeadlineSeconds` to change it. | | `lockTimeoutSeconds` | `300` | The default, passed to the runner as `--lock-timeout`. Set `spec.jobs.lockTimeoutSeconds` to change it. | The runner’s stop timeout is the grace period less a 30-second margin (`--stop-timeout=570s`), which leaves time to write the result. See [Deadlines]() for how the deadline and the lock timeout interact. ## Main container environment The Job sets these variables on the main container. A module sees them unchanged, except `HOME`, which the runner sets again to the same value. | Name | Value | What it does | | --- | --- | --- | | `TF_IN_AUTOMATION` | `1` | Tells the runtime it runs unattended, so it leaves out interactive follow-up hints. | | `TF_INPUT` | `0` | Disables interactive prompts. The runner also passes `-input=false`. | | `HOME` | `/captf/work` | The runner’s working directory. The image’s own `HOME` may not be writable. | | `TMPDIR` | `/tmp` | The Job’s `/tmp` volume, since the image’s root filesystem is read-only. | | `KUBE_NAMESPACE` | the object’s namespace | The namespace of the `kubernetes` state backend, so state Secrets land beside the owning object. | | `CHECKPOINT_DISABLE` | `1` | Stops Terraform from calling `checkpoint-api.hashicorp.com` on every command. A pod that holds cloud credentials otherwise makes that call, and it can stall on its timeout when a namespace drops egress silently. | The Job builds `env` from these six, then appends the entries of `spec.jobs.env` that do not begin with `TF_` or `KUBE_`; see [`spec.jobs.env` rejected names](<#specjobsenv-rejected-names>). The Job does not set `TF_DATA_DIR` or `TF_CLI_CONFIG_FILE`. The runner sets both for the runtime; see [What the runner changes](<#what-the-runner-changes>). An explicit `env` entry wins over `envFrom`. A credential key that is named like one of these six therefore never replaces the Job’s value. ## Identity credentials The identity’s credentials arrive in the main container two ways at once, both from the credentials mirror. The mirror is a Secret named `captf-creds-` in the object’s namespace, which the manager copies from the Secret that `TerraformClusterIdentity.spec.secretRef` names. - **`envFrom`.** Every key of the mirror becomes an environment variable with that name. A module’s providers read `AWS_ACCESS_KEY_ID` and the like from the environment as they do anywhere else. Kubernetes skips a key that is not a valid variable name. - **Files.** The `creds` volume mounts the mirror at `/var/run/captf/credentials`, read-only, one file per key, mode `0440`. A module reads a key file or certificate from there, for example by naming the path in a provider setting. > [!WARNING] > > **Every key in the identity Secret reaches the runtime** > > The mirror is a byte copy of the whole Secret, so every key becomes a variable and a file, whether or not the module uses it. A key that begins with `TF_` or `KUBE_` is removed from the runtime’s environment (see [What the runner changes](<#what-the-runner-changes>)) but stays available as a file. The mirror, its rotation and its revocation are covered in [Credentials](), and the identity itself in [Identities and Credentials](). ## Volumes and mounts The init container mounts only `runner`, read-write. The main container mounts the volumes below. The three Secret volumes use mode `0440`. | Mount path | Volume | Access | Contents | | --- | --- | --- | --- | | `/captf/bin` | `runner` | read-only | The runner binary, copied in by the init container. | | `/captf/work` | `work` | read-write | Scratch space: the generated root at `/captf/work/root`, the CLI configuration, plan files and `TF_DATA_DIR`. | | `/tmp` | `tmp` | read-write | General temporary storage (`TMPDIR`). | | `/captf/config` | `config` | read-only | The per-run Secret: the rendered `main.tf.json` and `terraform.tfvars.json`. For a restore Job, a projection of it plus the backup’s state chunks under `restore/`. | | `/var/run/captf/credentials` | `creds` | read-only | The identity’s credential files. | | `/captf/plan-key` | `plan-key` | read-only | The plan fingerprint key, as the single file `key`. Plan and apply Jobs only; other operations never see it. | `runner`, `work` and `tmp` are empty directories that live as long as the pod. The pod’s `fsGroup` defaults to `65532`, so a non-root image user reads the `0440` files through the group. `/captf/module`, `/captf/runtime` and the optional `/captf/providers` are not volumes: they are paths in your image. See [Image Contract]() for the layout and [Run Inputs and the Plan Key]() for the key. ## Default resources The init container only copies one static binary and never varies with the module, so its resources are fixed and not configurable. | Container | Kind | Values | | --- | --- | --- | | init (`runner`) | requests | `cpu=10m`, `memory=32Mi` | | init (`runner`) | limits | `cpu=100m`, `memory=64Mi` | | main (`source`), when `spec.jobs.resources` is unset | requests | `cpu=250m`, `memory=512Mi` | | main (`source`), when `spec.jobs.resources` is unset | limits | `memory=2Gi`, and no CPU limit | The main container has no default CPU limit because throttling a slow apply is worse than a slow apply. A pod with no requests at all runs at the BestEffort class, which Kubernetes evicts or OOM-kills first on a crowded node, and that is the wrong place for a Terraform process with several large providers. Setting `spec.jobs.resources` replaces the default as a whole; see [Resources](). ## Security contexts | Scope | Field | Value | | --- | --- | --- | | pod | `seccompProfile` | `RuntimeDefault` | | pod | `fsGroup` | `65532`, so a non-root image user reads the `0440` credential files through the group | | init (`runner`) | `allowPrivilegeEscalation` | `false` | | init (`runner`) | `capabilities.drop` | `ALL` | | init (`runner`) | `runAsNonRoot`, `runAsUser` | `true`, `65532` | | init (`runner`) | `readOnlyRootFilesystem` | `true` | | main (`source`) | `allowPrivilegeEscalation` | `false` | | main (`source`) | `capabilities.drop` | `ALL`, always, whatever `spec.jobs.securityContext` holds | | main (`source`) | `readOnlyRootFilesystem` | `true` | | main (`source`) | `runAsNonRoot`, `runAsUser` | not set | `spec.jobs.podSecurityContext` and `spec.jobs.securityContext` overlay these defaults field by field, within the limits the webhook enforces; see [Security contexts](). > [!NOTE] > > **The main container is not forced to run as non-root** > > A module image built from the `hashicorp/terraform` image runs as root, so the Job leaves `runAsNonRoot` and `runAsUser` unset. The webhook rejects an explicit root setting, but an image that runs as root by default still does. The container holds cloud credentials, so build the image to run as a non-root user. See [Pod security](). ## Runner command and args The main container’s command is `/captf/bin/runner run`. The Job passes these arguments, in this order. Every path is fixed by the [image contract](). The runner CLI’s other flags keep their defaults; see [Runner CLI](). | Flag | Value | What it does | | --- | --- | --- | | `--op` | `apply`, `destroy`, `refresh`, `drift`, `restore` or `plan` | The operation this Job runs. | | `--bin` | `/captf/runtime` | The image’s `tofu` or `terraform` binary. | | `--image` | the source image reference | Identifies the image that ran, in events and logs. | | `--module` | `/captf/module` | The image’s role module. | | `--providers` | `/captf/providers` | The image’s optional provider mirror. The runner checks whether it exists. | | `--workdir` | `/captf/work` | The parent of the generated root. | | `--config` | `/captf/config` | The per-run Secret’s mount. | | `--lock-timeout` | `300s`, or `spec.jobs.lockTimeoutSeconds` | How long the backend waits for the state lock before it fails. | | `--stop-timeout` | `570s` | How long the runner has after SIGTERM to finish in-flight work and write its result. | | `--backend-config=secret_suffix` | a per-attempt suffix | The suffix of the `kubernetes` backend’s Secret name. | | `--backend-config=namespace` | the object’s namespace | The backend’s namespace. | | `--backend-config=in_cluster_config` | `true` | Tells the backend to use the pod’s in-cluster credentials. | | `--backend-config=labels` | the state Secret labels, as an HCL object | The labels the backend puts on the state Secrets it manages. | | `--force-unlock` | a stale lock ID | Set only when the controller found a stale backend lock. The runner force-unlocks it after `init`. | | `--guard-deletes` | no value | Set for every `TerraformCluster` apply, and for a `TerraformMachinePool` apply that renders a change of the cluster’s exports. The runner stops before a plan that deletes or replaces a resource unless that is approved. | | `--inputs-hash` | the approval hash of the rendered inputs | Set with `--guard-deletes`. It is the hash an approval must name: a cluster’s inputs hash, or a pool’s inputs hash without `bootstrap_data`. | | `--allow-deletes-hash` | an approved hash | Set with `--guard-deletes` when the object carries `captf.io/approve-destructive-plan`. It allows a destructive plan for that exact input set. | | `--expect-plan` | an approved plan hash | Set with `--guard-deletes` under `applyPolicy: Manual`, from the cluster’s `captf.io/approve-plan`. The apply plans again and applies only if the plan matches. | | `--restore-chunks` | the chunk count | Restore Jobs only: how many backup state chunks the config volume projects. | | `--restore-resources` | the managed resource count | Restore Jobs only: the backup’s recorded count. The restore fails if `state list` shows none when this is not `0`. | | `--plan-key-file` | `/captf/plan-key/key` | Plan and apply Jobs only: the file that holds the key of the plan fingerprint. | | `--event-object` | `////` of the owner | Set only when the manager runs with `--runner-events`. It lets the runner report progress as events about the owning object. | | `--job-name` | the Job’s name | Set with `--event-object`: the Job the events relate to. | ## What the runner changes The runner starts from the container’s environment: the Job’s variables, the identity `envFrom`, the image’s own `ENV` and `spec.jobs.env`. Before it runs the first command, it builds a new environment for the runtime and leaves everything else out. Given a pod environment that carries every kind of credential-bearing or backend-changing variable: ```text TF_IN_AUTOMATION=1 TF_INPUT=0 KUBE_NAMESPACE=default CHECKPOINT_DISABLE=1 TF_LOG=DEBUG TF_VAR_region=us-east-1 TF_WORKSPACE=default TF_CLI_ARGS_apply=-parallelism=1 KUBE_CONFIG_PATH=/tmp/kubeconfig HOME=/root TMPDIR=/tmp AWS_ACCESS_KEY_ID=AKIA... ``` the runtime sees: ```text AWS_ACCESS_KEY_ID=AKIA... CHECKPOINT_DISABLE=1 HOME=/captf/work KUBE_NAMESPACE=default TF_DATA_DIR=/captf/work/.terraform TF_INPUT=0 TF_IN_AUTOMATION=1 TMPDIR=/tmp ``` and the runner logs the dropped names, never their values, as `Ignoring environment variables the runner owns`: `KUBE_CONFIG_PATH`, `TF_CLI_ARGS_apply`, `TF_LOG`, `TF_VAR_region` and `TF_WORKSPACE`. The rules are: - `TF_DATA_DIR` is always `/captf/work/.terraform`, whatever the pod inherited. - `HOME` is always `/captf/work`. - `TMPDIR` is kept when the pod sets it, and the Job sets it to `/tmp`. Otherwise it is `/captf/work/tmp`, which the runner creates. - `TF_CLI_CONFIG_FILE` is set only when the image ships a provider mirror; see [Provider mirror and CLI configuration](<#provider-mirror-and-cli-configuration>). - Every other `TF_*` and `KUBE_*` variable is dropped. Only `TF_IN_AUTOMATION`, `TF_INPUT` and `KUBE_NAMESPACE` stay, because the Job sets them and an explicit `env` entry beats `envFrom`, so they never carry a credential’s or an image’s value. Each drop closes a way to change a run from outside the module. `TF_WORKSPACE` moves state, `TF_CLI_ARGS*` rewrites every command, `TF_VAR_*` overrides the rendered inputs, `TF_LOG` logs provider traffic with its credentials, and `KUBE_*` redirects the `kubernetes` backend. ## Provider mirror and CLI configuration The Job cannot know whether the image ships providers, so the runner decides. When `/captf/providers` exists, it writes `/captf/work/cli.tfrc` with mode `0600`: /captf/work/cli.tfrc ```hcl provider_installation { filesystem_mirror { path = "/captf/providers" include = ["*/*/*"] } direct { exclude = ["*/*/*"] } } ``` and sets `TF_CLI_CONFIG_FILE=/captf/work/cli.tfrc`. `init` then installs every provider from the image and contacts no registry, and it fails with a clear error for a provider the mirror lacks instead of downloading it. The three-segment `*/*/*` pattern matches every provider on every registry host. The two-segment form matches only the default one. Without `/captf/providers`, the runner writes no file and sets no variable, and `init` falls back to the runtime’s default direct installation, which needs registry egress. The runner never sets `TF_PLUGIN_CACHE_DIR`, since a plugin cache must not coincide with a filesystem mirror. The mirror layout is in [Image Contract](). ## `spec.jobs.env` rejected names `spec.jobs.env` adds variables to the main container, up to 64 entries. The Job reserves every name that begins with `TF_` or `KUBE_`: the runner and the Job own those two prefixes. That covers the three the Job sets itself (`TF_IN_AUTOMATION`, `TF_INPUT` and `KUBE_NAMESPACE`) and every setting a person might reach for, such as `TF_LOG`, `TF_VAR_*`, `TF_WORKSPACE`, `TF_CLI_ARGS*`, `TF_CLI_CONFIG_FILE` and `KUBE_CONFIG_PATH`. > [!WARNING] > > **A reserved name is dropped without an error or an event** > > The API accepts an entry with a `TF_` or `KUBE_` name and the Job builder leaves it out. Nothing reports it, and the entry stays in the object’s spec. If a variable has no effect, check its prefix first. Other names are applied, including `HOME`, `TMPDIR` and `CHECKPOINT_DISABLE`, which the Job also sets. The runner forces `HOME` again, so setting it has no effect on the runtime. There is no way to turn on provider debug logging from `spec.jobs.env`. To debug a provider, run the image’s `/captf/runtime` by hand with `TF_LOG` set; [Runtime Environment]() describes the recipe. > [!NOTE] > > **See also** > > - [Tuning Jobs]() — the `spec.jobs` fields that change this Job. > - [Inside the Job]() — the Secrets the pod mounts and the permissions it holds. > - [Runtime Environment]() — the commands the runner runs and how a failure is reported. > - [Runner CLI]() — every runner flag and default. > - [Image Contract]() — the paths your image must provide. # Runner CLI `runner` is the program inside every CAPTF Job. The manager does not run Terraform or OpenTofu itself: it creates a Job, and the runner in that Job prepares the working directory, runs the module’s runtime step by step, and writes a small JSON result that the manager reads back when the Job ends. The binary is part of the manager image (`/runner`). You never start the runner by hand in normal use. The manager builds the Job so that: 1. An init container, from the manager image, runs `/runner copy /captf/bin/runner`. This puts a copy of the binary on a shared volume, so the module’s own image does not need to contain it. 2. The main container, from the module’s source image, runs `/captf/bin/runner run` with the flags for one operation. Run it yourself only to reproduce a failed Job in a container you control, or to read its flags while debugging a Job spec. See the [image contract]() for the paths the runner expects and [Choosing the Operation]() for when the manager picks each operation. ## Synopsis ```text runner [global flags] copy runner [global flags] run --op= [flags] runner [global flags] version runner --version[=raw] ``` A bare `runner` prints its usage to standard error and exits with code `2`. ## Global flags Every subcommand accepts these. They are the logging and version flags that every Kubernetes component registers, so they behave as they do in `kube-apiserver` or `kubectl`. | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--feature-gates` | `key=value,...` | | Enable alpha or beta logging features. The gates are `AllAlpha`, `AllBeta`, `ContextualLogging`, `LoggingAlphaOptions` and `LoggingBetaOptions`. | | `--log-flush-frequency` | `duration` | `5s` | Longest interval between log flushes. | | `--log-json-info-buffer-size` | `quantity` | `0` | Alpha. Buffer info messages in JSON format with split streams. `0` disables buffering. Needs the `LoggingAlphaOptions` gate. | | `--log-json-split-stream` | `bool` | `false` | Alpha. In JSON format, write errors to standard error and info messages to standard output. Needs the `LoggingAlphaOptions` gate. | | `--log-text-info-buffer-size` | `quantity` | `0` | Alpha. The same buffering for text format with split streams. | | `--log-text-split-stream` | `bool` | `false` | Alpha. The same stream split for text format. | | `--logging-format` | `string` | `text` | Log format: `text` or `json`. The `json` format needs the `LoggingBetaOptions` gate, which is on by default. | | `-v, --v` | `Level` | `0` | Log verbosity. | | `--version` | `version` | `false` | Print version information and exit. `--version=raw` prints the full build information. `--version=vX.Y.Z` sets the reported version. | | `--vmodule` | `pattern=N,...` | | Per-file verbosity. Works with the text format only. | Invalid logging flags fail with exit code `2`. For `run`, they also write a result (see [Result error kinds](<#result-error-kinds>)). ## copy ```text runner copy ``` Copies the running binary to `` and sets it executable (mode `0755`). `` is the one required argument and takes no flags. The Job’s init container uses it to place the runner at `/captf/bin/runner`. ```sh /runner copy /captf/bin/runner ``` It exits `0` on success, `1` when it cannot read or write a file, and `2` when it does not get exactly one argument. ## run ```text runner run --op= [flags] ``` Runs one operation in the Job’s main container. It takes no arguments, only flags. `--op` is required, and takes one of these: | Operation | What the runner does | | --- | --- | | `apply` | Runs `init`, `validate`, then `apply`. With `--guard-deletes` or `--expect-plan`, it plans first, checks the plan, and applies the saved plan only when the checks pass. | | `destroy` | Runs `init`, then `destroy`. | | `refresh` | Runs `init`, then a refresh-only apply to update state from the real resources. | | `drift` | Runs `init`, a refresh-only apply, then a plan without refresh, and reads the plan to report whether anything differs. | | `plan` | Runs `init`, `validate`, then `plan`, and summarizes the plan for review. Nothing is applied. | | `restore` | Runs `init`, then pushes a state backup assembled from `/restore` and lists the result as a check. | Every operation runs `init` first, with `-backend-config` values from `--backend-config` and the lock timeout from `--lock-timeout`. When `--force-unlock` is set, `force-unlock` runs right after `init`. The result goes to `--result` when the run ends, whether it succeeded or not. ### Paths and environment | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--bin` | `stringArray` | `/captf/runtime` | The runtime command, one flag per element. Repeat it to pass a command with arguments. | | `--config` | `string` | `/captf/config` | Directory with the rendered root module and tfvars: the mount of the per-run Secret. | | `--module` | `string` | `/captf/module` | The module directory. | | `--providers` | `string` | `/captf/providers` | The provider mirror directory. Optional. | | `--workdir` | `string` | `/captf/work` | The writable work directory. | | `--image` | `string` | | The source image reference, echoed in the result. | | `--result` | `string` | `/dev/termination-log` | Where to write the result. | ### Runtime behavior | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--op` | `string` | | The operation: `apply`, `destroy`, `refresh`, `drift`, `restore` or `plan`. Required. | | `--backend-config` | `stringArray` | | A value for `init -backend-config`. Repeatable. | | `--lock-timeout` | `duration` | `5m0s` | How long a step waits for the state lock. | | `--stop-timeout` | `duration` | `1m0s` | How long an interrupted step has to stop before the runner sends it `SIGKILL`. | | `--force-unlock` | `string` | | A stale lock ID to force-unlock after `init`. | ### Approval and plan checks | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--guard-deletes` | `bool` | `false` | For `apply`: stop before a plan that deletes or replaces a resource, unless `--allow-deletes-hash` equals `--inputs-hash`. | | `--inputs-hash` | `string` | | The hash of the inputs the Job renders. An approval of a destructive plan must name it. | | `--allow-deletes-hash` | `string` | | The hash approved for a destructive plan. | | `--expect-plan` | `string` | | For `apply`: the approved plan hash. Stops with error kind `plan-changed` unless the plan’s hash matches. | | `--plan-key-file` | `string` | `/captf/plan-key/key` | For `plan` and an approved `apply`: the file with the key of the plan fingerprint. Both fail when it cannot be read. | See [The Destructive-Plan Guard]() and [Manual Plan Approval]() for how the manager sets these. ### Restore and events | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--restore-chunks` | `int` | `0` | For `restore`: the number of backup chunks under `/restore`. | | `--restore-resources` | `int` | `0` | For `restore`: the backup’s count of managed resources. When it is not `0`, `state list` must show at least one. | | `--event-object` | `string` | | Emit progress events about this object, as `////`. No events when unset. | | `--job-name` | `string` | | The Job the events relate to. | ### Example A plan Job for a cluster, as the manager would start it, trimmed to the flags that matter: ```sh /captf/bin/runner run \ --op=plan \ --image=registry.example.com/acme/cluster:v1.2.0 \ --inputs-hash= ``` - `` is the hash the manager computed for the inputs it rendered into `--config`. The manager supplies the rest through the defaults above and the mounts it builds. A run that finishes writes its result to the container’s termination log, which the manager reads from the Pod status. ## version ```text runner version ``` Prints the program name and its build information, then exits `0`. `runner --version=raw` prints the full build information instead. ```sh runner version ``` ## Exit codes | Code | Name | Meaning | | --- | --- | --- | | `0` | `ExitOK` | Every step succeeded, or an apply’s plan had no changes to apply. | | `1` | `ExitFailure` | A step failed, was interrupted, stopped before a destructive plan (blocked), or found that its approved plan had changed. Also returned when preflight or preparation fails, or when the result cannot be written. When the failing step has its own exit code of `1` or more, the runner returns that code. | | `2` | `ExitUsage` | Bad input: an unknown `--op`, a bad flag, unexpected arguments or invalid logging flags. | ## Result error kinds `runner run` always writes a result document, to `--result` (the termination log by default). A successful run has `error: null`. A failed run carries an `error` object with a `kind`, the `step` that failed and a bounded `tail` that summarizes the failure. The `kind` is one of: | Kind | Meaning | | --- | --- | | `image-layout` | The image does not follow the [image contract](): the module or runtime is not where the contract puts it. Fix the image. | | `step` | A runtime step (`init`, `plan`, `apply` and so on) failed, or the runner’s own setup failed, such as preparing the work directory or assembling a restore. A bad flag also reports this kind, with step `flags`. | | `interrupted` | The Job was canceled while a step ran, for example by a drain, an eviction or a deletion. It does not count toward retry backoff. | | `blocked` | A guarded `apply` stopped before a plan that deletes or replaces resources, because no approval names the inputs hash. It changed nothing, and the manager does not retry it until the inputs or the approval change. | | `plan-changed` | An `apply` approved for one plan (`--expect-plan`) found a different plan. It changed nothing, carries the new plan and waits for it to be approved. | > [!NOTE] > > **Results are small by design** > > The kubelet truncates a termination message at 4096 bytes. The runner drops optional detail, such as the plan’s resource list, before it drops the counts and hashes the manager needs. The full output of every step is in the Pod log. > [!NOTE] > > **See also** > > - [Image Contract]() for the paths the runner uses. > - [Retries and Backoff]() for how the manager treats each failure. > - [Choosing the Operation]() for what selects each `--op`. # tfcapi-lint CLI `tfcapi-lint` checks a Terraform or OpenTofu module, or the OCI image built from it, against the CAPTF module contract and image contract. It reads the module’s files and the image’s layers. It never runs `init`, `plan` or `apply`, and it needs neither `terraform` nor `tofu`. This page is the reference: every command, flag, exit code and check. For a walkthrough, registry credentials and CI patterns, see [tfcapi-lint](). ## Install Each provider release attaches one static binary per platform and a checksum file: `tfcapi-lint-linux-amd64`, `tfcapi-lint-linux-arm64`, `tfcapi-lint-darwin-amd64`, `tfcapi-lint-darwin-arm64`, `tfcapi-lint-windows-amd64.exe` and `tfcapi-lint-checksums.txt` (SHA-256). The binary has the same version as the provider release. There is no Windows ARM64 build, and `go install` does not work because the repository is a Go workspace. ```sh curl -fsSLO "https://github.com/captf-io/cluster-api-provider-terraform/releases/download//tfcapi-lint--" install -m 0755 "tfcapi-lint--" /usr/local/bin/tfcapi-lint tfcapi-lint version ``` - `` is the provider release, for example `v0.1.0`. - `` and `` pick a binary from the list above. The [tfcapi-lint]() guide shows how to verify the checksum. ## Synopsis ```text tfcapi-lint image --role= [flags] tfcapi-lint module --role= [flags] tfcapi-lint version [--json] tfcapi-lint --version[=raw] ``` `` is `cluster`, `machine` or `machinepool`. A bare `tfcapi-lint` prints its usage and exits `3`. ## Global flags | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--version` | `version` | `false` | Print version information and exit. `--version=raw` prints the full build information. `--version=vX.Y.Z` sets the reported version. | ## module ```text tfcapi-lint module --role= [flags] ``` Lints a module directory against the contract. It reads the `.tf`, `.tf.json`, `.tofu` and `.tofu.json` files in `` and in the local modules it calls, and runs the [module checks](<#module-checks>) for the role. | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--role` | `string` | | The module role: `cluster`, `machine` or `machinepool`. Required. | | `--contract` | `string` | `v1alpha1` | The contract version to lint against. | | `--json` | `bool` | `false` | Print a JSON report instead of text. | | `--strict` | `bool` | `false` | Count warnings as errors for the exit code. | | `--allow-warning` | `stringSlice` | | Downgrade this check ID’s warnings to info. Repeatable, or a comma-separated list. Errors cannot be allowed. | ```sh tfcapi-lint module --role cluster --strict ./modules/cluster ``` ## image ```text tfcapi-lint image --role= [flags] ``` Lints a built source image against the image contract. It pulls the manifest and layers without running the image, runs the [image checks](<#image-checks>), and runs the module checks on the module in `/captf/module`. `` is a registry reference such as `registry.example.com/acme/machine:v1.0.0`, or `oci:` for a local OCI image layout. It authenticates with the same credential files as `docker` and `podman`. | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--role` | `string` | | The module role. Required. | | `--contract` | `string` | `v1alpha1` | The contract version to lint against. | | `--json` | `bool` | `false` | Print a JSON report instead of text. | | `--strict` | `bool` | `false` | Count warnings as errors for the exit code. | | `--allow-warning` | `stringSlice` | | Downgrade this check ID’s warnings to info. Repeatable. Errors cannot be allowed. | | `--platform` | `string` | `linux/amd64` | The platform to check in a multi-platform image. | | `--all-platforms` | `bool` | `false` | Check every platform the image publishes. | | `--insecure` | `bool` | `false` | Allow a plain-HTTP registry. | ```sh tfcapi-lint image --role machine --strict registry.example.com/acme/machine:v1.0.0 ``` Lint a local image without a registry by saving it as an OCI layout first: ```sh podman save --format oci-dir -o ./image tfcapi-lint image --role machine --strict oci:./image ``` - `` is the local image name, for example `localhost/acme/machine:dev`. If you set `--platform` and the image is for a different platform, the result is an `image/platform` error. An image with no manifest for the requested platform gets the same error. ### Output The text format prints one line per finding, then a summary: ```text error input/required -:0 contract input bootstrap_data is not declared as a variable 0 error(s), 1 warning(s), 0 info (read terraform files) ``` A finding that concerns the whole module or image shows `-` as its file. With `--json`, the report has a `findings` array and a `summary`. Each finding has `id`, `severity` (`error`, `warning` or `info`), `file`, `line` and `message`. `summary` counts `error`, `warning` and `info`. An image report also describes the image it checked. Findings are ordered by file, line and ID, so output is stable between runs. `--allow-warning` changes a warning to `info` and says so in its message. ## version ```text tfcapi-lint version [--json] ``` Prints the program name and version, then exits `0`. | Flag | Type | Default | Description | | --- | --- | --- | --- | | `--json` | `bool` | `false` | Print JSON with `version`, `commit`, `date` and `contract`, the list of contract versions this build can lint against. | ```sh tfcapi-lint version --json ``` ## Exit codes | Code | Name | Meaning | | --- | --- | --- | | `0` | `ExitOK` | No errors, and with `--strict`, no warnings. Info findings never fail a run. | | `1` | `ExitFindings` | At least one error, or a warning under `--strict`. | | `2` | `ExitUnparsable` | The module could not be read or parsed, or the image could not be pulled or extracted. | | `3` | `ExitUsage` | A bad command line: a missing `--role`, an unknown role or contract version, a wrong number of arguments, or an invalid image reference. | In CI, treat any non-zero code as a failed step. Use `--strict` to fail on warnings, which is the recommended default. ## Checks Every check has an ID of the form `/`, a severity and a fix. Run `--json` to see IDs in a report. A finding’s severity decides the exit code: | Severity | Effect | | --- | --- | | `error` | Fails the run. Cannot be allowed. | | `warning` | Fails the run only with `--strict`. `--allow-warning ` downgrades it to `info`. | | `info` | Never fails the run. | The `module` command runs the module checks. The `image` command runs the image checks and the module checks on the image’s module. All module checks apply to every role unless a row says otherwise. ### Module checks #### Inputs The contract’s inputs are the variables the controller passes to your module. See the [common contract]() and the role pages for the full lists. | ID | Severity | What it checks | How to fix | | --- | --- | --- | --- | | `input/required` | error | A contract input is not declared as a variable. | Declare a `variable` for it. | | `input/type` | error | A contract input’s type does not accept what the generated root passes. A missing type, or a type the linter cannot read, is a warning. | Declare a type that accepts the contract’s type. | | `input/default` | warning | An input the controller always sets to a non-null value has a default, which would hide a controller mistake. Nullable inputs may default to `null`. | Remove the `default`. | | `input/sensitive` | warning | `bootstrap_data` is declared without `sensitive = true`, though it carries the bootstrap payload. | Add `sensitive = true`. | | `input/reserved` | error | A variable uses the reserved `captf_` prefix but is not a contract input. | Rename the variable. | | `input/tags-declared` | error | The module does not declare `captf_tags`, the common input every module must accept. | Declare `captf_tags`. | | `input/tags-unused` | warning | `captf_tags` is declared but never used, directly or through a local module that uses it. | Apply `var.captf_tags` to every resource that can carry tags. Allow the warning if the provider cannot tag anything. | | `input/user-variable-default` | warning | A variable outside the contract has no default, so an object that does not set it in `spec.variables` or `variablesFrom` fails to apply. | Give it a `default`. | #### Outputs | ID | Severity | What it checks | How to fix | | --- | --- | --- | --- | | `output/required` | error | A contract output is not declared. | Declare an `output` for it. | | `output/health` | error | The `health` output is not declared. Every module must report its health. | Declare `health` as the contract describes. | | `output/reserved` | warning | An output uses the reserved `captf_` prefix. | Rename the output. | | `output/endpoint-never-set` | warning | Cluster role only. `control_plane_endpoint` is a literal `null`, so a `KubeadmControlPlane` cluster with no user-set endpoint waits forever. | Output the endpoint your infrastructure creates. | | `output/provider-id-list-shape` | error or warning | Machinepool role only. It is an error when `provider_id_list` is not declared. It is a warning when the expression filters on health or state, or is not sorted with `sort()`. | Declare it, list every non-terminated member, and wrap the list in `sort()`. | | `pool/autoscaling-ignore-changes` | warning | Machinepool role only. The module uses `var.autoscaling`, but no resource ignores changes to its desired capacity, so each apply resets the cloud autoscaler. | Add the desired-capacity argument to `lifecycle { ignore_changes }`. | #### Module source CAPTF owns the backend and the credentials of a run, so a module must not configure them. | ID | Severity | What it checks | How to fix | | --- | --- | --- | --- | | `module/backend` | error | The root or a nested module declares a `terraform { backend }` block. | Remove it. The generated root owns the backend. | | `module/cloud` | error | The root or a nested module declares a `terraform { cloud }` block. | Remove it, for the same reason. | | `module/provider-config` | warning | A provider block sets a credential argument, such as `access_key`, `token` or `password`, to a string literal. A reference such as `var.token` is fine. | Take credentials from the identity, not from the module source. | | `module/source-escape` | error | A local module call resolves outside the root module directory, directly or through a symlink, or does not resolve. A module file that is a symlink out of the root is also reported. Such code is never linted. | Keep every local module inside the root directory. | | `module/tofu-shadow` | warning | A `.tofu` file shadows a `.tf` file, and OpenTofu loads different declarations than Terraform does. | Make the two files agree, or ship only one. | | `module/version` | info | The module declares no `required_version`. | Declare the runtime versions you tested with. | ### Image checks The image checks read the image’s files and config. They apply to every role. The [image contract]() describes the paths and labels they enforce. #### Content | ID | Severity | What it checks | How to fix | | --- | --- | --- | --- | | `image/module-present` | error | `/captf/module` has no `.tf`, `.tf.json`, `.tofu` or `.tofu.json` file at its top level. | Copy the module into `/captf/module`. | | `image/module-link` | error | A link under `/captf/module` does not end at a regular file inside it. | Replace the link with the file. | | `image/module-readable` | warning | With a numeric non-root user, a path under the module or provider mirror is not readable or traversable. The finding lists up to five paths. | Fix the file modes or ownership. | | `image/runtime-present` | error | `/captf/runtime` is missing, is not a regular file, or is not executable by the image user. | Install the runtime there, executable. | | `image/runtime-version` | warning | The runtime’s file or link name looks like a different runtime than `io.captf.runtime` says. | Correct the label or the runtime. | | `image/reserved-paths` | error | Files exist under `/captf/work`, `/captf/bin`, `/captf/config` or `/var/run/captf/credentials`, which the Job mounts over. | Remove them. | | `image/entrypoint` | info | `ENTRYPOINT` or `CMD` is set. The runner replaces both. | None needed. | | `image/platform` | error | The image is not for the requested platform, or the index has no manifest for it. | Build for the platform, or pass `--platform`. | #### Provider mirror The mirror at `/captf/providers` is optional. Without it, `init` needs registry access when the Job runs. | ID | Severity | What it checks | How to fix | | --- | --- | --- | --- | | `image/providers-absent` | info | The image has no provider mirror. | Add a mirror to run without registry access. | | `image/providers-layout` | error | An entry is neither a packed (`HOST/NS/TYPE/*.zip` with `*.json`) nor an unpacked (`HOST/NS/TYPE/VERSION/TARGET/`) mirror entry. | Use the mirror layout. | | `image/providers-complete` | error | A mirror link is broken, or the mirror has no package for the checked platform for a provider the module requires. Built-in providers are exempt. | Mirror every required provider for the platform. | #### Labels and user | ID | Severity | What it checks | How to fix | | --- | --- | --- | --- | | `image/label-role` | info or error | `io.captf.role` is not set (info), or names a different role than `--role` (error). | Set the label to the role. | | `image/label-contract` | info or warning | `io.captf.contract` is not set (info), or names a different version than `--contract` (warning). | Set the label to the contract version. | | `image/label-capacity` | info or error | Machine role. `io.captf.capacity` and `io.captf.node-info` do not parse. Missing labels are info: no scale-from-zero capacity. Other roles get info only. | Make the labels valid JSON of the documented shape. | | `image/user-root` | warning | The image runs as root, so the Job cannot run under the restricted Pod Security Standard. | Set a numeric non-root `USER`. | | `image/user-unresolved` | warning | The image user is a name the linter cannot resolve, so permission checks are approximated. | Use a numeric uid. | > [!TIP] > > **Allow a warning only with a reason** > > `--allow-warning` hides a warning from `--strict` but not from the report, where it shows as `info`. Pass one flag per ID you have decided to accept, and keep the list in the CI file so a reviewer sees each exception. > [!NOTE] > > **See also** > > - [tfcapi-lint]() for installing, strict mode and CI. > - [Module Contract]() and [Image Contract]() for what the checks enforce. # Tags Every page carries tags for its reader (cluster users, module authors, operators, contributors, evaluators), its kind (tutorial, how-to, concept, runbook, reference) and its topic. Pick a tag to see every page that has it. # Contributing # Contributing This page is for anyone changing CAPTF itself: the repository layout, the prerequisites, the build/lint/test/verify loop, the conventions the checks enforce, and running the manager against a real cluster. ## Repository layout | Path | Holds | | --- | --- | | `api/` | The `v1alpha1` Go types: its own module in `go.work`, so the CRD types carry no dependency on controller-runtime or the manager. | | `cmd/manager`, `cmd/runner`, `cmd/tfcapi-lint` | The three binaries. Each has an `app` package with its wiring and an `app/options` or equivalent for flags, so `main.go` stays a thin entry point. | | `internal/` | Everything the binaries share, one package per concern: `controllers` (one subpackage per reconciled kind, plus `shared` for the common reconcile flow and `sweep` for the orphan RBAC sweep), `jobs`, `runner`, `state`, `identity`, `rbac`, `runlease`, `inputs`, `render`, `outputs`, `locks`, `conditions`, `contract`, `hash`, `ownership`, `plankey`, `webhooks`, `lint`, `feature`, `imageinspect`, `manager`, `metrics`, `strutil`. | | `config/` | Kustomize bases: `crd`, `rbac`, `webhook`, `manager`, `certmanager`, assembled by `default`; `network-policy` and `prometheus` are separate optional overlays, and `samples` holds example custom resources. | | `templates/` | The `clusterctl generate` templates and flavors, and the Go tests that decode them. | | `hack/` | Build and verify tooling: pinned tool installers under `hack/tools`, the `verify-*.sh`/`check-*.sh` scripts `make verify` runs, `hack/godoccheck`, `hack/covercheck` with its per-package floors in `hack/coverage-floors.txt`, and the dev scripts behind `make run`. | | `test/` | A third module in `go.work` for the e2e-tagged test environment and suites: `framework/` (kind, images, providers and diagnostics, unit-tested with fakes), `env/lifecycle/` (the `testenv-*` operations) and `e2e/` (the foundation and no-op suites). See [Testing](). | | `internal/*/testdata` | Golden files and frozen fixtures that unit tests read, such as real captured Terraform and OpenTofu state under `internal/state/testdata/fixtures`. | The no-op demo modules are not in this repository: they live in [captf-io/noop-modules](). The documentation lives in a separate repository, [captf-io/captf-io.github.io](), with the rest of captf.io, published at [https://captf.io/docs/](); see [Writing Documentation](). Each Go module (`.`, `api` and `test`) is listed in `go.work`. `hack/verify-modules.sh` (`make verify`) checks that none carries a `replace` directive and that they agree on the Kubernetes and controller-runtime versions they share. Because `api` only exists inside the workspace, `go mod tidy` run from the repository root does not update its `go.mod`; add or bump one of its dependencies by editing `api/go.mod`’s `require` block directly, then run `go mod tidy` inside `api/` with `GOWORK=off`. > [!NOTE] > > **Before you begin** > > - Go, matching the version `go.work` declares, `make`, `git`, `jq` and `curl`. > - `podman` (the default `CONTAINER_TOOL`) or `docker`, to build images. > - `python3`, with the `jsonschema` and `PyYAML` packages, for `make verify-schemas`, `make verify-metadata`, `make promtool-check`, `make promtool-test` and `make release-assets`. > - `podman` or `docker`, running, for the [test environment]() (`make testenv-up` and the `e2e-*` targets). Everything else — `controller-gen`, `kustomize`, `golangci-lint` (plus its kube-api-linter build), `clusterctl`, `promtool`, `goreleaser`, `goimports` and `gotestsum` — is a pinned binary `make` downloads for you. ## The build, lint, test and verify loop 1. Install the pinned tools once: `make tools`. Each lands in `hack/tools/bin` as `-`, plus an unversioned symlink; bumping a version in the `Makefile` re-downloads it. 2. After changing `api/v1alpha1` or a controller’s markers, regenerate the deepcopy code and the CRD/RBAC/webhook manifests: `make generate manifests`. `make verify-gen` (part of `make verify`) fails if either is stale. 3. Format and lint: `make fmt lint`. `fmt` runs `gofmt -s` and `goimports -local github.com/captf-io/cluster-api-provider-terraform`; `lint` runs `golangci-lint` (`.golangci.yml`) in every module, then kube-api-linter (`.golangci-kal.yml`) on `api/`. `lint` and `vet` also check the `test` module’s e2e-tagged code. `make lint-fix` reruns both linters with their auto-fixers. 4. Run the unit tests: `make test`. See [Testing]() for what it covers and how to run less than everything. CI runs `make test-cover` and `make cover-check` instead, which also enforce a coverage floor per package. 5. Before sending a change, run `make verify`: every check listed in [Make Targets](), including that generated code is current and that e2e code stays out of the default test run. A change to anything the documentation’s reference pages describe (a flag, a condition, an event, a metric, an annotation, a field of a kind) also needs that page updated in `captf-io/captf-io.github.io`: see [Reference pages](). 6. `make build` compiles every `cmd/*` binary to `bin/`; never to the repository root. To see a change work, run the manager against a cluster: see [Running the manager](<#running-the-manager>). The full target list, grouped the same way, is in [Make Targets](). ## Conventions - **Commits**: an imperative subject in `: ` form, such as `runner: name the exit codes` — check `git log` for the subsystem names already in use. - **License header**: every hand-written Go file starts with the Apache 2.0 header the other files in its package carry; `hack/boilerplate.go.txt` is the header `controller-gen` writes on `api/v1alpha1/zz_generated.deepcopy.go`. - **Documentation comments**: `hack/godoccheck` (`make verify-godoc`, part of `make verify`) enforces one rule per declaration and one per package, everywhere except generated files and `hack/tools`: - Every function, method and type has a doc comment starting with its name; every interface method and top-level `const` or `var` group needs only a doc comment, not one starting with its name. - Every named parameter is mentioned by name in that comment. Receivers, parameters named `_`, and a `Test`/`Benchmark`/`Fuzz` function’s conventional `*testing.T`/`*testing.B`/`*testing.F` parameter are exempt. - A function or method that returns anything says what it returns, using “return”, “returns”, “returned” or “reports”. - Every non-generated, non-external-test package has a `doc.go` whose package comment is a real overview of at least 400 characters. - **Generated code**: `zz_generated.deepcopy.go` and the CRD/RBAC/webhook manifests come from `make generate manifests`. Never hand-edit a generated file: `make verify-gen` fails when the deepcopy code or the manifests no longer match their source. The documentation’s reference pages are written by hand (see [Reference pages]()). - **Test tiers**: code that needs a real cluster carries the `e2e` build tag and lives only in `test/e2e/` or `test/env/lifecycle/`; `make verify-test-tiers` fails otherwise. See [Testing](). - **Coverage**: a package’s coverage must not fall below its floor in `hack/coverage-floors.txt`. When a package’s coverage rises, raise its floor in the same change; never lower one without a reason on its line. ## Common changes Each change below also needs its reference page in `captf-io/captf-io.github.io` updated by hand, in the same change; run `tools/check_reference.py` there, with `--provider` pointed at this checkout, to list any name a page is missing (see [Reference pages]()). - **Adding a manager flag**: register it in `cmd/manager/app/options/options.go`, then document it in [Manager Flags](). - **Adding a condition reason**: condition types, their reasons and each reason’s doc comment live in `api/v1alpha1/conditions_consts.go`; `ConditionReasons` returns the full type-to-reason table, and every reason constant must be in it. A reason is set from `internal/conditions` when it applies across controllers, or from `internal/controllers/shared` when it belongs to one reconcile flow. Afterward, run `make generate manifests` to regenerate the CRD schema and the deepcopy code, and document the reason in [Conditions](). `make verify-gen` and `make test` fail until the generated files and the reason table are complete. An event reason follows the same shape: it lives in `internal/controllers/shared/events.go`, or `internal/runner/events.go` for a runner event, with a doc comment. Document it in [Events](). A `tfcapi-lint` check ID in `internal/lint` needs a doc comment too, and an entry in [tfcapi-lint CLI](). ## Running the manager There are two ways to run the manager you are changing. Both need a cluster that has the CAPTF CRDs, cert-manager and Cluster API. - **The test environment** is the closest to a real install. `make testenv-up` builds the manager image from your tree, creates a kind cluster and installs cert-manager, Cluster API and CAPTF. After a change, `make testenv-reload` rebuilds the image and rolls the manager Deployment over to it. `make testenv-logs` collects diagnostics and `make testenv-down` deletes the cluster. The same image is the runner image, so a change under `internal/runner` or `cmd/runner` reaches Jobs on the next reload. - **`make run`** builds `bin/manager` and runs it on your machine against the cluster in your current `kubeconfig`, with leader election off and a self-signed webhook certificate in `bin/dev-webhook-certs`. It is quick, but the Jobs it creates run the image in `RUNNER_IMAGE` (default `IMG`), not your working tree: run `make docker-build` and make that image available to the cluster first when you change the runner. Pass extra manager flags with `ARGS`. > [!WARNING] > > **Point `make run` at a cluster you can break** > > `make run` uses whatever cluster your current `kubeconfig` context names. The test environment never reads `~/.kube/config`; it writes its own `bin/testenv//kubeconfig` and an `env.sh` that exports it. > [!NOTE] > > **See also** > > - [Testing]() > - [Releasing]() > - [Writing Documentation](), [Contributing to the Docs]() and [Updating the Website]() for changes to captf.io itself. # Testing This page describes CAPTF’s test suite: what `make test` runs, its tiers, the end-to-end test environment, coverage floors, golden files, and running less than the whole suite. ## Running the tests `make test` runs `go test -race -count=1 ./...` in every module (`.`, `api` and `test`). That is the unit tier, and nothing it runs creates a Kubernetes cluster, calls a real cloud API, or invokes `terraform` or `tofu`. - Controllers and webhooks run against the controller-runtime fake client, never a real API server; there is no [envtest]() in this repository. - Job execution is faked the same way: a reconciler test asserts on the `batch/v1.Job` CAPTF would create, and a runner test drives the runner’s own logic directly, without a pod ever starting. - `internal/state`, `internal/outputs` and `internal/locks` instead read real `kubernetes`-backend state: [`internal/state/testdata/fixtures`]() holds state Secrets and lock Leases captured once from real Terraform and OpenTofu runs, checked in as frozen data. Regenerating them needs a capture setup that does not exist in this repository; treat those files as read-only. - `go test ./templates/` decodes every shipped `templates/*.yaml` object strictly into its API type and checks the values `hack/verify-templates.sh` renders; it also runs CAPI’s own `ClusterClass` admission webhook and topology generator, applying the class patches, against `clusterclass-noop.yaml`. - The packages under `test/framework` (kind, images, providers, waits, diagnostics) are unit-tested with fakes. They never start a process or touch `podman`. ## Test tiers | Tier | What | How it runs | | --- | --- | --- | | Unit | Everything above, including reconcilers driven through several packages at once against the fake client, such as `internal/controllers/shared`’s `TestBringUpJobCount`, which reconciles a whole cluster bring-up and asserts the resulting Job count. | `make test`, `make test-cover` and CI, always. | | E2E | `test/e2e/...` and `test/env/lifecycle`, against a real kind cluster on `podman` or `docker`. | Only through the `testenv-*` and `e2e-*` targets, or `go test -tags=e2e` with `-run`. CI does not run them yet. | Three guards keep the tiers apart: - **The build tag.** Every Go file under `test/e2e/` and `test/env/lifecycle/` starts with `//go:build e2e`, so `go test ./...` never compiles it. `make verify-test-tiers` (part of `make verify`, and of CI) fails if one lacks the tag, or if any other Go file uses it. `hack/verify-test-tiers_test.sh` tests that script against fixtures; it is not part of `make test`, and `verify-test-tiers` runs it first. - **`-run`.** Each e2e package’s `TestMain` refuses to run unless `-run` selects a test, so a bare `go test -tags=e2e ./...` does nothing. - **Compiled without running.** `make vet` and `make lint` add a `-tags e2e` pass over the `test` module, so tagged code cannot rot unnoticed. Besides `go test`, two `make verify` checks exercise the release assets without a real cloud: `make verify-templates` renders `templates/` with the pinned `clusterctl`, and `make verify-local-repository` runs `clusterctl generate provider` and `generate cluster` against a local repository of the release assets, offline. `hack/check-metadata_test.sh` and `hack/version_test.sh` test `hack/check-metadata.sh` and `hack/version.sh` against fixtures under `hack/testdata`; `make verify-metadata` and `make verify-version` run them before the check itself. ## End-to-end tests The e2e tier needs `podman` or `docker`, and network access on the first run to download the pinned kind, Cluster API and cert-manager assets. Every version and image is pinned in `test/framework/versions.go`. ### The test environment `make testenv-up` builds the manager image from your working tree and brings up a kind cluster with cert-manager, Cluster API (core, kubeadm bootstrap and kubeadm control plane) and CAPTF. It takes about 6 minutes cold and about 15 seconds when it reuses a cluster. Work in a loop: ```sh make testenv-up source bin/testenv/captf-test-dev/env.sh kubectl get pods -A make testenv-reload make testenv-logs make testenv-down ``` `env.sh` points `KUBECONFIG` at the test cluster. `testenv-reload` rebuilds the image and rolls the manager over to it; `testenv-logs` writes a diagnostics bundle without Secrets; `testenv-down` deletes the cluster and keeps the artifacts. Everything for a cluster lives in `bin/testenv//`. [Make Targets]() lists the targets and their variables. The environment is built so it cannot touch anything else on your host: - Cluster names must start with `captf-test-`, and `testenv-down` only ever deletes such clusters. - The nodes join a dedicated `captf-test` network, not kind’s default one. - It never reads `~/.kube/config`; every child process gets the cluster’s own kubeconfig. - It never changes host sysctls or limits, and it never deletes a cluster implicitly: `testenv-up` refuses an existing one unless `CAPTF_TESTENV_REUSE=1`. ### The e2e suites | Target | Suite | What it proves | | --- | --- | --- | | `make e2e-foundation` | `TestFoundation` in `test/e2e/foundation` | In six ordered stages, a failed one stopping the run: the `captf-test-e2e` cluster builds, the base components and the CAPTF install are healthy, a real reconcile of a `TerraformClusterIdentity` works and the cluster holds still for a stability window, then a green light is written. | | `make e2e-noop` | `TestNoop` in `test/e2e/noop` | The published no-op modules, driven through real Cluster API objects, move data end to end with no cloud: the CAPI spec into the module inputs, the outputs into CAPTF status and on into CAPI, the cluster’s exports into the machine and pool inputs, the pinned digests into every later Job, and deletion into a full cleanup. | | `make e2e-down` | | Deletes the e2e cluster; the artifacts are kept. | Run `e2e-foundation` first. `e2e-noop` fails at once, and never skips, unless the foundation suite’s green light is still valid: the cluster exists, its pins and manager image match this build, every stage passed, and the light is younger than `CAPTF_E2E_GREENLIGHT_MAX_AGE` (24 hours by default). The foundation suite keeps a passing cluster for later tests unless `CAPTF_E2E_TEARDOWN=1`, and on failure it keeps the cluster too and writes a diagnostics bundle into `bin/testenv//artifacts/`. Each `e2e-noop` run uses a fresh namespace `e2e-noop-` and removes what it created. Set `CAPTF_E2E_NOOP_BAD_DIGEST=1` to run its negative check instead, in which one machine gets a nonexistent image digest and the suite must fail on the image pull. `test/README.md` in the provider repository has the full stage-by-stage description, the variables and the allowlists of known-benign log lines. ## Coverage `make test-cover` runs the unit tests with `-race` and `-covermode=atomic` through `gotestsum`, and writes a profile and a JUnit report per module into `bin/`. `make cover-check` then reads the profiles and fails when a package falls below its floor in `hack/coverage-floors.txt` (the checker is `hack/covercheck`). CI runs both and writes the per-package table to the job summary. ```sh make test-cover make cover-check go tool cover -html=bin/cover-root.out -o bin/cover.html ``` The floors are a ratchet, set per package: - A `default` line sets the floor for every package without its own line, currently 80 percent. - A ` ` line sets one package’s floor. Packages are import paths relative to the module path. - A ` exempt` line reports a package but never gates it. Each entry carries its reason in a comment; the `cmd/*` packages are exempt because the logic they call lives in their `app` packages. When a package’s coverage rises, raise its floor in the same change. Never lower a floor or add an exemption without a reason on the line. ## Golden files Several packages pin an exact rendered output as a checked-in file and compare against it on every run, rather than asserting field by field: | Package | What it pins | | --- | --- | | `internal/contract` | The rendered JSON of fully populated contract inputs. | | `internal/jobs` | The `batch/v1.Job` of every operation, as YAML. | | `internal/outputs` | Rendered output fixtures. | | `internal/render` | The generated root module. | | `cmd/tfcapi-lint` | Its `--json` output. | Run the affected test with `UPDATE_SNAPSHOTS=1` to rewrite its golden files. `make verify-schemas` also validates the golden inputs and outputs of `internal/contract` and `internal/outputs` against the contract schemas. > [!WARNING] > > **Read the diff before committing a rewrite** > > Read the diff before committing it: a passing rewrite is not the same as a correct one. ## Running one package, or one test ```sh go test -race ./internal/jobs/... go test -race -run TestGoldenJobs ./internal/jobs/... UPDATE_SNAPSHOTS=1 go test ./internal/jobs/... -run TestGoldenJobs ``` `go.work` makes the `api` module’s tests reachable from the repository root too: `go test ./api/...` needs no `cd`. The `test` module’s unit tests run from its own directory: `cd test && go test ./framework/...`. > [!NOTE] > > **See also** > > - [Contributing]() > - [Writing Documentation]() > - [Releasing]() # Releasing This page is for whoever cuts a CAPTF release: what a release consists of, the checklist, and installing the assets before or instead of publishing them. A release is a tag `vX.Y.Z` (or `vX.Y.Z-rc.N`) on a clean `main`, the manager image `ghcr.io/captf-io/cluster-api-provider-terraform:vX.Y.Z`, and a GitHub release with the clusterctl assets and the `tfcapi-lint` binaries. The same image is the runner image. The no-op demo module images are not part of a CAPTF release: they are built and tagged from [captf-io/noop-modules](). > [!WARNING] > > **Tags never move** > > Nothing is force-pushed: a bad release candidate gets a new `-rc.N`. > [!NOTE] > > **Before you begin** > > - Write access to push a signed tag and create a GitHub release. > - Registry push access for `ghcr.io/captf-io/cluster-api-provider-terraform`. > - `gh` authenticated against this repository, for `make release-github`. > - `skopeo`, for `make release` to read the pushed image’s registry digest. ## Assets | Asset | Built by | | --- | --- | | `infrastructure-components.yaml` | `make manifests-release`: `config/default` with the release image, and `CAPTF_MANAGER_IMAGE` set to the same image. | | `metadata.yaml` | The repository root file; `hack/check-metadata.sh` enforces an append-only `releaseSeries`. | | `cluster-template.yaml`, `cluster-template-clusterclass.yaml`, `clusterclass-noop.yaml`, `identity.yaml` | [`templates/`](). | | `tfcapi-lint--`, `tfcapi-lint-checksums.txt` | GoReleaser (`.goreleaser.yaml`), run by `make release-assets`, for Linux and macOS on amd64 and arm64 and Windows on amd64; see [tfcapi-lint](). | GoReleaser only builds the `tfcapi-lint` binaries and checksums: it does not build the image or publish the release. `make release` builds and pushes the image, and `make release-github` publishes. ## Checklist 1. If this release starts a new minor series, append it to `metadata.yaml` `releaseSeries`. Never remove or change an existing series. 2. Update the [contract changelog]() for anything that changes what a module sees or must implement, and [Upgrades]() for anything an operator needs to do when moving to this release. Make sure the documentation’s reference pages ([Reference pages]()) match the code going into the release: from a `captf-io/captf-io.github.io` checkout, `tools/check_resources.py` and `tools/check_reference.py`, pointed at this checkout with `--provider`, list anything they miss. 3. `make lint test verify` is green, and CI is green on the commit you tag; CI also runs `make cover-check`, the per-package coverage gate. `verify` includes `verify-local-repository`, which generates the provider and the default and clusterclass flavors from a clusterctl local repository of the release assets, offline. 4. Tag and push the tag: `git tag -s vX.Y.Z && git push origin vX.Y.Z`. 5. Build and push the image and build the assets: ```sh make release VERSION=vX.Y.Z ``` `release-preflight` refuses a dirty tree, an untagged HEAD, a malformed version, or a metadata change that is not append-only. `release` then builds and pushes the manager image, reads its registry digest with `skopeo`, and builds the assets with the image pinned by that digest, so the published components never follow a moved tag. The assets land in `out/release/`. 6. Smoke-test from a local repository against a real cluster, by hand: see [Installing from a local repository](<#installing-from-a-local-repository>). `make e2e-foundation e2e-noop` also runs the opt-in e2e suites on a kind cluster, but they build the manager from your tree rather than install the release assets, so they do not replace this step; see [Testing](). 7. Publish: `make release-github VERSION=vX.Y.Z` writes `out/release/notes.md` from the commits since the previous tag and runs `gh release create` with every asset. ## Installing from a local repository To try the assets before publishing, or offline, copy `out/release/*` to `~/local-repository/infrastructure-terraform/vX.Y.Z/`: the directory name is the provider label `infrastructure-terraform`. Point a `clusterctl` config’s `url` at the local path instead of a release URL (see [Register the provider]() for the rest of the config entry): clusterctl.yaml ```yaml url: file:///home//local-repository/infrastructure-terraform/vX.Y.Z/infrastructure-components.yaml ``` > [!NOTE] > > **A `file://` URL needs the absolute path** > > `` is your username on this machine; a `file://` URL needs the absolute path, so `~` does not work here. Pin the version at install time: ```sh clusterctl init --config clusterctl.yaml --infrastructure terraform:vX.Y.Z ``` `hack/verify-local-repository.sh` does the same layout in a temp directory and runs `clusterctl generate provider` and `generate cluster` (the default and clusterclass flavors) against it. > [!NOTE] > > **See also** > > - [Writing Documentation]() for the book’s own build and verification commands. > - [Contributing]() for the general build/lint/test/verify loop. # Contributing to the Docs This page walks through a change to captf.io, from a local preview to a pull request. The site’s source is the [captf-io/captf-io.github.io]() repository: the documentation under `docs/docs/`, the blog under `docs/blog/` and the landing page, all built together by [Zensical](). For how to write a page, see [Writing Documentation](); for the landing page, blog, navigation and theme, see [Updating the Website](). > [!NOTE] > > **Before you begin** > > - `git` and `make`. > - [`uv`](). It installs the Python version and the pinned Zensical release the site needs on first use, so there is nothing else to install. > - For the browser checks only: the Playwright Chromium headless shell (`uv run --with playwright playwright install chromium-headless-shell`). ## Layout | Path | What it holds | | --- | --- | | `docs/docs/` | The documentation, one Markdown file per page, published under `/docs/` | | `docs/blog/posts/` | Blog posts, published under `/blog/` | | `docs/index.md` | The landing page’s data (its copy is in `overrides/home.html`) | | `docs/assets/` | Brand marks, the social card and the landing page’s images | | `zensical.toml` | Site configuration: the navigation, plugins, Markdown extensions, header and footer content | | `overrides/` | Theme overrides: templates, stylesheets and scripts | | `includes/abbreviations.md` | Acronyms that get a hover tooltip on every page | | `tools/` | Generators and checks, described below | ## Preview the site 1. Clone your fork and change into it. 2. Start the preview server: ```sh make serve ``` It builds the site, serves it at [http://127.0.0.1:8001/]() and rebuilds when a page changes. To reach it from another machine, set `ADDR` to an address of yours, on the command line or in a `local.mk` file, which is not committed: `ADDR =
`. 3. Open the page you are changing. > [!NOTE] > > **Restart after template or configuration changes** > > The preview rebuilds pages as you save them. A change to `zensical.toml` or to a template under `overrides/` needs the server stopped and started again. ## Edit a page Find the page’s file from its URL: `/docs/user-guide/drift/` is `docs/docs/user-guide/drift.md`, and a section’s own page (`/docs/cloud-modules/`) is its `README.md`. Pages keep the paths they were first published under even where the navigation groups them differently, so the URL, not the tab, tells you where a file is. Keep to the [house style](). In particular, do not rename a heading other pages link to: links point at the anchor its text makes. Check for inbound links before renaming one: ```sh grep -rn '#' docs/ ``` `` is the heading in lowercase, with spaces as hyphens and punctuation dropped. The reference pages describe what the provider repository defines; see [Reference pages]() for keeping them in step with it. ## Add a page 1. Create the Markdown file in the directory of the part of the docs it belongs to, such as `docs/docs/user-guide/`. Its tags come from that directory’s `.meta.yml`. 2. Give it front matter with a `description:` and an H1, following [Page shape](). 3. Add it to the `nav` in `zensical.toml`, where readers will look for it. The entry’s text is the page’s title in the navigation. 4. Give it a sidebar icon and subtitle: add an entry to `tools/nav_meta.json`, then run: ```sh uv run python tools/apply_nav_meta.py --write ``` It refuses to write if a page has no entry, an icon does not exist or a subtitle is longer than 34 characters. 5. List the page in `SKIP` in `tools/gen_redirects.py`. That script keeps the URLs of the first edition of the docs (`/docs/.html`) working; a new page has no such URL. 6. Run `make gen` and `make build`. Avoid moving or renaming a published page: its path is its URL, and every link to it, inside the site and out, would break. ## Run the checks | Command | What it does | | --- | --- | | `make build` | Builds the site in strict mode; any warning fails it, including a broken link or anchor and a missing snippet file | | `make gen` | Rewrites the generated parts of `zensical.toml` and the redirect stubs; run it after adding, moving or re-describing a page or post, and commit what it changes | | `uv run --with playwright python tools/layout_audit.py` | With `make serve` running: tables squeezed too narrow and font sizes, at desktop, laptop and phone widths | | `uv run --with playwright python tools/mobile_audit.py ` | With `make serve` running: the header, tab row and navigation drawer at 25 phone, tablet and desktop sizes, with a screenshot of each state in `` | | `uv run python tools/check_resources.py --provider ` | Every field of every CAPTF kind is named on its page under `reference/resources/`; `` is a provider repository checkout | `make build` is the one to run on every change. Run the audits when a change touches the navigation, a template or a stylesheet, or adds a wide table. Both audits read the preview at [http://127.0.0.1:8001/](); set `CAPTF_SITE` to another address if yours listens elsewhere. ## Send the change 1. Branch from `main` and make one logical change per commit. 2. Write the commit subject as `: `, imperative and about 50 characters, then a body, wrapped at about 72 columns, that says why. The subsystem is the part of the docs: `docs` for changes across sections or the site itself, or the section’s directory, such as `user-guide`, `operator-guide`, `runbooks`, `concepts`, `cloud-modules`, `contract` or `reference`. Check `git log` for the names in use. 3. Push and open a pull request against `main`, filling in the template. Open an issue before writing a change to the module contract pages, or to any page that describes the API or the security model: such a change starts in the provider repository. The project’s [CONTRIBUTING.md]() covers licensing and conduct for every repository. > [!NOTE] > > **See also** > > - [Writing Documentation]() for the house style and the reference pages. > - [Updating the Website]() for the landing page, blog, navigation, header, footer and theme. > - [Contributing]() for changes to the provider itself. # Writing Documentation This page is the house style for the CAPTF documentation, published at [https://captf.io/docs/](). The site is built with [Zensical](): the pages are Markdown under `docs/docs/`, and `zensical.toml` holds the configuration and the navigation. Every page follows this style, and the strict build enforces the parts a tool can check. To preview, check and send a change, see [Contributing to the Docs](). For the landing page, blog posts, navigation and theme, see [Updating the Website](). ## Where things go The documentation is split by reader, one navigation tab each. Pick the tab by who reads the page, then the page type by what they need. | Tab | Reader | Page types | | --- | --- | --- | | Overview | Anyone meeting CAPTF for the first time | The introduction, Quick Start, architecture, glossary and what to know before adopting | | User Guide | People creating clusters with CAPTF | How-to: one task per page, in steps | | Cloud Modules | People using the reference cloud modules | What each cloud’s modules create and how to use them | | Module Authors | People writing Terraform or OpenTofu modules for CAPTF | The normative contract, plus how-to pages | | Operations | People installing and running the CAPTF manager | How-to pages and runbooks | | Troubleshooting | Anyone with something not working | Diagnosis by symptom, then the fix | | How It Works | Anyone who needs to understand how CAPTF works | Explanation: how and why, no step lists | | Reference | Everyone | Every resource field by field, and lookups: conditions, events, alerts, metrics, flags, keys | | Contributing | People changing CAPTF itself | How-to and conventions | The navigation in `zensical.toml` decides where a page appears, not its path: pages keep the paths they were first published under, so links and anchors keep working. Every fact lives on exactly one page. Other pages link to it instead of restating it. The page that owns a topic is the one whose title names it; when two pages could own a fact, the more specific one does. ## Reference pages Every page under `reference/` is written by hand: the [custom resources](), field by field, and the conditions, events, alerts, metrics, manager flags, Job environment, the runner and `tfcapi-lint` command lines, annotations, labels and finalizers, clusterctl variables and make targets. They explain each item (what it means, its default, what to do about it), not only list it. What they describe is defined in the provider repository ([cluster-api-provider-terraform]()), so a change there that adds, removes or renames one of these items needs the page updated in the same change. Two checks catch a page that falls behind; run them from this repository with a provider checkout: - `tools/check_resources.py` reads the provider’s CRDs and fails on any field of any kind that its page under `reference/resources/` does not name. - `tools/check_reference.py` reads the provider’s source (condition and event constants, alert rules, metric names, the commands’ flags, the Job’s environment, lint check IDs, `captf.io/` keys and finalizers, template variables and make targets) and fails on any name its page does not mention. Flags that come from libraries (logging, feature gates) are not in the provider’s source, so it does not check those. Both print each missing name. Neither can tell that a default or a meaning changed, so read the provider change and check the page’s wording too. Other pages link to the reference pages for field lists, flags, reasons, events, metrics and keys, and never copy those tables. A few files are copies of provider files rather than pages: the contract schemas under `module-author/contract/v1alpha1/schemas/`, the no-op machine module under `getting-started/examples/noop/machine/` and the Containerfiles under `module-author/examples/`. Refresh them from the provider when its versions change. ## Page shape - One H1, the page title, matching its entry in the `nav` of `zensical.toml`. - Front matter with a `description:`: one sentence of at most 160 characters saying what the reader gets from the page, always in double quotes. It feeds search results, link previews and `llms.txt`. - An `icon:` (a Lucide icon, `lucide/`) and a `subtitle:` (at most 34 characters) for the sidebar. Set these in `tools/nav_meta.json` and run `tools/apply_nav_meta.py --write`, not in the page. - A first paragraph that says what the page covers and who it is for. - For how-to pages and runbooks: a “Before you begin” box of prerequisites, then numbered steps, then how to confirm it worked. - A closing “See also” box when related pages exist. - Headings in sentence case: “Rotate the credentials”, not “Rotate The Credentials”. Navigation entries use title case. - Heading text stays stable once published, since links and alert `runbook_url`s point at the anchors derived from it. A how-to page, start to finish: docs/docs/user-guide/rotate-credentials.md ````markdown --- description: "Rotate the cloud credentials a TerraformClusterIdentity points at, without touching the clusters that use it." --- # Rotate Credentials Replace the credentials behind an identity. For operators who manage the credentials Secret. !!! info "Before you begin" - `kubectl` access to the identity's namespace. ## Replace the Secret 1. Write the new credentials to the Secret: ```sh kubectl create secret generic --from-file= \ --dry-run=client -o yaml | kubectl apply -f - ``` `` is the Secret the identity references, and `` holds the new credentials. ## Confirm it worked !!! success "" The next Job for each cluster that uses the identity succeeds. !!! related "See also" - [Identities and Credentials](identities.md) ```` The title in the H1 is the page’s name; its navigation entry in `zensical.toml` uses the same words. The page’s tags come from its directory’s `.meta.yml`, not from the page. ## Voice and wording - Address the reader as “you”. Use the present tense and the active voice. - Say what happens, not what “should” happen. Reserve MUST, MUST NOT, SHOULD and MAY (RFC 2119, in capitals) for the normative module contract. - American English: behavior, labeled, license, canceled. - “CAPTF” is the project. Spell out “Cluster API Provider Terraform (CAPTF)” on the introduction page only. - “Terraform or OpenTofu” in prose; `terraform` and `tofu` in code for the command-line tools. “OpenTofu” is always written with a capital O and T. - Kinds use their exact names in code spans on first mention in a section: `TerraformCluster`, `TerraformMachine`, `TerraformMachinePool`, their `*Template` kinds, and `TerraformClusterIdentity`. In running prose, “machine pool” is fine. - Expand an abbreviation on first use on each page: KubeadmControlPlane (KCP). Common ones (CAPI, KCP, RBAC and others in `includes/abbreviations.md`) also get a hover tooltip on every page. - No references to source files, functions or line numbers on Overview, User Guide or Operations pages. Describe the behavior. Module Author and Contributing pages may name source files when the reader needs them. - No `§` section references. Link to the heading instead. - No dates, review notes, “TODO”, “pending” or “planned” statements. State what is true now. If something does not exist yet, say that it does not exist. ## Markdown - Wrap prose at 80 columns. Tables, code blocks and headings are exempt. - Indent nested lists, and code blocks or paragraphs inside a list item, by 4 spaces: with 2 or 3, Python-Markdown renders them flat. - Fence every code block with a language: `sh`, `yaml`, `hcl`, `json`, `text`, `promql` or `dockerfile`. Add `title="main.tf"` when the block is a whole file with a known name. - Shell examples show commands only, without a `$` prompt, and use `` placeholders the reader replaces. Explain each placeholder below the block. - Outside fenced code blocks, a placeholder always goes in a code span: `` `` ``. A bare `` in prose or a table cell is read as an HTML tag and disappears from the page. - Tables use the compact style: `| a | b |` with a `| --- |` separator row. - Link to other pages with relative links to the `.md` file, including the anchor when you mean a section: `[drift](../concepts/drift-and-health.md#drift)`. Relative links stay inside `docs/`. - Link to files in the provider repository with a full `https://github.com/captf-io/cluster-api-provider-terraform/blob/main/...` URL. - Include real files instead of pasting them, with a snippet line whose path is relative to `docs/docs/`: `--8<-- "module-author/examples/Containerfile.opentofu"` inside a fenced block. - Boxes are admonitions: `!!! info "Before you begin"`, `!!! success ""` for the body of “Confirm it worked”, `!!! related "See also"`, and `danger`, `warning`, `note`, `tip` or `example` for caveats, each with a title that states the point. Indent the body by 4 spaces. Never nest them or put two in a row. - Show alternatives a reader picks one of (Terraform or OpenTofu, one YAML per kind) as content tabs, `=== "Label"`. Never tab steps a reader must do in sequence. - Diagrams are Mermaid code blocks (a `mermaid` fence). ## Checks From the root of the documentation’s source: | Command | Checks | | --- | --- | | `make build` | The site builds in strict mode: any warning fails it, including a broken link or anchor and a missing snippet file | | `make gen` | Rewrites the generated parts of the configuration: the redirects from the original page URLs, the `llms.txt` sections and the recent posts; run it after adding, moving or describing a page or post | | `tools/layout_audit.py` | Tables squeezed too narrow and font sizes, per window size | | `tools/mobile_audit.py` | The header, tabs and navigation drawer at every phone and tablet size | | `tools/check_resources.py` | Every field of every CAPTF kind is named on its page under `reference/resources/`; reads the CRDs from a provider checkout (`--provider `) | | `tools/check_reference.py` | Every condition, event, alert, metric, flag, variable, annotation, target and check in the provider’s source is mentioned on its reference page; reads a provider checkout (`--provider `) | Both audits run against `make serve` with `uv run --with playwright python tools/.py`. The provider repository’s code, messages and Markdown link to pages here by `https://captf.io/docs/` URL, some with a heading’s anchor. Nothing checks those links from this side, so keep page paths and headings stable, and search the provider for a page’s URL before moving it. Preview the site with `make serve`; it serves on [http://127.0.0.1:8001/]() with live reload. Set `ADDR` (on the command line, or in an uncommitted `local.mk`) to listen on another address. # Updating the Website captf.io is one site, built from one source by [Zensical](): the landing page at `/`, the documentation at `/docs/` and the blog at `/blog/`. This page covers the parts that are not documentation pages: the landing page, blog posts, the navigation, the header and footer, the generated files, and the theme. For the change workflow and the checks, see [Contributing to the Docs](). The site has one look, dark only: there is no light theme or theme switch. ## The landing page The landing page is `docs/index.md`, rendered by the `overrides/home.html` template with `overrides/assets/stylesheets/landing.css` and `overrides/assets/javascripts/landing.js`. The Markdown file has no body: - the page’s copy (hero, sections, questions) is in `home.html`. It is mirrored word for word from the organization profile in [captf-io/.github](), so change the two together; - the feature grid is data, in the `features:` list of `docs/index.md`’s front matter, in display order. Each entry has a `title`, a `body` (which may use ``), an `href` (a page of this site, or an external URL), the link text `more`, and an `icon`, one of the line icons `home.html` defines: `globe`, `module`, `cluster`, `lock`, `pulse`, `shield`, `checks`, `chart`, `layers`, `link`, `inbox`, `scale`, `home`, `job`, `book`, `bolt`, `github`, `terraform`, `opentofu` or `kubernetes`; - the `anchors:` list in the same front matter is what the navigation drawer shows on this page: one entry per landing section, with the `id` of the section’s heading in `home.html`, a `title`, a `subtitle` and a Lucide `icon`. Add an entry when you add a section to `home.html`. Every claim on the landing page must be one the documentation backs up. The diagram and social card are `docs/assets/landing/how-it-works.svg` and `og.jpg`, also shared with the organization profile. ## Blog posts A post is a Markdown file in `docs/blog/posts/`, named `-.md`: docs/blog/posts/2026-10-02-reference-cloud-modules.md ```markdown --- date: 2026-10-02 slug: reference-cloud-modules title: Reference modules for five clouds description: "Reference module sets for AWS, Google Cloud, Azure, OCI and OpenStack: what each creates, and their pre-release status." authors: - maintainers categories: - Modules --- # Reference modules for five clouds The opening paragraph, shown on the blog's index page. The rest of the post. ``` - The post’s URL is `/blog///
//`, from `date` and `slug`, so both are required, and neither changes once published. - Everything above `` is the excerpt on the index and in the feeds. The marker is required: the build fails without it. - `authors` names entries in `docs/blog/.authors.yml`. - Use one category, an existing one where it fits: `Documentation` or `Modules`. Each category gets its own index page. - State only what the docs or the project’s history back up, and link to the docs for detail rather than repeating it. After adding or retitling a post, run `make gen`: the navigation drawer lists the newest posts on blog pages from data it writes into `zensical.toml`. The RSS and JSON feeds, the archive and the post’s social card are built automatically. ## Navigation and tabs The `nav` in `zensical.toml` is the whole site’s navigation. Its top-level entries are the documentation’s tabs, one per reader (see [Where things go]()), plus the blog. A group without a page of its own shows as a heading; a group whose first entry is a `README.md` opens on that page. Each top-level section has an entry in `[project.extra.sections]`, keyed by its title: a Lucide `icon`, a `subtitle` of at most 34 characters and, when the title is long, a `short` label. The tab row and the drawer both use it, so a section looks the same in each. > [!WARNING] > > **The tab row is full** > > Nine tabs with icons fill the tab row at the narrowest desktop width, and on smaller windows the row steps down to titles, then short labels, then icons. A tenth tab, or a longer section title, needs the row checked at every size: run `tools/mobile_audit.py`. A page’s own sidebar icon and subtitle come from `tools/nav_meta.json`; see [Add a page](). The navigation drawer, on windows narrower than the sidebar layout, follows where the reader is: the docs sections on docs pages, the blog and its recent posts on blog pages, and the landing page’s sections on the home page. Its top row switches between Home, Docs and Blog. ## Header, footer and banner | Part | Where to change it | | --- | --- | | Site name | `site_name` in `zensical.toml`: the full name, for the browser title, link previews and feeds | | Brand name | `extra.brand`: the short name in the header, the drawer and the footer | | Home, Docs and Blog links | `overrides/partials/header.html` and, for the drawer, `overrides/partials/nav.html` | | Footer columns, tagline, status | `[project.extra.footer]`; a link `href` without `://` is a page of this site | | Social links | `[[project.extra.social]]` | | Pre-release banner | The `announce` block in `overrides/main.html`; readers can dismiss it | | Contract chip at the end of the tab row | `[project.extra.tabs]` | | Hover tooltips for acronyms | `includes/abbreviations.md`, one `*[ABBR]: Expansion` line each | | Logo and favicon | `docs/assets/brand/` | ## Generated files `make gen` rewrites three things. Never edit them by hand: the next run overwrites the change. | Generator | Writes | Reads | | --- | --- | --- | | `tools/gen_redirects.py` | A redirect page at each URL of the first edition of the docs (`/docs/.html`), pointing at the page’s current URL | Every page under `docs/docs/` not listed in its `SKIP` | | `tools/gen_llms.py` | The `llms.txt` sections, between markers in `zensical.toml` | The `nav` and each page’s `description:` | | `tools/gen_blog_nav.py` | The recent posts for the drawer, between markers in `zensical.toml` | The posts’ front matter | Search, the RSS and JSON feeds, the sitemap, the social cards, `llms.txt` itself and a Markdown copy of every page are built with the site. ## The theme The site uses Zensical’s own theme, with templates under `overrides/` replacing some of its parts: | Override | What it changes | | --- | --- | | `main.html` | Page metadata (Open Graph, structured data, feeds), the browser title, the banner | | `home.html` | The landing page | | `partials/header.html` | The header: brand, the Home, Docs and Blog links, search | | `partials/tabs.html`, `partials/tabs-item.html` | The docs tab row, its icons and the contract chip | | `partials/nav.html`, `partials/nav-item.html` | The sidebar and drawer: icons, subtitles, status badges, the drawer’s switch and lists | | `partials/footer.html` | The footer and the previous and next page links | | `partials/source-file.html` | The page’s dates and authors | | `partials/comments.html` | The share links under blog posts | The theme’s look is changed in `overrides/assets/stylesheets/extra.css`, in numbered sections with a comment on each saying what it is for. Colors are variables at its top. > [!WARNING] > > **Overrides are copies** > > Each overridden template except `home.html` starts as a copy of the theme’s own and names the Zensical version it was copied from. When the pinned Zensical version in `pyproject.toml` changes, compare each override with the new theme file and carry the theme’s changes across. Two rules keep the theme working across page changes: - Clicking a link inside the site swaps the page’s content without reloading, and the header stays. So anything in the header that depends on the page (which site link is active, for example) is set from a script on each page change (`overrides/assets/javascripts/sitenav.js`), and its links are absolute URLs. A part of the page is only swapped in if the page being left had it too, so such parts are rendered on every page, empty where unused. - Do not put a `title` attribute on an element: the theme turns it into a tooltip that can stay on screen after a page change. Use `aria-label`. Check a theme change with `make build`, then `tools/mobile_audit.py`, clicking through from the home page as well as loading pages directly. > [!NOTE] > > **See also** > > - [Contributing to the Docs]() for previewing, checking and sending a change. > - [Writing Documentation]() for the house style. # Make Targets The provider repository, `cluster-api-provider-terraform`, drives its build, checks, test environment and releases with `make`. Every tool is a pinned version that `make` installs into `hack/tools/bin` on first use, so you need only Go, `make`, and for some targets `podman`, `python3`, `git` or `jq`. `make` with no target, and `make help`, print the targets with their one-line descriptions, grouped as they are here. For how the targets fit into a contribution, see [Contributing](). ## The everyday loop For most changes, run this before you push: ```sh make lint test verify ``` | Target | What it gives you | | --- | --- | | `make lint` | Style and correctness findings from `golangci-lint` in every module, including the e2e-tagged test code, and from `kube-api-linter` on `api/`. | | `make test` | The unit tests of every Go module, with the race detector. | | `make verify` | Every other consistency check: generated files, manifests, schemas, templates, alert rules and the test-tier guard. | After you change `api/v1alpha1` or a controller’s kubebuilder markers, run `make generate manifests` first. `make fmt` formats Go code before you lint it. To try a change against a real cluster, run `make testenv-up` once, then `make testenv-reload` after each change. See [Test environment](<#test-environment>). > [!WARNING] > > **`make verify` fails on unstaged changes** > > Its `verify-gen` step regenerates code and manifests, then runs `git diff --exit-code` over the whole tree. Stage your work (`git add`) before you run it, or it reports your own edits as a failure. ## Variables Set a variable on the command line, as in `make docker-build IMG=`. Unless the table says otherwise, a variable only affects the targets named in its row. | Variable | Default | Effect | | --- | --- | --- | | `IMG` | `ghcr.io/captf-io/cluster-api-provider-terraform:dev` | The manager and runner image that `docker-build`, `docker-buildx` and `docker-push` build or push. | | `CONTAINER_TOOL` | `podman` | The tool that builds and pushes images. `docker` works for `docker-build` and `docker-push`. `docker-buildx` needs `podman`. | | `PLATFORMS` | `linux/amd64,linux/arm64` | The platforms of the manifest list that `docker-buildx` builds. | | `GO_VERSION` | `1.26` | The Go version used by image builds. | | `VERSION` | from `hack/version.sh` | The version stamped into binaries, and the tag that release targets use. It is always a valid semantic version: `v0.0.0-dev.g` before the first release tag. Release targets require `vX.Y.Z` or `vX.Y.Z-rc.N`. | | `RELEASE_REPO` | `ghcr.io/captf-io/cluster-api-provider-terraform` | The repository of the release image. | | `RELEASE_IMG` | `$(RELEASE_REPO):$(VERSION)` | The image `manifests-release` writes into the components file. `make release` passes the pushed image by digest. | | `RELEASE_DIR` | `out` | Where `manifests-release` writes its files. Release assets go to `out/release`. | | `SKOPEO` | `skopeo` | The `skopeo` binary that `release-image-digest` runs. | | `RUNNER_IMAGE` | `$(IMG)` | The runner image `make run` gives the manager. | | `WEBHOOK_CERT_DIR` | `bin/dev-webhook-certs` | Where `make run` keeps its self-signed webhook certificate. | | `ARGS` | | Extra flags that `make run` passes to the manager. | | `GOTESTSUM_FORMAT` | `pkgname` | The `gotestsum` output format of `test-cover`. CI sets `github-actions`. | | `TESTENV_NAME` | `captf-test-dev` | The cluster name for the `testenv-*` targets. It must start with `captf-test-`. | | `TESTENV_ENGINE` | auto-detected | `podman` or `docker`, for the `testenv-*` and `e2e-*` targets. Auto-detection prefers `podman`. | | `TESTENV_WORKERS` | `0` | The number of kind worker nodes for `testenv-up`, from 0 to 5. | | `CAPTF_TESTENV_REUSE` | off | Set to `1` to make `testenv-up` reuse a matching cluster. | | `TESTENV_ALL` | off | Set to `1` to make `testenv-down` delete every `captf-test-*` cluster. | | `CAPTF_E2E_CLUSTER` | `captf-test-e2e` | The cluster the `e2e-*` targets use. | | `CAPTF_E2E_REUSE` | off | Set to `1` to make `e2e-foundation` reuse an existing cluster. | | `CAPTF_E2E_TEARDOWN` | off | Set to `1` to make `e2e-foundation` delete the cluster after a passing run. | | `CAPTF_E2E_STABILITY` | `2m` | The length of the foundation suite’s stability window. | | `CAPTF_E2E_WORKERS` | `0` | The number of kind worker nodes for `e2e-foundation`. | | `CAPTF_E2E_GREENLIGHT_MAX_AGE` | `24h` | How old the green light may be before `e2e-noop` rejects it. | | `CAPTF_E2E_NOOP_BAD_DIGEST` | off | Set to `1` to run `e2e-noop`’s negative check: one machine gets a nonexistent image digest, and the suite must fail on the image pull. | An empty value takes the default in the table. ## General | Target | Description | | --- | --- | | `help` | Print every target with its description, grouped. This is the default target. | ## Development | Target | Description | | --- | --- | | `generate` | Regenerate the deepcopy code for `api/` with `controller-gen`. Run it after you change an API type. | | `manifests` | Regenerate the CRD, RBAC and webhook manifests into `config/`. Run it after you change API types or kubebuilder markers. | | `fmt` | Format Go code with `gofmt -s` and `goimports`. | | `vet` | Run `go vet` in every Go module, then on the `test` module’s e2e-tagged code. | | `lint` | Run `golangci-lint` in every module and on the `test` module’s e2e-tagged code, then `lint-api`. | | `lint-api` | Run `kube-api-linter` on the `api/` module. | | `lint-fix` | Run `lint` with the auto-fixers of both linters on. | ## Test | Target | Description | | --- | --- | | `test` | Run the unit tests of every Go module with `-race -count=1`. It never compiles the e2e-tagged code. | | `test-cover` | Run the unit tests with `-race` and coverage through `gotestsum`, and write a profile and a JUnit report per module into `bin/`. CI runs this instead of `test`. | | `cover-check` | Check each package’s coverage in the `bin/cover-*.out` profiles against its floor in `hack/coverage-floors.txt`. Run it after `test-cover`. | See [Testing]() for what the tests cover, the coverage floors, and how to run one package or one test. ## Test environment The `testenv-*` targets manage a kind cluster on `podman` (or `docker`) with cert-manager, Cluster API and CAPTF built from your working tree. Use it to try a change against a real API server, without a cloud. Each target runs one operation of the e2e-tagged package `test/env/lifecycle`, with the pinned `clusterctl` and `kustomize`. Everything for a cluster `` lands in `bin/testenv//`: its `kubeconfig`, an `env.sh` that exports `KUBECONFIG`, a `state.json` and the diagnostics bundles. | Target | Description | | --- | --- | | `testenv-up` | Build the manager image, create the kind cluster, install the providers and wait until they are ready. About 6 minutes cold. | | `testenv-reload` | Rebuild the manager image from the tree, load it into the nodes and roll the manager Deployment over to it. Run it after each change to the manager or runner. | | `testenv-status` | Show whether the cluster exists, its nodes, the pods that are not ready and a summary of `state.json`. | | `testenv-logs` | Collect pod logs, events, CAPI and CAPTF objects and node logs into `bin/testenv//artifacts//`. Secrets are not collected. | | `testenv-down` | Delete the cluster. The `artifacts/` directory is kept. | ```sh make testenv-up source bin/testenv/captf-test-dev/env.sh kubectl get pods -A make testenv-reload make testenv-down ``` Only cluster names that start with `captf-test-` are accepted, and the targets never read `~/.kube/config`. `testenv-up` refuses an existing cluster unless `CAPTF_TESTENV_REUSE=1` is set. For example, to run a cluster with one worker node next to the default one: ```sh make testenv-up TESTENV_NAME=captf-test-pool TESTENV_WORKERS=1 ``` ## End-to-end The `e2e-*` targets run the opt-in e2e suites under `test/e2e/`. They need `podman` or `docker`. `make test` never compiles them, and CI does not run them. See [Testing](). | Target | Description | | --- | --- | | `e2e-foundation` | Build the manager from your tree and the `captf-test-e2e` cluster, check it stage by stage, and write a green light for it. Run it first. It keeps the cluster after a passing run unless `CAPTF_E2E_TEARDOWN=1`. | | `e2e-noop` | Run the no-op data-flow suite on the green-lit cluster: Cluster API objects drive the published no-op modules, and the suite checks that data flows from the spec to the module and back and that deletion cleans up. It fails unless `e2e-foundation` passed recently. | | `e2e-down` | Delete the e2e cluster, `captf-test-e2e` or `CAPTF_E2E_CLUSTER`. The artifacts are kept. | ## Build | Target | Description | | --- | --- | | `build` | Build every `cmd/*` binary into `bin/`. | | `manager` | Build the static manager into `bin/manager`, and check that it is statically linked. | | `runner` | Build the static Job runner into `bin/runner`, and check that it is statically linked. | | `run` | Build the manager and run it out of cluster against your current `kubeconfig`. It turns leader election off and makes a self-signed webhook certificate if needed. | | `docker-build` | Build the manager and runner image `$(IMG)` for the host platform. | | `docker-buildx` | Build a multi-architecture manifest list `$(IMG)` for `$(PLATFORMS)`, with `podman`. | | `docker-push` | Push `$(IMG)`. A manifest list from `docker-buildx` is pushed with all its images. | `make run` points the manager at `$(RUNNER_IMAGE)`, which is `$(IMG)` by default. Build the image first when a Job needs a runner you changed: ```sh make docker-build IMG= make run RUNNER_IMAGE= ARGS="" ``` - `` is an image reference the cluster can pull, for example `localhost/captf:dev` on a cluster that shares your image store. - `` are extra manager flags. Leave out `ARGS` if you have none. ## Verify `make verify` runs every check below. Each one also runs alone, which is faster when only one has failed. Several need `python3`, `git` or `podman`. | Target | Description | | --- | --- | | `verify` | Run all the checks in this table. | | `verify-godoc` | Check that every declaration, parameter, return value and package has a doc comment. | | `verify-gen` | Run `generate` and `manifests`, then fail if `git diff` shows a change. It needs `git`. | | `verify-modules` | Check `go.work` and `go.mod` pins, and that no module uses `replace`. | | `verify-schemas` | Validate the contract JSON Schemas, their examples and the golden inputs and outputs. It needs `python3` with `jsonschema`. | | `verify-components` | Check the clusterctl components built from `config/default`. | | `verify-metadata` | Validate `metadata.yaml`, and check that `releaseSeries` only grows compared with the previous tag. | | `verify-version` | Check that `hack/version.sh` prints a valid semantic version for every checkout state. | | `verify-templates` | Render `templates/` with the pinned `clusterctl`. | | `verify-local-repository` | Generate the provider and both flavors from a clusterctl local repository of the release assets, offline and without a cluster. | | `verify-test-tiers` | Check that e2e code carries the `e2e` build tag and lives only in `test/e2e/` and `test/env/lifecycle/`. | | `check-licenses` | Check that no MPL-2.0 dependency of `tfcapi-lint` applies Exhibit B. | | `promtool-check` | Check the alert rules with `promtool`, and build the Prometheus component on `config/default`. | | `promtool-test` | Unit-test the alert rules. | `verify` does not run `cover-check`, which needs the profiles `test-cover` writes. CI runs the two as separate steps. ## Release These targets build and publish a release. They need a clean tree with `HEAD` tagged `$(VERSION)`. A tag never moves, so a bad release candidate gets a new one. [Releasing]() has the full procedure. Publishing is for maintainers. | Target | Description | | --- | --- | | `release-preflight` | Check that the tree is clean, `HEAD` carries the tag `$(VERSION)`, and `metadata.yaml` only grows. | | `release` | Run the preflight, build and push the manager image, and build every asset into `out/release` with the image pinned by digest. Set `VERSION=vX.Y.Z`. | | `release-image-digest` | Print the registry digest of `$(RELEASE_REPO):$(VERSION)`. It needs `skopeo`. | | `manifests-release` | Build `out/infrastructure-components.yaml` for `$(RELEASE_IMG)`, and copy `metadata.yaml` and `templates/*.yaml` next to it. | | `release-assets` | Build every release asset for `$(VERSION)` into `out/release`. It needs the git tag. | | `release-notes` | Write `out/release/notes.md` from the commits since the previous tag. | | `release-github` | Create the GitHub release for `$(VERSION)` from `out/release`. This publishes: run it as the maintainer only. | | `release-lint-snapshot` | Build the `tfcapi-lint` release assets and checksums into `dist/` as a GoReleaser snapshot, to try a release without a tag. | | `release-lint-binaries` | An alias of `release-lint-snapshot`. | | `release-lint` | Build the `tfcapi-lint` release assets from the current git tag. | ## Tools | Target | Description | | --- | --- | | `tools` | Install every pinned tool into `hack/tools/bin`. Run it once, and again after a version bump. | ## Cleanup | Target | Description | | --- | --- | | `clean` | Remove the build output in `bin/` and `dist/`. This also removes `bin/testenv/`, so run `testenv-down` first if you want the cluster gone. | # Third-party licenses CAPTF is Apache-2.0. The modules below are linked into `tfcapi-lint`, and this file lists those whose license is not Apache-2.0. The full dependency set is in `go.mod`; a license check over the full set is not yet automated. | Module | Version | License | Used for | | --- | --- | --- | --- | | `github.com/hashicorp/terraform-config-inspect` | `v0.0.0-20260904064934-75d64de68c31` | MPL-2.0 | `tfcapi-lint` reads a module’s variables and outputs without running Terraform | | `github.com/hashicorp/hcl/v2` | `v2.20.1` | MPL-2.0 | the second lint pass: `backend`/`cloud` blocks, provider attributes, `.tofu` files, expressions | | `github.com/hashicorp/hcl` | `v0.0.0-20170504190234-a4b07c25de5f` | MPL-2.0 | legacy HCL parser required by terraform-config-inspect | | `github.com/zclconf/go-cty` | `v1.14.4` | MIT | HCL’s type and value system (the lint pass reads literal values with it) | | `github.com/mitchellh/go-wordwrap` | `v1.0.0` | MIT | diagnostic formatting in HCL | ## MPL-2.0 MPL-2.0 is file-level copyleft. An Apache-2.0 binary may link unmodified MPL-2.0 code (MPL-2.0 section 3.3, “Distribution of a Larger Work”), provided the MPL-covered source stays available: these modules are published, and CAPTF does not modify them. The combination would not be permitted for a file carrying Exhibit B, the “Incompatible With Secondary Licenses” notice. `hack/check-licenses.sh` (part of `make verify`) checks the three MPL modules at their pinned versions: the phrase must appear only in their `LICENSE` files, as the license’s own definitions (sections 1.5, 3.3 and 10.4), and in no source file. It last passed on 2026-09-26. This remains subject to license review before a release. # Blog # Reference modules for five clouds The CAPTF project now maintains reference modules for five clouds: AWS, Google Cloud, Azure, Oracle Cloud Infrastructure (OCI) and OpenStack. Each set implements the `v1alpha1` module contract and ships as module images you reference from a `TerraformCluster`, `TerraformMachineTemplate` or `TerraformMachinePool`. Use them as they are, or fork them as the starting point for your own. ## The sets | Cloud | Repository | Roles | | --- | --- | --- | | [AWS]() | [`captf-io/aws-modules`]() | cluster, machine, machinepool | | [Google Cloud]() | [`captf-io/gcp-modules`]() | cluster, machine, machinepool | | [Azure]() | [`captf-io/azure-modules`]() | cluster, machine, machinepool | | [OCI]() | [`captf-io/oci-modules`]() | cluster, machine, machinepool | | [OpenStack]() | [`captf-io/openstack-modules`]() | cluster, machine | All five follow one set of conventions, so they behave the same way wherever the cloud allows it: an internal API endpoint by default, the same traffic rules, node identities, bootstrap delivery and health outputs. The [Shared Behavior]() page describes them once; each cloud’s pages describe only what differs. ## Status The modules are pre-release. They pass static analysis, mocked unit tests on Terraform and OpenTofu, and image smoke tests, but none has yet been applied to a real cloud. Each repository’s `DESIGN.md` lists the facts the first real apply must confirm. Read it before you rely on a module, and pin a release tag. Start with the [Cloud Modules overview](). # The CAPTF book is online The documentation for Cluster API Provider Terraform now has a home of its own at [captf.io/docs](). It covers what CAPTF is and how it works, how to run clusters with it, how to write a module for it, and how to operate the manager. ## What is in it - **Getting started.** A [Quick Start]() that installs the provider and brings up a cluster with the no-op modules, on any Kubernetes cluster and with no cloud account, and a tutorial that takes you through [your first module](). - **Concepts.** The [architecture](), the kinds, Terraform state, secrets, approvals, jobs and deletion, each explained from the controller’s side. - **Guides.** How to provide credentials, pass module variables, use ClusterClass templates, run machine pools, and handle drift and plan approval. - **The module contract.** The [`v1alpha1` contract]() every module is written to, with its changelog, and integration guides for the kubeadm and RKE2 control planes. - **Operations and runbooks.** Installation, configuration, observability, upgrades and recovery, and [runbooks]() for failing Jobs, stuck destroys, stale locks and more. - **Reference.** The API, conditions, events, metrics, alerts and flags, generated from the provider’s source so they stay in step with it. ## Status Every kind and the module contract are `v1alpha1` and pre-release: the contract is frozen for implementation but may still change. CAPTF has no end-to-end tests yet. The [Known Limitations]() page lists everything it does not do.