Common
Applies to every role. Inputs here are injected by the controller; modules MUST declare them (they may ignore them). Outputs here MUST be declared by every role.
Inputs
| Name | Type | Required | Value source |
|---|---|---|---|
captf_contract | string | yes | Literal "v1alpha1"; lets a module assert the contract it was generated for |
captf_cluster | object (below) | yes | The owning CAPI Cluster |
captf_object | object (below) | yes | The Terraform* object being reconciled |
captf_cluster_outputs | any | machine, machinepool only | The cluster role’s exports output, read from the cluster state (below) |
captf_tags | map(string) | yes | Fixed tag keys the controller always sets (below) |
variable "captf_contract" { type = string }
variable "captf_cluster" {
type = object({
name = string
namespace = string
})
}
variable "captf_object" {
type = object({
kind = string # TerraformCluster | TerraformMachine | TerraformMachinePool
name = string
namespace = string
})
}
variable "captf_cluster_outputs" {
type = any
default = null
}
variable "captf_tags" { type = map(string) }
captf_cluster_outputs
The value of the cluster role’s exports output
(cluster.md), read from the cluster state. Nothing
else from the cluster state is injected.
- Not passed to the cluster role. Its generated root declares no such
variable and its tfvars carry no key, so a cluster module either leaves it
undeclared or declares it with
default = null, as the skeleton below does. A declaration without a default failsvalidate. {}when the TerraformCluster is externally managed (thecluster.x-k8s.io/managed-byannotation; the provider then never runs the cluster module, so there is no cluster state to read).- A machine/pool whose owning Cluster’s
spec.infrastructureRefis not kindTerraformClusteris never rendered at all: it setsDependenciesReady=False/ClusterNotTerraformand does nothing, since there is no cluster state to readexportsfrom. - This is how a machine/pool module gets network ids, security groups,
subnet ids and similar cluster-wide values. The cluster module puts them
in its
exportsoutput; the controller reads the cluster state and injects exactly that value.exportsmay be markedsensitive: the generated root re-exports every contract output with"sensitive": trueregardless, so a sensitiveexportsnever failsplan, and sensitivity has no effect on the controller — the state Secret still stores the value in cleartext; only CLI display is affected (seeREADME.md“Type conventions”). Machine/pool apply is gated on the cluster being provisioned andexportsreadable, so the value is complete by then. - Because
captf_cluster_outputsisany, a change in its content is not treated as a spec change for machines (immutable) but is for pools and clusters (a value change triggers a re-apply). See each role’s Lifecycle section.
captf_tags
Fixed keys the controller always sets: captf.io/cluster (Cluster name),
captf.io/namespace, captf.io/kind (the Terraform* kind),
captf.io/name (the Terraform* object name), captf.io/managed-by =
captf, and captf.io/template.
captf.io/templateis the value of the object’scluster.x-k8s.io/cloned-from-nameannotation — the template it was cloned from byexternal.GenerateTemplate(util.go,common_types.go) — when present, else"". The key is always present: it is read from the Terraform* object’s own metadata, so an object created directly gets"". It is hashed like the other keys even though annotations in general are not:GenerateTemplatewrites it once at clone time, but for a TerraformCluster or TerraformMachinePool managed by topology, a ClusterClass rebase onto a differently named template rewrites the annotation through SSA, which re-applies that object once; for immutable machines it is pinned at first apply.- Modules MUST apply these tags to every cloud resource they create that
supports tags or labels, so resources can be attributed and
garbage-collected out of band. This follows the CAPI provider best
practices: a tagging/labeling mechanism for identifying cloud objects
“MUST always be provided”
(
best-practices.md;security-guidelines.md, housekeeping).
What is hashed
The controller re-applies mutable kinds when the hash of the rendered
inputs changes. Rendered inputs are hashed inputs: the hash covers a
canonical struct of exactly the spec-derived inputs the module receives,
and nothing is rendered outside it (the one special case is pool replicas
while autoscaling.enabled, rendered as the module’s own observed desired
capacity; see machinepool.md).
captf_cluster and captf_object therefore carry only kind, name and
namespace:
- no
uid— aclusterctl movere-creates objects with new UIDs, and hashing them would re-apply every cluster and pool after a move; - no
generation— it changes whenever the controller itself writesspec.providerID,providerIDListorcontrolPlaneEndpoint; - no
annotations— they toggle around every Job (clusterctl.cluster.x-k8s.io/block-move) or are edited by GitOps tools; - no
labels— a metadata edit must never reconcile live infrastructure.
Cloud tags come only from captf_tags, whose keys are fixed. Drift for
mutable kinds renders the current hashed set; immutable machines render
drift and destroy from the durable inputs Secret below.
Durable inputs
The rendered inputs (including bootstrap_data) are stored in a Secret
owned by the Terraform* object (captf-inputs-<kindshort>-<name>), so
drift and destroy of immutable machines re-feed exactly the values used at
apply, and the Secret moves with the object. The resolved image digest
(captf.io/image-digest, see
image-contract.md
“Versioning and pinning”) and the identity are pinned there too, so an
immutable machine’s drift and destroy always run the exact image that
created it.
User variables
Everything a module is parameterized by beyond the contract inputs
(instance type, disk size, subnets, a database password) is an ordinary
Terraform variable of the module. The object’s owner sets it through
spec.variables (inline JSON) or spec.variablesFrom (ConfigMaps and
Secrets labeled captf.io/variables=true) on the TerraformCluster,
TerraformMachine or their templates; see
Module Variables for the operator side.
For module authors:
- Declare each one as an ordinary variable, with a type and a
default. The generated root passes a user variable as a named argument ofmodule "role", declared in the root without a type, so your declaration’s type converts the value: anumbervariable accepts40inline and"40"from a ConfigMap. A variable nobody sets gets itsdefault; without one, every object that does not set it failsplanon the missing argument.tfcapi-lintreports such a variable asinput/user-variable-default(a warning: a module may deliberately require it). - Mark secrets
sensitive = truein the module. The root declares a variablesensitiveonly when its value came from a Secret; your declaration makes the value sensitive inside the module whatever its source. A value that arrives sensitive stays sensitive in everything derived from it: an output built from it must be declaredsensitive = true(orplanfails with “Output refers to sensitive values”), and it cannot drivefor_each. Either way the value is stored in cleartext in the inputs Secrets and in the state, like every input (see Module Variables); credentials the module runs with come from the identity, never from a variable. - Names. User variables are Terraform identifiers
(
^[a-zA-Z_][a-zA-Z0-9_-]*$). Reserved, and rejected before anything is rendered: everycaptf_name, every contract input of the module’s role (machine_nameon a machine,control_plane_initializedon a cluster, and so on) and the module meta-argumentssource,version,providers,count,for_each,depends_on,lifecycle,locals. Do not rely on a user variable to carry a contract input’s value. - Unknown names fail the apply. A variable the object sets and the
module does not declare is rejected by Terraform/OpenTofu (
Unsupported argument; for the JSON root,Extraneous JSON object property). The controller does not check names against the image. - Hashing. User variables are part of the inputs hash only when the
object sets some, so modules and objects without them hash exactly as
before. For a TerraformCluster a changed variable re-applies (guarded by
the destructive-plan approval, see
cluster.md); a TerraformMachine reads its variables until it is provisioned, and its later runs use the ones pinned in its durable inputs Secret, like every other input. - The role input schemas (
schemas/*-inputs.json) describe the contract inputs only; the renderedterraform.tfvars.jsoncarries the user variables next to them.
Outputs
| Name | Type | Required | Maps to |
|---|---|---|---|
health | object (below) | yes | Ready condition; machine remediation; drift reporting |
output "health" {
value = {
state = "running" # closed enum: pending | running | degraded | stopped | terminated | unknown
healthy = true
message = null # optional string, surfaced in condition message
reasons = [] # optional list(string), machine-readable
}
}
state is a closed enum. Any other string is
OutputsValid=False/OutputsInvalid.
Controller-side semantics. health is written to the
InfrastructureHealthy condition. InfrastructureHealthy is one of the
explicitly listed inputs to the Ready summary; so is Deleting (negative
polarity, so Ready=False while a destroy runs). Drift results, drift-Job
outcomes, Paused and DeletionBlocked are not, so Ready — and therefore
the Cluster/Machine InfrastructureReady mirror that MachineHealthCheck
acts on — reflects infrastructure health, not controller housekeeping.
state | healthy | InfrastructureHealthy | Effect |
|---|---|---|---|
pending | any | False/InstancePending | Provisioning in progress; object not marked provisioned even if other outputs are set. While pending, the controller refreshes (apply -refresh-only) regardless of drift.intervalSeconds, so provisioning never waits for a drift tick: for cluster and machine, 30s after the last refresh, then 1m, 2m, 4m and at most every 5m while the readings stay pending (plus the object’s jitter); for a machinepool, every 30s flat, since a pending pool is also not converged (see machinepool.md “Membership refresh”) |
running | true | True/Healthy | Ready=True once role outputs are satisfied |
running | false | False/InstanceUnhealthy | machine: remediation candidate |
degraded | any | False/InstanceDegraded | machine: remediation candidate |
stopped | any | False/InstanceStopped | machine: remediation candidate |
terminated | any | False/InstanceTerminated | machine: remediation candidate; object stays provisioned and spec.providerID is kept |
unknown | any | Unknown/HealthUnknown | Ready=Unknown; no remediation |
Once provisioned, the Machine controller treats an empty InfraMachine
spec.providerID as “waiting” and logs it; it is not a hard error, but
clearing it would stall the Machine
(machine_controller_phases.go
reconcileInfrastructure), so a terminated reading keeps providerID
rather than clearing it.
healthis re-read on every refresh (apply -refresh-only, which updates state and root output values to match remote objects; see OpenTofu docs:cli/commands/plan/“Planning Modes”): the pending refresh loop above, every drift tick, and for pools the membership refresh (seemachinepool.md). For immutable machines that is the only way it changes after apply. Modules that cannot observe health return{state="running", healthy=true}(the no-op module does this) orunknown.- Out-of-band termination. If, after provisioning, a refresh makes
provider_idcome backnull(the resource vanished from state) or the refresh itself fails because the instance is gone, the controller treats that asterminated:InfrastructureHealthy=False/InstanceTerminated,spec.providerIDkept. Modules SHOULD still reportterminatedexplicitly where the cloud API exposes it, because detection through a vanished resource is slower (drift interval plus a Job) and CAPI’s Node-based checks will usually fire first. - A group with zero members (pool role,
replicas == 0) reports{state="running", healthy=true}. - Only the machine role triggers remediation. Cluster/pool
healthfeeds conditions only.
Role-specific inputs
kubernetes_version is a role input, not a common one: for machines/pools
it comes from Machine.spec.version / MachinePool.spec.template.spec.version
(authoritative for nodes); for the cluster role it comes from
Cluster.spec.topology.version (null without ClusterClass; see
cluster.md). cluster_network,
control_plane_initialized and the non-module control_plane_endpoint
(including one set later by a control-plane provider) are cluster-role
inputs (see cluster.md); machine/pool modules get
endpoint-derived values through exports.