Skip to content

Stuck Destroy

CAPTF has no skip-destroy annotation, only an abandon one (see Abandon instead): if a destroy Job keeps failing, the object stays with ApplyJobSucceeded=False/DestroyFailed (so Ready=False) and the controller retries with backoff forever. This applies the same way to a TerraformCluster, a TerraformMachine and a TerraformMachinePool. This page walks through recovering: back up what the destroy would remove, clean up the cloud resources another way, then remove the object’s finalizer by hand.

Removing the finalizer deletes the only record of your cloud resources

Removing the finalizer garbage-collects the state Secrets, the state backups and the durable inputs Secret through their owner references, which are the only record of the live cloud resources. Back them up, or un-own them, before you do that.

If the destroy is not failing but never starts because the state is missing or unreadable, deletion is held on purpose; see deleting while state is unreadable before removing anything by hand.

A stuck control-plane destroy blocks the control plane

A stuck destroy on a control-plane machine also blocks KubeadmControlPlane/RKE2ControlPlane remediation, scale and upgrade until it is resolved.

Abandon instead

captf.io/abandon-infrastructure releases a deleting TerraformCluster, TerraformMachine or TerraformMachinePool whose destroy cannot run. It covers four cases:

  • The deletion is held because the state is missing or unreadable (StateReadable is False).
  • The last destroy Job failed (status.lastRun is a destroy and ApplyJobSucceeded is False).
  • The destroy cannot be rendered because the durable inputs Secret is gone (ApplyJobSucceeded is False/DestroyFailed).
  • The destroy cannot start because the identity no longer allows the namespace or was deleted (IdentityNotAllowed), or the runner credentials cannot be prepared (the Deleting condition says the destroy waits for its credentials).

Set it to the object’s metadata.uid; any other value is ignored. The controller removes the finalizer without a destroy and records an InfrastructureAbandoned Warning event naming the cause. Whatever the module created keeps running, untracked, and the state backups are deleted with the object, so back up what you need and clean the cloud resources up yourself (steps 2 and 4 below). The controller checks the annotation before restore and destroy, so it also works while a restore is pending.

An object whose state reads and whose destroy can start is destroyed as usual, even with the annotation set: it takes effect once that destroy fails or turns out unable to start. For the commands, see deleting while state is unreadable.

Terminating namespaces

A deleting object needs credentials only when it runs a Job. With the namespace terminating, an object that never applied finishes at once, and a held one can be abandoned. A destroy or restore waiting on credentials shows the Deleting condition message The <op> Job waits for its credentials: … and retries.

Before you begin

  • get, patch and delete access to Secrets and to the stuck object, in its namespace, on the management cluster.
  • Replace <ns>, <name> and <kind> below with the object’s namespace, name and Kind (terraformcluster, terraformmachine or terraformmachinepool; the lowercase singular works with kubectl). Run every command against the management cluster.

1. Find the state suffix

State Secret and backup Secret names are keyed by a suffix the controller derives from the object’s namespace, kind and name. Read it back from the object’s own status rather than recomputing it:

kubectl get <kind> -n <ns> <name> -o jsonpath='{.status.stateSecretSuffix}'

Save the output as <suffix> for the commands below.

2. Back up the state and inputs Secrets

kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix> -o yaml > backup-state.yaml
kubectl get secret -n <ns> -l captf.io/state-backup=true,captf.io/state-backup-suffix=<suffix> -o yaml > backup-state-backups.yaml
kubectl get secret -n <ns> captf-inputs-<kindshort>-<name> -o yaml > backup-inputs.yaml

<kindshort> is c for a TerraformCluster, m for a TerraformMachine, mp for a TerraformMachinePool. The first selector matches the base state Secret and every -part-N chunk of a large Terraform state in one call; the second matches every backup Secret (and its chunks) taken of this object — see Terraform State and the state restore runbook. Both kinds of Secret are owned by the object and are garbage-collected along with it once its finalizer is removed.

backup-inputs.yaml’s captf.io/image-digest annotation (see Annotations, labels and finalizers) names the exact image that ran the last successful apply — you will need it in step 4.

3. Un-own the Secrets, if you want them to survive finalizer removal

kubectl patch takes names, not a label selector, so patch each Secret by name:

for s in $(kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix> -o name); do
  kubectl patch -n <ns> "$s" --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]'
done
for s in $(kubectl get secret -n <ns> -l captf.io/state-backup=true,captf.io/state-backup-suffix=<suffix> -o name); do
  kubectl patch -n <ns> "$s" --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]'
done
kubectl patch secret -n <ns> captf-inputs-<kindshort>-<name> \
  --type=json -p '[{"op":"remove","path":"/metadata/ownerReferences"}]'

Skip this if you would rather rely on step 2’s backup-*.yaml files: with those in hand you don’t need the live Secrets to survive finalizer removal.

4. Clean up the cloud resources

Either clean up out of band (the cloud console, or the module’s own tooling), or run the pinned image’s terraform/tofu binary by hand against the backed-up state and inputs. The image’s own binary lives at /captf/runtime (Image Contract); the Job’s own use of it, including its exact -backend-config flags, is in Job Environment.

docker run --rm --entrypoint /captf/runtime \
  -v "$PWD/root:/captf/work/root" -w /captf/work/root \
  <image>@<digest> init -input=false -no-color \
  -backend-config=secret_suffix=<suffix> -backend-config=namespace=<ns> \
  -backend-config=in_cluster_config=true -backend-config=labels=<labels>

docker run --rm --entrypoint /captf/runtime \
  -v "$PWD/root:/captf/work/root" -w /captf/work/root \
  <image>@<digest> destroy -auto-approve -input=false -no-color \
  -var-file=terraform.tfvars.json
  • <image>@<digest> is the repository from backup-inputs.yaml’s captf.io/image annotation (drop any :tag) plus @ and the digest from its captf.io/image-digest annotation.
  • <labels> must be the same HCL object the Job itself would pass, or init reads an empty state: the backend selects state by this whole map. Copy it from backup-state.yaml’s base Secret .metadata.labels, dropping tfstate, tfstateSecretSuffix and tfstateWorkspace (the backend sets those itself), as {"key"="value",...}; see Annotations, labels and finalizers for what each remaining key means.
  • ./root needs main.tf.json and terraform.tfvars.json from backup-inputs.yaml’s data (the durable inputs Secret’s rendered root module and tfvars), and the identity’s credentials as environment variables (-e <KEY>=<value> for each key of the mirrored credentials Secret, or a file for a file-based one — see Identities and credentials).
  • -backend-config=in_cluster_config=true reads the pod’s own ServiceAccount token, so this only works run from inside the cluster (for example, a debug pod using the captf-runner ServiceAccount). Running it from a workstation needs a kubeconfig-based backend configuration instead.
  • Restoring backup-state.yaml first (recreate the Secrets it holds) and letting the kubernetes backend read that live state also works, and skips reconstructing the backend configuration by hand.

5. Remove the finalizer

Stripping the finalizer skips the controller

No destroy, no cleanup and no event run. Finish steps 2 and 4 first: the cloud resources the object created keep running, untracked, and the state backups go with the object.

kubectl patch <kind> -n <ns> <name> --type=json \
  -p '[{"op":"remove","path":"/metadata/finalizers"}]'

Use the lowercase, plural CRD resource name (for example, terraformmachines) if your kubectl version needs it instead of the Kind. This removes the whole metadata.finalizers array: safe for a Terraform* object, which carries only CAPTF’s own finalizer, but check .metadata.finalizers first if something else may have added one, and remove that entry by index instead.

Confirm it worked

kubectl get <kind> -n <ns> <name>
kubectl get secret -n <ns> -l tfstate=true,tfstateSecretSuffix=<suffix>

The first command reports NotFound once the object is gone. The second shows nothing unless you un-owned the Secrets in step 3, in which case they are exactly what you chose to keep.

The durable inputs Secret is missing

If the object reports ApplyJobSucceeded=False/DestroyFailed with the message “The durable inputs Secret is missing, so destroy cannot be rendered; see https://captf.io/docs/operator-guide/runbooks/stuck-destroy.html”, the durable inputs Secret (captf-inputs-<kindshort>-<name>: the rendered root module, tfvars and pinned image from the last successful apply — see Job Inputs) was deleted or never written.

TerraformMachine is the persistent case: it is immutable and never falls back to re-rendering current inputs for a destroy, since the Machine and its bootstrap Secret a rebuild would need are usually already gone by the time destroy runs, so once its durable Secret is gone the condition never clears on its own. TerraformCluster and TerraformMachinePool (both mutable) fall back to building current inputs instead; they show the same message only while that build is gated (for example, waiting on a dependency that is itself being deleted), and it usually clears once the gate does. The controller retries forever either way; it never invents inputs to destroy with.

To recover:

  • If you have a backup (backup-inputs.yaml from a previous run of step 2 above, or any earlier copy of the durable inputs Secret), recreate it with kubectl apply -f backup-inputs.yaml after removing metadata.uid and metadata.resourceVersion from the YAML: reapplying the exact object restores its owner reference, labels and annotations, including the pinned captf.io/image and captf.io/image-digest. The next reconcile reads it and starts the destroy Job.
  • If you have no backup, the controller cannot destroy the object’s resources. Back up the state (step 2), then clean up out of band (step 4): without the rendered main.tf.json and terraform.tfvars.json, running the module by hand needs reconstructing them, but the state lists every resource the module created. Instead of removing the finalizer by hand, abandon releases the object without a destroy and records the cause. status.source.image and status.source.imageDigest still name the image the last Job ran. Then remove the finalizer (step 5). Removing the finalizer without cleaning up the cloud resources first abandons them: they stay in the cloud, unmanaged, with nothing in Kubernetes recording that they ever existed.