Skip to content

Runbook: state that will not read

StateReadable reports whether the controller could read a TerraformCluster, TerraformMachine or TerraformMachinePool’s state this reconcile. While it is False or Unknown, no plan, apply, drift check or refresh Job runs for the object, and Ready follows it down. This runbook covers every reason StateReadable (and its True counterpart) can carry, what each means, and what to do. For the CAPTFStateUnreadable alert itself, see Observability.

Replace <ns>, <kind> and <name> below with the object’s namespace, kind (terraformcluster, terraformmachine or terraformmachinepool) and name.

Find the reason

kubectl get <kind> -n <ns> <name> \
  -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}: {.message}{"\n"}{end}'

Also check for a Warning event: a reason that turns False emits one, named either StateLost, StateLocked or, for the three read errors below, StateUnreadable.

kubectl events --for <kind>/<name> -n <ns>

Deleting while state is unreadable

If the object ever applied, deletion is held while its state is missing or unreadable: the destroy cannot run safely without the state, and dropping the finalizer would leave the infrastructure running with no record of it. “Ever applied” means the object was provisioned, a state backup exists, the durable inputs Secret pins an image digest, or that Secret carries the captf.io/applied: "true" marker (see the applied marker). The controller keeps the finalizer, StateReadable stays False with the reason StateLost, StateEncrypted, StateCorrupt or StateInconsistent, a Warning event is emitted, and every state backup is preserved. An object that never applied has nothing to lose and still deletes immediately.

Two ways out:

  • Restore, then destroy. Set captf.io/restore-state to a serial from status.stateBackups, as in the state restore runbook. On a held delete the restore runs first, and the normal destroy follows once the state reads again.
  • Abandon the infrastructure. Set captf.io/abandon-infrastructure to the object’s metadata.uid. The controller checks it before restore and destroy, so it works even while a restore is pending. It removes the finalizer without running a destroy and emits a Warning event, InfrastructureAbandoned, naming the cause.

Abandoning leaves the infrastructure running and untracked

The infrastructure keeps running and is no longer tracked by anything in Kubernetes; delete it through the cloud provider. The state backups are garbage-collected with the object, so copy out anything you want to keep first (see the stuck destroy runbook).

uid=$(kubectl get <kind> -n <ns> <name> -o jsonpath='{.metadata.uid}')
kubectl annotate <kind> -n <ns> <name> captf.io/abandon-infrastructure="$uid"

The value must equal the UID exactly; the controller ignores any other value, so a stale annotation copied from another object does nothing.

A held delete is one of four cases the annotation releases; the stuck destroy runbook lists all of them. An object whose state reads and whose destroy can start is destroyed as usual, even with the annotation set.

The applied marker

The durable inputs Secret captf-inputs-<kindshort>-<name> carries captf.io/applied: "true" once an apply succeeded or a state with an inputs hash was read. Unlike status, it survives clusterctl move, so after a move an object whose state is lost still holds with StateLost, deleting or not, instead of being deleted without a destroy or applied again from scratch. One limit: deleting the namespace removes that Secret along with the state and backups, so a moved object in a deleted namespace cannot be told apart from one that never applied.

StateRead

The healthy, True reason: the state was read this reconcile. Nothing to do.

StateNotFound

Unknown. No state Secret exists yet, because the object has never completed an apply, or (for a mutable kind) it was deleted along with the state on a previous destroy. This is not an error: a new object reports it until its first successful apply, and it never fires CAPTFStateUnreadable. If an object that used to be provisioned reports this instead of StateRead, its state Secret is gone; see StateLost below, which is the reason an object that applied before, with a missing state Secret, carries instead.

StateLost

False. Either the state Secret of an object that applied before is missing (it need not still be provisioned), or its state has no recorded inputs hash although the object is provisioned (only possible for TerraformMachine, which is immutable and so never rebuilds inputs to re-apply from). Either way, applying again is not safe: for the missing-Secret case, a second apply would create a second set of resources next to the live ones instead of managing the ones the lost state recorded; for the missing-inputs-hash case, there is nothing to re-apply. No Job runs, and the condition does not clear on its own.

Fix: restore the object’s newest usable state backup with the captf.io/restore-state annotation. See the state restore runbook, including how to list the backups in status.stateBackups. If no backup exists, StateLost never clears; the object’s resources still exist in the cloud but nothing in Kubernetes can manage them until a state is restored, or reconstructed out of band with manual recovery. The Total State Loss and Import runbook gives the whole flow, including recreating the object with import blocks.

StateLocked

False. The state’s lock Lease is held by something other than this object’s own runner — for example a workstation running terraform state rm, or a Job that died without releasing it. The condition’s message names the holder, its operation and when it took the lock; every Job for the object waits lockTimeoutSeconds (spec.jobs.lockTimeoutSeconds, defaulting to 300 seconds — see Job tuning) for it and then fails.

The controller already force-unlocks a lock whose holder pod is gone, the next time it starts a Job for the object; most StateLocked conditions clear on their own. See the stale-lock runbook for how that detection works, how to tell a live holder from a stale one by hand, and manual force-unlock.

StateEncrypted

False. The state Secret carries OpenTofu’s client-side state encryption envelope. CAPTF v1 cannot read an encrypted state at all — not the outputs, not the resource count, nothing — so this condition never clears on its own and the object gets no drift checks, health checks or further applies. Because an encrypted state fails the reader before it is ever parsed, the manager never took a backup of it either: there is no CAPTF-side backup to restore.

Fix: disable client-side state encryption for this object’s module (drop the encryption configuration so future writes are plain state), and use a workstation that holds the decryption key to decrypt the current state and push the plaintext back with terraform state push (or tofu state push) against the same kubernetes backend configuration — the backend config, not the state format, is what CAPTF’s reader needs to match; see manual recovery for the exact secret_suffix, namespace and labels the object’s backend uses. Once the pushed state is unencrypted, the next reconcile reads it normally.

StateCorrupt

False. The reader could not treat the Secret’s payload as gzip-compressed Terraform state JSON: the gzip or JSON decoding failed, the decompressed size exceeded the reader’s cap, or (for a multi-Secret chunked state) there were more chunks than the reader accepts. An unsupported state file version is reported the same way. The condition’s message names which of these it was. See Chunking and size caps for the reader’s exact limits.

A corrupt state is never backed up by the manager (backups copy only a state the reader could parse), so a backup taken before the corruption is your most recent recoverable copy. Fix: restore the newest backup from before the corruption with the state restore runbook. If nothing wrote a backup before the state became corrupt, there is no CAPTF-side recovery: rebuild the state out of band with manual recovery, using the resource list from the last known-good backup or the cloud provider’s own inventory as your guide.

StateInconsistent

False. The set of state Secrets for the object’s suffix does not add up to one complete, contiguous state: a chunk is missing or duplicated, a Secret in the set has an unexpected name, or the base Secret has no tfstate data key. The condition’s message names which of these it was.

This can be transient: a runner Job writes a multi-chunk state one Secret at a time, so a reconcile that reads mid-write sees an incomplete set and the next reconcile, after the Job finishes, usually reads a complete one. If it persists past the Job finishing, something outside CAPTF edited or deleted one of the chunk Secrets by hand. As with StateCorrupt, an inconsistent state is never backed up, so restore the newest backup from before the inconsistency appeared with the state restore runbook; the reader’s own limits are at Chunking and size caps.

Confirm it worked

kubectl get <kind> -n <ns> <name> \
  -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}{"\n"}{end}'

Expect True/StateRead. status.observedStateSerial moves to the restored or rebuilt state’s serial, and the object’s next drift check or apply proceeds from it.