Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Runbook: state that will not read

StateReadable reports whether the controller could read a TerraformCluster, TerraformMachine or TerraformMachinePool’s state this reconcile. While it is False or Unknown, no plan, apply, drift check or refresh Job runs for the object, and Ready follows it down. This runbook covers every reason StateReadable (and its True counterpart) can carry, what each means, and what to do. For the CAPTFStateUnreadable alert itself, see Observability.

Replace <ns>, <kind> and <name> below with the object’s namespace, kind (terraformcluster, terraformmachine or terraformmachinepool) and name.

Find the reason

kubectl get <kind> -n <ns> <name> \
  -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}: {.message}{"\n"}{end}'

Also check for a Warning event: a reason that turns False emits one, named either StateLost, StateLocked or, for the three read errors below, StateUnreadable.

kubectl events --for <kind>/<name> -n <ns>

StateRead

The healthy, True reason: the state was read this reconcile. Nothing to do.

StateNotFound

Unknown. No state Secret exists yet, because the object has never completed an apply, or (for a mutable kind) it was deleted along with the state on a previous destroy. This is not an error: a new object reports it until its first successful apply, and it never fires CAPTFStateUnreadable. If an object that used to be provisioned reports this instead of StateRead, its state Secret is gone; see StateLost below, which is the reason a provisioned object with a missing state Secret carries instead.

StateLost

False. Either the state Secret of a provisioned object is missing, or its state has no recorded inputs hash although the object is provisioned (only possible for TerraformMachine, which is immutable and so never rebuilds inputs to re-apply from). Either way, applying again is not safe: for the missing-Secret case, a second apply would create a second set of resources next to the live ones instead of managing the ones the lost state recorded; for the missing-inputs-hash case, there is nothing to re-apply. No Job runs, and the condition does not clear on its own.

Fix: restore the object’s newest usable state backup with the captf.io/restore-state annotation. See the state restore runbook, including how to list the backups in status.stateBackups. If no backup exists, StateLost never clears; the object’s resources still exist in the cloud but nothing in Kubernetes can manage them until a state is restored, or reconstructed out of band with manual recovery.

StateLocked

False. The state’s lock Lease is held by something other than this object’s own runner — for example a workstation running terraform state rm, or a Job that died without releasing it. The condition’s message names the holder, its operation and when it took the lock; every Job for the object waits lockTimeoutSeconds (spec.jobs.lockTimeoutSeconds, defaulting to 300 seconds — see Job tuning) for it and then fails.

The controller already force-unlocks a lock whose holder pod is gone, the next time it starts a Job for the object; most StateLocked conditions clear on their own. See the stale-lock runbook for how that detection works, how to tell a live holder from a stale one by hand, and manual force-unlock.

StateEncrypted

False. The state Secret carries OpenTofu’s client-side state encryption envelope. CAPTF v1 cannot read an encrypted state at all — not the outputs, not the resource count, nothing — so this condition never clears on its own and the object gets no drift checks, health checks or further applies. Because an encrypted state fails the reader before it is ever parsed, the manager never took a backup of it either: there is no CAPTF-side backup to restore.

Fix: disable client-side state encryption for this object’s module (drop the encryption configuration so future writes are plain state), and use a workstation that holds the decryption key to decrypt the current state and push the plaintext back with terraform state push (or tofu state push) against the same kubernetes backend configuration — the backend config, not the state format, is what CAPTF’s reader needs to match; see manual recovery for the exact secret_suffix, namespace and labels the object’s backend uses. Once the pushed state is unencrypted, the next reconcile reads it normally.

StateCorrupt

False. The reader could not treat the Secret’s payload as gzip-compressed Terraform state JSON: the gzip or JSON decoding failed, the decompressed size exceeded the reader’s cap, or (for a multi-Secret chunked state) there were more chunks than the reader accepts. An unsupported state file version is reported the same way. The condition’s message names which of these it was. See Chunking and size caps for the reader’s exact limits.

A corrupt state is never backed up by the manager (backups copy only a state the reader could parse), so a backup taken before the corruption is your most recent recoverable copy. Fix: restore the newest backup from before the corruption with the state restore runbook. If nothing wrote a backup before the state became corrupt, there is no CAPTF-side recovery: rebuild the state out of band with manual recovery, using the resource list from the last known-good backup or the cloud provider’s own inventory as your guide.

StateInconsistent

False. The set of state Secrets for the object’s suffix does not add up to one complete, contiguous state: a chunk is missing or duplicated, a Secret in the set has an unexpected name, or the base Secret has no tfstate data key. The condition’s message names which of these it was.

This can be transient: a runner Job writes a multi-chunk state one Secret at a time, so a reconcile that reads mid-write sees an incomplete set and the next reconcile, after the Job finishes, usually reads a complete one. If it persists past the Job finishing, something outside CAPTF edited or deleted one of the chunk Secrets by hand. As with StateCorrupt, an inconsistent state is never backed up, so restore the newest backup from before the inconsistency appeared with the state restore runbook; the reader’s own limits are at Chunking and size caps.

Confirm it worked

kubectl get <kind> -n <ns> <name> \
  -o jsonpath='{range .status.conditions[?(@.type=="StateReadable")]}{.status}/{.reason}{"\n"}{end}'

Expect True/StateRead. status.observedStateSerial moves to the restored or rebuilt state’s serial, and the object’s next drift check or apply proceeds from it.

See also