Retries and Backoff¶
CAPTF retries a failed operation forever, with a delay that grows to a cap. Nothing limits the number of attempts; what the failure limit bounds is the delay and the remediation retries. This page defines what counts as a failure, the delay, and the cases that deliberately do not back off.
Counting failures¶
For each operation separately, the controller walks the object’s finished Jobs newest first and counts the failed ones until it meets a success of that same op. A failing drift check therefore never delays an apply.
| Outcome of a finished Job | Counts as a failure |
|---|---|
| Failed: a step failed, the image would not pull, the image broke the contract | Yes |
Killed by activeDeadlineSeconds | Yes, even if the runner reported itself interrupted |
| Stopped from outside: a drain, an eviction, a Job deletion | No (JobInterrupted) |
| An apply the runner blocked before a destructive plan | No (DestructivePlanBlocked) |
| An approved apply whose plan no longer matched | No (PlanChanged) |
| A plan Job that succeeded but whose plan could not be read | Yes: the next plan backs off instead of looping |
| A lease wait | No: no Job existed |
Two details explain the table:
- Interrupted versus the deadline. On a deadline, Kubernetes sends the runner SIGTERM, as it does for a drain, and the runner may report the run as
interrupted. The controller still counts it, because a step that always hangs would otherwise retry at once forever and never reach the cap. Only an interruption that is not a deadline kill is free. - Blocked and plan-changed. Both stopped before changing anything, and both are waiting for a person, not a timer. They count toward no backoff and no retry number;
DecideOpwaits for the approval or for new inputs instead (see Choosing the operation).
The delay¶
The delay after n consecutive failures of an op is
counted from when the newest failed Job finished. Once n reaches the failed-Jobs history limit, it is the 10-minute cap straight away. The limit is spec.jobs.failedJobsHistoryLimit, default 3, and at least 1 (a limit of 0 counts like 1, because pruning always keeps the newest unresolved failure). The timing is compiled in: one minute base, ten minute cap.
| Limit | Delay after failure 1, 2, 3, 4, 5 |
|---|---|
| 0 or 1 | 10m, 10m, 10m, 10m, 10m |
| 3 (default) | 1m, 2m, 10m, 10m, 10m |
| 5 | 1m, 2m, 4m, 8m, 10m |
Pruning keeps at most the limit’s worth of failures, so n cannot exceed it; reaching the limit means “at least that many”. A success of the op resets the count. During the delay the decision is a requeue with the reason <reason>Backoff (for example InputsChangedBackoff); there is no way to skip it by hand.
Retries by reason¶
| What failed | What retries it | Delay |
|---|---|---|
| An apply of a mutable kind | LastApplyFailed, until an apply succeeds | Backoff |
An apply of a TerraformMachine | NoState or StateWithoutInputsHash (state has no inputs hash until an apply succeeds) | Backoff |
| A drift remediation apply | The remediation cap, below | Backoff |
| An apply Job deleted while it ran (cluster, pool) | An apply of the current inputs stays due; see below | Backoff |
refresh, drift | The schedule: still due, so due again | Backoff |
destroy | Deleting, forever | Backoff; see Destroy |
plan | The plan flow | Backoff |
restore | Nothing: not retried for the same serial | None, see below |
The remediation cap¶
An apply that remediates drift (captf.io/drift-remediation on the Job) is not retried as an unconverged apply: LastApplyFailed ignores it. Instead, DecideOp counts the failed apply Jobs that finished after the last successful drift check, stopping at an apply success. Remediation starts only while that count is below the failed limit (at least 1). At the limit, the drift stays pending (DriftDetected stays True) and no more remediation runs until the next successful drift check resets the count. Blocked and plan-changed applies do not count. See Drift.
An apply Job deleted while it ran¶
If an apply Job of a TerraformCluster or TerraformMachinePool is deleted while it runs, it never finishes and bookkeeping never reads its result, but it may have applied part of its change. CAPTF confirms with live reads (the Job is NotFound and the object’s live status.activeJob still names it) and records captf.io/interrupted-apply=<job> on the durable inputs Secret. An apply of the current inputs then stays due, even when the inputs equal the state’s, and is guarded where the destructive guard applies. ApplyJobSucceeded is False/ApplyFailed: Job <name>: disappeared while it ran and may have applied part of its change; an apply of the current inputs is due. It clears when an apply started afterwards succeeds. A stuck Job that CAPTF deleted itself, and a TerraformMachine, record nothing. See Machine pools.
Restore¶
A failed restore is not retried for the same serial
A failed restore is not retried for the same backup serial and counts toward no backoff. To try again, remove the annotation, wait for status.lastRestoredSerial to clear, and set it again. See Other manual actions.
What does not back off¶
These end the pass with a requeue, not a failure:
| Wait | Requeue |
|---|---|
A lease (WaitingForRunLease, WaitingForClusterOperation, WaitingForMachineOperations) | 30 s |
| Credentials, dependencies or an owner not ready | 30 s |
| An unreadable or lost state | 1 min |
| A Job still running | 1 min fallback; the Job watch wakes the reconcile sooner |
| The Job cache behind the API server | 5 s |
| An approval for a blocked apply or a plan | At most 10 min; an annotation wakes the reconcile |
JobPolicyInvalid, InputsTooLarge, missing durable inputs | 10 min; a change wakes the reconcile |
Metrics and events¶
Each finished Job is counted once. captf_job_attempts records its retry number (one plus the earlier failed Jobs of the op since the last success, not counting blocked and plan-changed ones). The events are JobFailed, JobInterrupted, JobDeadlineExceeded and DestructivePlanBlocked, each once per Job. See Metrics and Events.