# Jobs, Retries and Concurrency

Every Terraform or OpenTofu run CAPTF does is one Kubernetes Job, started by the controller, watched to the end and counted. This chapter follows a Job from the decision to run it to the bookkeeping after it finishes, and explains the machinery that keeps those Jobs from colliding: the names that make creation idempotent, the retry rules, the deadlines, the leases, the cache-lag checks, the state lock and leader election. It is the model behind the “nothing is happening” questions that [Slow Jobs](<https://captf.io/docs/operator-guide/runbooks/slow-jobs/index.md>) and [Failing Jobs](<https://captf.io/docs/operator-guide/runbooks/job-failures/index.md>) answer one symptom at a time, and it goes deeper than [The Reconcile Lifecycle](<https://captf.io/docs/concepts/lifecycle/index.md>), which it extends.

The pages:

- **Job names, attempts and history**

  ---

  Deterministic names, adoption and pruning.
- **Choosing the Operation**

  ---

  The priority order of `DecideOp`.
- **Retries and backoff**

  ---

  What counts as a failure and what does not.
- **Deadlines and lock timeouts**

  ---

  The two clocks on a Job.
- **Leases and the operation gate**

  ---

  Who may run when.
- **Cache lag**

  ---

  The live checks that cover a stale cache.
- **The state lock and stale locks**

  ---

  Terraform’s lock against CAPTF’s leases.
- **Leader election and failover**

  ---

  One active manager, and what a failover does to Jobs in flight.

Two more pages sit alongside them:

- [Requeue intervals and schedules](<https://captf.io/docs/concepts/jobs/schedules/index.md>).
- [Nothing is happening](<https://captf.io/docs/concepts/jobs/troubleshooting/index.md>): from a wait reason to the action.

This page covers the life of a Job, the operations and the flags.

## The life of a Job

```
%%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%%
flowchart TD
    A[Reconcile: bookkeeping,<br/>no active Job] --> B[DecideOp picks an operation]
    B --> C{Backoff,<br/>approval or gate?}
    C -- wait --> W[Requeue]
    C -- run --> D[Take the run lease,<br/>and the cluster lease]
    D --> E{Lease free?}
    E -- no --> W
    E -- yes --> F[Create the Job,<br/>then its per-run Secret]
    F --> G[Job runs]
    G --> H[Bookkeeping: read the pod,<br/>delete the per-run Secret]
    H --> I[Set conditions and lastRun,<br/>release the leases, prune]
    I --> A
```

Four rules shape the diagram:

- **At most one Job per object.** The reconcile does nothing else while a Job of the object runs, except to watch it. The run lease makes that hold across managers and across crashes.
- **Creation is idempotent.** Names are deterministic, so a retry after a crash or a stale read finds the Job it already created.
- **The controller owns retries.** A Job has `backoffLimit: 0`, a pod `restartPolicy: Never` and no TTL; every retry is a new Job, named with a new attempt number, started only after the backoff.
- **Results are read once.** After a Job finishes, one pass reads its pod, deletes its per-run Secret, records the outcome on the object and marks the Job as bookkept. Later passes read the Job’s annotations, not its pod.

## The operations

| Operation | What it runs | Chosen when |
| --- | --- | --- |
| `apply` | `validate`, then `apply` of the rendered inputs | No state, state without an inputs hash, changed inputs, a failed last apply, or drift remediation |
| `plan` | `validate`, then `plan` for review | A `TerraformCluster` under `applyPolicy: Manual` needs a plan to approve |
| `destroy` | `destroy` of the durable inputs | The object is deleting |
| `refresh` | `apply -refresh-only` | Once after an apply, and on the health, membership and pending schedules |
| `drift` | `apply -refresh-only`, then a plan without refresh | The drift interval is due |
| `restore` | `state push` of a backup, then `state list` | `captf.io/restore-state` names a complete backup |

Which of these wins when several are possible is [the priority order](<https://captf.io/docs/concepts/jobs/operations/index.md>). Every op is a Job of the same shape, differing in the runner’s `--op`.

## Manager flags and where the rest lives

The manager has no flags for backoff or for the failure limit. The retry timing is compiled in, and the limit is the per-object `spec.jobs.failedJobsHistoryLimit`. The flags that matter here:

| Flag | Default | What it does | Page |
| --- | --- | --- | --- |
| `--cluster-operation-gate` | `true` | A `TerraformCluster`’s apply, destroy or restore and its machines’ operations exclude each other | [Leases](<https://captf.io/docs/concepts/jobs/leases/index.md>) |
| `--sync-period` | `10m` | The informers’ resync, and the orphan sweep’s interval | [Schedules](<https://captf.io/docs/concepts/jobs/schedules/index.md>) |
| `--drift-default-interval` | `30m` | Drift interval for objects that set none | [Schedules](<https://captf.io/docs/concepts/jobs/schedules/index.md>) |
| `--terraformcluster-concurrency` and the machine, template and pool counterparts | `10` each | Reconciles in flight per kind | [Leader election](<https://captf.io/docs/concepts/jobs/leader-election/index.md>) |
| `--leader-elect` and its lease, renew and retry flags | `false`; `15s`, `10s`, `2s` | One active manager | [Leader election](<https://captf.io/docs/concepts/jobs/leader-election/index.md>) |
| `--runner-events` | `true` | Runner progress events | [Events](<https://captf.io/docs/reference/events/index.md>) |
| `--state-backups` | `5` | Backups kept per object | [Backups](<https://captf.io/docs/concepts/secret-management/backups/index.md>) |

The per-object knobs are in `spec.jobs`: `activeDeadlineSeconds`, `lockTimeoutSeconds`, `successfulJobsHistoryLimit` and `failedJobsHistoryLimit`. See [Tuning Jobs](<https://captf.io/docs/user-guide/job-tuning/index.md>). The full flag list is [Manager Flags](<https://captf.io/docs/reference/manager-flags/index.md>).

> [!NOTE]
>
> **See also**
>
> - [The Reconcile Lifecycle](<https://captf.io/docs/concepts/lifecycle/index.md>).
> - [Deletion and Teardown](<https://captf.io/docs/concepts/deletion/index.md>) for the destroy path.
> - [Approvals and Gates](<https://captf.io/docs/concepts/approvals/index.md>) for the plan flow.
> - [Drift and Health](<https://captf.io/docs/concepts/drift-and-health/index.md>).
