Skip to content

Deadlines and Lock Timeouts

A Job runs against two clocks. activeDeadlineSeconds bounds the whole Job. lockTimeoutSeconds bounds how long one runner step waits for the Terraform state lock. They are set per object in spec.jobs, inherited by machines and pools from their cluster’s spec.defaults.jobs field by field, and the second must stay below the first.

Setting Default Range Enforced by
activeDeadlineSeconds 3600 1 to 86400 Kubernetes, on the Job
lockTimeoutSeconds 300 0 to 3600 Terraform or OpenTofu, as -lock-timeout

Neither is a manager flag. See Tuning Jobs for how to set them.

The deadline

Kubernetes counts activeDeadlineSeconds from the Job’s start. When it passes, the pod is sent SIGTERM and then, after the grace period, SIGKILL. The Job sets a termination grace period of 600 seconds so that a SIGTERM is not a SIGKILL in 30 seconds: the runner interrupts the runtime, which finishes the provider calls in flight (a VM being created), writes the state and releases the lock. The runner’s own stop timeout is the grace period less a 30 second margin, which leaves time to write its result.

A deadline kill is a failure, with three condition reasons by op:

Op Condition Reason
apply, plan, destroy ApplyJobSucceeded=False JobDeadlineExceeded
refresh, drift DriftJobSucceeded=False DriftJobDeadlineExceeded
restore RestoreJobSucceeded=False RestoreFailed

A deadline kill counts toward backoff

It counts toward backoff like any other failure, even though the runner reports the interruption: see Retries.

An image that cannot be pulled also ends on the deadline, but its more specific reason, ImagePullFailed, is checked first.

Raise the deadline for a slow module

Raising it also lengthens the lease backstop, because a lease is held up to the holder’s deadline plus 15 minutes.

The lock timeout

lockTimeoutSeconds is passed to the runner as --lock-timeout, and the runner adds -lock-timeout=<n>s to the steps that take or wait for the state lock:

Step Gets -lock-timeout
init Yes: it succeeds while the lock is held
apply, plan, destroy Yes
apply -refresh-only (refresh, and the first step of drift) Yes
plan -refresh=false (the second step of drift) Yes
state push (restore) Yes
force-unlock, validate, show -json, state list No

A step that cannot get the lock within the timeout fails, and the Job fails with it. That counts toward backoff. The lock itself is the Terraform state lock, which is a different mechanism from CAPTF’s leases: see The state lock and stale locks.

A timeout of 0 means do not wait: fail at once on a held lock.

The merged-policy check

The deadline must be longer than the lock wait, or a Job could reach its deadline while still waiting for a lock. The admission webhook checks one policy at a time against the built-in default of the field the policy does not set. It cannot see the field-wise merge of a machine’s policy over its cluster’s defaults, so the controller checks the merged policy again before it starts a Job:

lockTimeoutSeconds (default 300) must be less than
activeDeadlineSeconds (default 3600)

When it fails, the object reports ApplyJobSucceeded=False/JobPolicyInvalid with a message naming both effective values and whether each is configured or the built-in default, no Job starts, and the reconcile requeues in ten minutes (a change to the object triggers it sooner). The message is on ApplyJobSucceeded whichever op was due: a refresh or drift that cannot start reports there too.

Two exceptions:

  • A destroy never waits on this check, so a teardown cannot wedge on a policy error. See The destroy Job.
  • A restore does not run it either: the restore path starts its Job without the check.

To fix it, lower lockTimeoutSeconds or raise activeDeadlineSeconds in the policy that sets it, or in the cluster’s spec.defaults.jobs.