Failing Jobs
This page helps you diagnose a Job that failed, or an object whose apply,
destroy, drift check or health refresh keeps failing. It covers the
CAPTFJobFailing and CAPTFNoRecentSuccess alerts and reading a failure
directly from an object’s status, for a TerraformCluster,
TerraformMachine or TerraformMachinePool.
Before you begin
kubectlaccess to read the failing object and its Job’s pod logs, in its namespace, on the management cluster.- The object’s kind, namespace and name.
CAPTFJobFailingcarries onlykindandop, so find the object through its conditions (step 1);CAPTFNoRecentSuccessalready carriesnamespaceandname.
1. Find the failing objects
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(any(.status.conditions[]?; (.type=="ApplyJobSucceeded" or .type=="DriftJobSucceeded") and .status=="False")) | "\(.kind) \(.metadata.namespace)/\(.metadata.name)"'
ApplyJobSucceeded covers apply and destroy; DriftJobSucceeded covers
drift checks and health refreshes. CAPTFNoRecentSuccess fires when a
scheduled drift check or refresh has not succeeded for six hours, which
usually means one of these Jobs is failing repeatedly rather than missing
once.
2. Read the condition’s reason
The reason on ApplyJobSucceeded or DriftJobSucceeded says where to
look next; see Conditions for the full
list of reasons each condition can carry. The ones that matter here:
ApplyFailed,DestroyFailedorDriftJobFailed: a Job ran and its runtime failed partway through.status.lastRuncarries the detail — go to step 3.ImagePullFailed: a container stayed inErrImagePullorImagePullBackOffuntil the Job’s deadline. No step ever ran, sostatus.lastRuncarries nothing useful — go to Image pull failures.ImageInvalid: the image does not follow the module contract. Go to Image layout errors.JobDeadlineExceeded: the Job’s pod ran pastactiveDeadlineSecondswithout finishing. Go to Deadline exceeded.DestructivePlanBlockedorPlanChanged: not a failure — an apply stopped on purpose, waiting for an approval, and does not count towardCAPTFJobFailing. See Plan Approval.
3. Read status.lastRun for a completed run
When the runner actually started and produced a result — every reason
above except ImagePullFailed and a JobDeadlineExceeded that hit before
any step ran — status.lastRun.error.kind classifies what happened. See
the full field list in API Reference.
step: a runtime command failed.error.stepnames it — usuallyinit,validate,plan,show-json,apply,apply-refresh-onlyordestroy; occasionallyforce-unlock(retrying past a stale lock left by a previous run; see the stale state lock runbook) orprepare(the runner’s own environment setup, before any runtime command ran) — anderror.summarygives the runner’s curated one-line reason, at most 512 bytes — the module’s or the cloud provider’s own error, never raw output. Read the rest from the Job’s pod logs (step 4).image-layout: the image does not follow the module contract; the same cause asImageInvalidabove. See Image layout errors.interrupted: the runner was sentSIGTERMbefore it could finish — a node drain, an eviction, the Job being deleted, or the deadline. It is retried without counting toward backoff; see Retry backoff.blockedorplan-changed: see Plan Approval.
status.lastRun.steps lists every step the runner completed, with its
exit code and duration — useful for finding a slow step even when the run
as a whole succeeded.
4. Read the Job’s pod logs
kubectl get job -n <ns> <job-name>
kubectl logs -n <ns> job/<job-name> -c source
<job-name> is status.lastRun.job for a finished run, or
status.activeJob.name for one still running. The source container
runs the pinned image and the module; a separate init container only
copies the runner binary into it and rarely fails on its own (that
failure surfaces as an image pull failure on the manager’s --runner-image
instead, a cluster-wide problem rather than one specific to this object).
status.lastRun.error.summary is a curated, size-limited copy of the
failing step’s output, not the raw log, because the object’s status is
readable by anyone who can get it — the full log is only in
kubectl logs.
By default the newest three failed Jobs of each operation are kept
(spec.jobs.failedJobsHistoryLimit; see
Tuning Jobs), so the pod and its logs
are usually still there for a fresh alert.
kubectl events --for <kind>/<name> -n <ns> --types=Warning also carries a
curated summary of each failure as a JobFailed or StepFailed event,
even after the Job itself is pruned.
Image pull failures
ImagePullFailed means a container was stuck pulling its image until the
Job’s deadline — never that the module ran and failed. Two images can be
at fault:
- the
sourcecontainer’s image,spec.source.image— check for a typo, a missing pull secret, or a tag that was deleted from the registry; - the runner init container’s image, the manager’s
--runner-image(defaults to the manager’s own image) — a cluster-wide problem, not specific to this object.
kubectl get pods -n <ns> -l batch.kubernetes.io/job-name=<job-name>
kubectl describe pod -n <ns> <pod-name>
The Events section names the image and the pull error
(ErrImagePull/ImagePullBackOff). Fix the image reference or the pull
credentials (spec.jobs.imagePullSecrets; see
Tuning Jobs), then let the next
reconcile retry — no manual retrigger is needed.
Image layout errors
ImageInvalid (status.lastRun.error.kind: image-layout) means the
runner started but the image itself does not follow the module contract:
it is missing the module directory the contract requires, or its entry
point is not executable. This is a problem with the image, not with the
object’s spec or a transient failure — no retry fixes it. See
Image Contract for the layout an
image must follow. Republish the image and point spec.source.image at
the new tag or digest to try again.
Deadline exceeded
JobDeadlineExceeded means the Job’s pod ran past
spec.jobs.activeDeadlineSeconds (one hour unless set) without
finishing. Kubernetes sends the runner SIGTERM; it gets most of the
Job’s termination grace period to finish an in-flight provider call and
write state before being killed, so infrastructure created before the
deadline is not lost even though the Job is marked failed. Common causes:
- a module that is simply slow relative to its deadline — raise
spec.jobs.activeDeadlineSeconds; see Tuning Jobs; - a state lock wait: the first step that locks the state (
plan,apply,apply-refresh-onlyordestroy—initnever takes the lock) waits at mostspec.jobs.lockTimeoutSeconds(five minutes unless set) for it before failing on its own — well inside the default deadline, so a lock wait alone should rarely be the cause; see Slow Jobs and the stale state lock runbook if it is; - a provider call that never returns — the module’s or the cloud API’s problem; the log up to where it stopped is the only record when the runner is killed before it can write a result.
Retry backoff
A Job that failed for step, image-layout or a pull failure counts
toward that operation’s retry backoff; a blocked or plan-changed Job
never counts, since it waits for an approval instead of a timer. See
The Reconcile Lifecycle for
the delay itself. A deadline counts too, unless the runner caught its
SIGTERM in time to report itself interrupted
(status.lastRun.error.kind: interrupted) — the common case, which
retries on the next reconcile with no wait, the same as any other
interruption. There is no way to skip a real backoff by hand; retrying
immediately against an unresolved cause (a bad image, a still-held lock)
only fails again the same way.
Confirm it worked
kubectl get <kind> -n <ns> <name> -o json \
| jq '.status.conditions[] | select(.type=="ApplyJobSucceeded" or .type=="DriftJobSucceeded")'
The condition returns to status: "True", and the next
status.lastRun.error is empty.
See also
- Slow Jobs — a Job that runs but takes far longer than expected.
- Stale State Lock — a lock held long enough to fail a Job on its own.
- Plan Approval — approving a blocked or changed plan.
- Tuning Jobs — deadlines, lock timeouts, history limits and pull secrets.