Tuning Jobs¶
Every operation CAPTF runs — apply, destroy, drift, refresh, restore and plan — is one Kubernetes Job. spec.jobs on a TerraformCluster, TerraformMachine or TerraformMachinePool tunes that Job: its resources, deadlines, history, environment, pull secrets, ServiceAccount and security contexts. Every field is optional and has a built-in default, so you only set the fields you need to change. For the full field list, types and validation see spec.jobs in Common Fields; for the Job’s exact shape (containers, mounts, args) see Job Environment.
Before you begin
- Know which object’s
spec.jobsyou are changing. ATerraformClusteralso hasspec.defaults.jobs, which itsTerraformMachines andTerraformMachinePools inherit; see Inheriting from a cluster’s defaults below. - A
TerraformCluster’s ownspec.jobsdoes not inheritspec.defaults.jobs: defaults are for its machines and pools, never for the cluster itself.
Resources¶
spec.jobs.resources sets the main container’s resources as a whole; when unset the controller applies its own default (250m CPU / 512Mi memory requested, 2Gi memory limit, deliberately no CPU limit — throttling a slow apply is worse than a slow apply). Raise it when a module pulls large provider plugins or holds a large plan in memory; the init container that copies the runner binary is not configurable, since it never varies with the module. See Default resources for the exact values.
Deadlines and lock waits¶
activeDeadlineSecondsbounds the whole Job, at most one day. Defaults to 3600 (one hour). Raise it for a module with a slow apply or many resources; a Job that hits its deadline is interrupted with SIGTERM, the same as a pod eviction or deletion. For an apply, destroy or plan Job,ApplyJobSucceededis set False with reasonJobDeadlineExceeded; for a drift or refresh Job,DriftJobSucceededis set False with reasonDriftJobDeadlineExceeded; for a restore Job,RestoreJobSucceededis set False with reasonRestoreFailed. A Job killed by its deadline is a failure: it counts toward retry backoff and the failed limit like any other failed Job. See Conditions.lockTimeoutSecondsis passed to the runtime as-lock-timeout. Defaults to 300 (five minutes). Raise it when Jobs commonly queue behind each other on the same state lock.lockTimeoutSecondsmust be less thanactiveDeadlineSeconds, including inherited and built-in defaults: the webhook rejects a policy where a lock wait alone could fill the whole deadline. A policy that sets only one of the two is checked against the built-in default of the other (300 seconds for the lock wait, 3600 for the deadline), solockTimeoutSeconds: 4000alone is rejected. The webhook sees one policy at a time, so it cannot see a machine’s or pool’s field-wise merge with a cluster’sspec.defaults.jobs. Reconcile checks the merged policy instead: an inconsistent result reportsApplyJobSucceeded=False/JobPolicyInvalidand no Job runs until you fix one of the two values.
History limits¶
successfulJobsHistoryLimit and failedJobsHistoryLimit cap how many finished Jobs of each kind are kept, both defaulting to 3. Each is scoped per object and per operation: apply, destroy, drift, refresh, restore and plan are pruned independently, so lowering one does not shrink another’s history. The newest Job of an operation is always kept, even at a limit of 0 — for a failed operation, until a newer Job of the same operation succeeds. failedJobsHistoryLimit also bounds retry backoff, since backoff looks at the same retained failures; see The reconcile lifecycle for how backoff is computed.
Environment variables¶
spec.jobs.env adds environment variables to the main container. Entries named TF_* or KUBE_* are reserved for the runner and the Job’s own environment (TF_IN_AUTOMATION, KUBE_NAMESPACE, and so on): See spec.jobs.env rejected names for the full list of names the Job already sets.
A reserved environment name is silently dropped
An entry using a TF_* or KUBE_* name is dropped instead of applied, with no error.
Image pull secrets¶
spec.jobs.imagePullSecrets covers both the role image (spec.source.image) and the runner’s own init image, so one list is enough even when they come from different registries.
ServiceAccount¶
spec.jobs.serviceAccountName overrides the runner ServiceAccount. Left unset, the controller creates and uses captf-runner, bound to the static captf-runner ClusterRole. An override must already exist and carry the label captf.io/runner=true; without it, no Job is created and RunnerRBACReady reports False with reason ServiceAccountNotOptedIn (see Conditions). Setting up and opting in your own ServiceAccount, and the RBAC CAPTF manages around it, is covered in RBAC.
Security contexts¶
spec.jobs.securityContext sets the main container’s securityContext. The container holds cloud credentials, so the webhook rejects privileged: true, allowPrivilegeEscalation: true, any capabilities.add, readOnlyRootFilesystem: false, a seccompProfile of Unconfined, procMount: Unmasked, windowsOptions.hostProcess and an explicit runAsUser: 0 or runAsNonRoot: false on it, whatever else the policy sets. Every other field you set is applied on top of the controller’s defaults (no privilege escalation, every capability dropped, a read-only root filesystem). capabilities always drops ALL, even when you supply a capabilities object.
An image that runs as root by default still runs as root
The webhook rejects only an explicit root setting; the Job sets neither runAsUser nor runAsNonRoot, so an image that runs as root by default is still admitted and runs as root.
spec.jobs.podSecurityContext sets the Job pod’s securityContext. The controller defaults its seccompProfile to RuntimeDefault and its fsGroup to the runner’s UID, so a non-root image user can read the identity credential files (mode 0440) through that group; set your own fsGroup to override it, or your own runAsNonRoot/runAsUser to run the main container as a specific non-root user. For the trust boundary these defaults protect — why the Job is treated like any other pod that holds cloud credentials — see The security model.
Rules apply on create and on policy change
The lockTimeoutSeconds and securityContext rules on this page apply on create and whenever the jobs policy changes. An existing object with an older, weaker policy still accepts unrelated updates and can always be deleted.
Inheriting from a cluster’s defaults¶
A TerraformMachine’s or TerraformMachinePool’s spec.jobs is merged field by field with its TerraformCluster’s spec.defaults.jobs: a field the machine or pool sets wins, an unset one falls back to the cluster’s default, and a field neither sets gets the built-in default above. Two fields merge instead of falling back as a whole:
envis merged by name — the machine’s or pool’s entries first, then any of the cluster default’s entries whose name they do not already use.imagePullSecretsis the union of both lists, without duplicates, the machine’s or pool’s first.
resources, securityContext and podSecurityContext are each replaced as a whole: setting any one of them on the machine or pool drops the cluster default’s value for that field entirely, rather than merging individual keys inside it. See What a cluster passes to its machines and pools for how this fits the rest of spec.defaults.
Defaults are resolved at reconcile time and never persisted, so raising a cluster’s spec.defaults.jobs reaches every existing machine and pool on their next reconcile.
Confirm it worked¶
kubectl get job -n <namespace> -l captf.infrastructure.cluster.x-k8s.io/owner-name=<object-name>
kubectl get job -n <namespace> <job-name> -o jsonpath='{.spec.activeDeadlineSeconds}{"\n"}{.spec.template.spec.serviceAccountName}{"\n"}'
<namespace> and <object-name> are the namespace and name of the TerraformCluster, TerraformMachine or TerraformMachinePool you changed; <job-name> is one Job name from the first command’s output. Compare the Job’s spec.template.spec.containers[0].resources, securityContext and env against what you set; a field you expected to change but that still shows the built-in default usually means it was set on the wrong object, or dropped by the merge rules above.