Skip to content

Events

CAPTF records a Kubernetes Event for each step in the life of an object: the manager once per transition, and a Job’s runner in real time while the Job runs. This page lists every reason, grouped by situation. Jump to: lifecycle, Jobs, ordering and leases, plans and approvals, state, drift and health, identity and credentials, deletion, pools and templates, runner events.

Reading events

Events are events.k8s.io/v1 objects recorded on the CAPTF object they are about. Read them in order with kubectl events:

kubectl events --for terraformcluster/<name> -n <namespace>
kubectl events --for terraformmachine/<name> -n <namespace>
kubectl events --for terraformmachinepool/<name> -n <namespace>
kubectl events --for terraformmachinetemplate/<name> -n <namespace>
kubectl events --for terraformclusteridentity/<name> -n default

kubectl describe shows the same events at the end of its output. A TerraformClusterIdentity is cluster-scoped, so its events land in the default namespace. Add -o wide to see which events the manager recorded (captf-manager) and which a runner did (captf.io/runner).

Three rules hold for every event:

  • The manager emits an event once per transition or occurrence, never on every reconcile. A reason that appears again means the thing happened again.
  • A condition that first appears emits an event only when it starts in its bad state, so a new object’s first reconcile is quiet.
  • Notes never carry credentials, tfvars, outputs, plan values or raw stderr.

Where a table says “any provisioned kind”, the event is recorded on a TerraformCluster, TerraformMachine or TerraformMachinePool. The Type column is the Kubernetes event type: Normal reports progress, Warning reports something that needs a look. For a Warning, the Action column says what to check next; the condition that backs the event is in Conditions.

Events expire

Kubernetes keeps events for about an hour by default. For history, read the object’s conditions and status.lastRun, or ship events to your log pipeline.

Lifecycle and readiness

Reason Type On Fires when Action
Provisioned Normal any provisioned kind Provisioned latched true: the first apply finished and the infrastructure exists. None.
InputsChanged Normal any provisioned kind The inputs hash differs from the state’s and an apply of the new inputs starts. None. See Inputs.
ControlPlaneEndpointSet Normal TerraformCluster spec.controlPlaneEndpoint was written from the module output. None.
FailureDomainsChanged Normal TerraformCluster status.failureDomains changed. None.
ProviderIDSet Normal TerraformMachine, TerraformMachinePool spec.providerID was written. A pool’s value can change later. None.
Paused Normal any provisioned kind The Paused condition became True. Reconciliation stops starting Jobs. Resume when ready: clear the pause on the object or its Cluster.
Resumed Normal any provisioned kind Paused went from True to False. Reconciliation resumes. None.
OutputsInvalid Warning any provisioned kind The module’s outputs broke the contract. Fix the module output named in the note. See OutputsValid.
ConditionChanged Normal or Warning any provisioned kind Any owned condition without a more specific reason changed status or reason. Warning when it moved into its bad state, Normal otherwise. For a Warning, read the condition named in the note.
DigestPinned Normal any provisioned kind An image digest was recorded on the durable inputs Secret, or pinned again after an apply of a mutable kind. None.
DigestUnknown Warning any provisioned kind No digest could be pinned, or an operation runs the spec reference for lack of one. Check that the registry is reachable and the reference resolves.

Jobs

Reason Type On Fires when Action
JobCreated Normal any provisioned kind A Job started. The note gives the operation, attempt, image and why. None.
JobSucceeded Normal any provisioned kind An apply, destroy, refresh or drift Job succeeded. None.
JobFailed Warning any provisioned kind A Job failed, or an apply or destroy could not start. Read the Job logs and status.lastRun. See the job failures runbook.
JobDeadlineExceeded Warning any provisioned kind A Job hit activeDeadlineSeconds. Find the slow step in the runner events, then raise the deadline or fix the module. See the slow jobs runbook.
JobInterrupted Warning any provisioned kind Something outside CAPTF stopped the Job, such as a node drain, an eviction or a deletion. CAPTF retries without backoff. None if it was planned. Otherwise check node pressure and preemption.
StuckJobDeleted Warning any provisioned kind A Job that could never start, because its per-run Secret is missing, was deleted so it can start again. None. Repeats mean something removes the Secret.

Ordering and leases

An object runs one operation at a time, and a machine waits for its cluster. These events say what an operation waits for. See Leases.

Reason Type On Fires when Action
WaitingForRunLease Normal any provisioned kind Another live Job holds the object’s run lease. Wait. If no Job is running, see the stale lock runbook.
WaitingForClusterOperation Normal TerraformMachine, TerraformMachinePool A machine’s apply or destroy waits for its TerraformCluster’s apply or destroy. Wait. Check the cluster if it never ends.
WaitingForMachineOperations Normal TerraformCluster A cluster’s apply or destroy waits for its machines’ applies and destroys in flight. Wait. Check the machines if it never ends.

Plans and approvals

These events belong to the destructive-plan guard and to applyPolicy: Manual. See Approvals.

Reason Type On Fires when Action
PlanReady Normal TerraformCluster A plan Job planned a change under applyPolicy: Manual. The note has the counts, the plan hash and the approve command. Once per plan Job. Review the plan, then approve it.
PlanApproved Normal TerraformCluster The apply of an approved plan started. None.
PlanApplied Normal TerraformCluster The approved plan was applied and its approval annotation removed. None.
PlanChanged Warning TerraformCluster An approved apply planned other changes and stopped before applying them. Once per such Job. Review the new plan and approve it again.
DestructivePlanBlocked Warning TerraformCluster An apply stopped before a plan that deletes or replaces resources. Once per blocked Job, in place of JobFailed. Read the plan and approve it if the deletes are intended. See Destructive guard.
DestructivePlanApprovalConsumed Normal TerraformCluster The approved destructive apply succeeded and its approval annotation was removed. None.

DestructivePlanBlocked is also recorded for a TerraformMachinePool apply that would change the cluster’s exports. The pool keeps applying with the exports of its last successful apply until you approve the change. If an earlier apply may have left such a change partly applied, the pool waits for the approval, as a cluster does.

State

See State and Backups.

Reason Type On Fires when Action
StateAdopted Normal any provisioned kind The state written by a successful apply was adopted with its new inputs hash. None.
StateBackedUp Normal any provisioned kind A new state serial was copied into a backup. Once per backup. None.
StateRestored Normal any provisioned kind A restore Job pushed a backup into the backend and the captf.io/restore-state annotation was removed. None.
StateRestoreFailed Warning any provisioned kind A restore Job failed. CAPTF does not retry it for the same serial. Read the restore Job logs. See the state restore runbook.
StateLocked Warning any provisioned kind Something else holds the state lock (StateReadable False, reason StateLocked). Wait, or find the holder. See the stale lock runbook.
ForceUnlocked Warning any provisioned kind CAPTF force-unlocked a stale state lock. Find out why the previous Job died. Alert: CAPTFForceUnlocks.
StateUnreadable Warning any provisioned kind The state could not be read. See the state unreadable runbook. Alert: CAPTFStateUnreadable.
StateLost Warning any provisioned kind A provisioned object’s state is gone or carries no inputs hash (StateReadable False, reason StateLost). Restore from a backup. See the total state loss runbook.
OwnerReferencesRepaired Normal any provisioned kind Secrets of the object (state, backups, durable inputs, plan key or its credential mirror entry) had no owner reference, or one to an earlier UID, as a management-cluster restore leaves them. They are owned by the object again. None. See the move runbook.

Drift and health

See Drift and health.

Reason Type On Fires when Action
DriftDetected Warning any provisioned kind A drift check found a difference. Once per finding. Read the drift plan summary. With drift.action: Report you decide; with Remediate CAPTF applies. Alert: CAPTFClusterDrift.
DriftRemediationStarted Normal any provisioned kind An apply that remediates drift started. None.
DriftResolved Normal any provisioned kind DriftDetected went from True to False. None.
InstanceHealthy Normal any provisioned kind InfrastructureHealthy became True. None.
InstanceUnhealthy Warning any provisioned kind InfrastructureHealthy became False for an unhealthy, degraded, stopped or terminated instance. Check the instance in the cloud console. See InfrastructureHealthy.
RemediationRequested Warning TerraformMachine CAPTF annotated the owner Machine with cluster.x-k8s.io/remediate-machine. Cluster API replaces the machine if a MachineHealthCheck acts on it.
RemediationWithdrawn Normal TerraformMachine The instance read healthy again and CAPTF removed the annotation it set. None.

Identity and credentials

See Identity and credentials.

Reason Type On Fires when Action
IdentityNotAllowed Warning any provisioned kind The identity does not allow the object’s namespace, or does not exist. Add the namespace to the identity’s allowed namespaces. See IdentityAllowed.
IdentitySecretFound Normal TerraformClusterIdentity The identity’s credentials Secret appeared. None.
IdentitySecretNotFound Warning TerraformClusterIdentity The identity’s credentials Secret went missing. Recreate the Secret.
MirrorCreated Normal any provisioned kind The credential mirror of the namespace was created on behalf of the object. None.
MirrorRemoved Normal any provisioned kind The credential mirror of the namespace was deleted on behalf of the object. None.

Deletion

See Deletion.

Reason Type On Fires when Action
DeletionStarted Normal any provisioned kind The first reconcile with a deletionTimestamp. None.
Destroyed Normal any provisioned kind The destroy succeeded and cleanup ran. None.
FinalizerRemoved Normal any provisioned kind The finalizer was removed. The object goes away. None.
InfrastructureAbandoned Warning any provisioned kind A deletion held on lost or unreadable state, or whose destroy failed or cannot start, was released by captf.io/abandon-infrastructure naming the object’s UID. The finalizer was removed without a destroy. Delete the orphaned cloud resources by hand. See Manual finalizer.

Pools and templates

Reason Type On Fires when Action
ReplicasWrittenBack Normal TerraformMachinePool An autoscaled pool’s observed replicas output was written to MachinePool.spec.replicas. The note reads X → Y. None.
ReplicasManagedExternally Warning TerraformMachinePool The autoscaler annotations are valid, but another owner holds replicas-managed-by. CAPTF does not write spec.replicas back. Decide which controller owns the replica count.
CapacityResolved Normal TerraformMachineTemplate The template’s capacity or nodeInfo changed from its image labels. None.
ImageInspectFailed Warning TerraformMachineTemplate The registry could not be read for capacity. Check registry access and credentials.

Runner events

A Job’s runner posts its own progress as events on the object the Job is for, related to the Job itself. They fire in real time, which makes them the quickest way to see which step a long Job is in. The manager enables them with --runner-events, which defaults to true; see Manager flags. The runner needs create on events in its ClusterRole, which the shipped manifests grant.

Reason Type Fires when
RunStarted Normal The runtime is ready and the first step is about to run. The note gives the operation, image reference and runtime version.
StepStarted Normal A runtime step started.
StepSucceeded Normal A runtime step finished.
StepFailed Warning A runtime step failed. The note is the runner’s curated failure summary, never raw stderr.
PlanSummary Normal The runner parsed a plan (drift, or a guarded cluster apply). The note has counts only.
ResourcesChanged Normal An apply or destroy step finished. The note has the counts from the runtime’s summary line.
RunFinished Normal or Warning The run ended. The note gives the result and total duration. Warning when the run did not succeed.

Runner events are best effort

Each request has a two-second timeout. After three consecutive failures the runner stops emitting for the rest of that run. Emission never fails or slows the run. Turn --runner-events off on a large fleet: a single scheduled drift check alone produces several events.

For where to read events during an incident, see Observability.