Skip to content

Troubleshooting by Condition

Every CAPTF object tells you what it is waiting for in its status conditions. This chapter is the lookup: given a condition and a reason, what it means, what usually causes it and what to do. It is organized as one page per kind of lookup, not per symptom, so use it when you already see a reason. When you only have a symptom, start from the runbooks.

Choose a starting point

This page covers where to start and how Ready is built.

Start here

Work from the top level down. Each step narrows the search.

  1. Read Ready. kubectl get <kind> -n <ns> <name> shows it. If it is True and nothing seems wrong, you are done. If it is False or Unknown, go on.
  2. Read its inputs. Ready summarizes other conditions, and its message names the ones that decide it. Which ones count depends on the kind and on whether the object has been provisioned (see below).
  3. Find the False or Unknown condition. List them:

    kubectl get <kind> -n <ns> <name> -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\t"}{.message}{"\n"}{end}'
    

    Three conditions have negative polarity, where True is the problem: Deleting, DriftDetected and DeletionBlocked. Every other condition is healthy when True. 4. Read its reason and message. The reason is what the table keys on. The message adds the specifics: a Job name, a key, an annotation to set. 5. Look it up in the conditions table, under the condition’s type. The row gives the cause, the action and the page that goes deeper. 6. Check the events. kubectl events -n <ns> --for <kind>/<name> shows what the controller did and when. The events page explains each Warning.

%%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%%
flowchart TD
    A[Ready] --> B{Status?}
    B -- True --> C[Healthy: check DriftDetected<br/>and ApplyJobSucceeded if unsure]
    B -- False or Unknown --> D[Read the Ready message<br/>and list the conditions]
    D --> E{Which input is not True?}
    E -- DependenciesReady --> F[An owner, the cluster or bootstrap data]
    E -- IdentityAllowed, CredentialsMirrored, RunnerRBACReady --> G[Credentials and RBAC]
    E -- ApplyJobSucceeded --> H[A Job, a lease, an approval or a gate]
    E -- StateReadable --> I[The state]
    E -- OutputsValid, InfrastructureHealthy --> J[The module or the instance]
    F --> K[Conditions table]
    G --> K
    H --> K
    I --> K
    J --> K

What feeds Ready

Phase TerraformCluster and TerraformMachine TerraformMachinePool
Before provisioned DependenciesReady, IdentityAllowed, CredentialsMirrored, RunnerRBACReady, ApplyJobSucceeded, StateReadable, OutputsValid, InfrastructureHealthy, Deleting The same nine
After provisioned InfrastructureHealthy, Deleting InfrastructureHealthy, ApplyJobSucceeded, Deleting

Two consequences follow:

  • After provisioning, a failing apply does not turn a cluster’s or machine’s Ready false. A failed re-apply of a cluster must not flip the Cluster’s InfrastructureReady, which would suspend every MachineHealthCheck. Check ApplyJobSucceeded, StateReadable and DriftDetected yourself when the object looks healthy but is not changing.
  • These conditions never feed Ready, so check them by symptom: Paused, RestoreJobSucceeded, DriftJobSucceeded, DriftDetected, DeletionBlocked, EndpointAvailable and AutoscalingActive. CapacityResolved belongs to the template kinds and Ready to the identity, which sets it from its Secret.

By symptom

You see Look at
A new object that never provisions DependenciesReady, then the credential conditions, then ApplyJobSucceeded
Nothing is happening and there is no Job ApplyJobSucceeded wait reasons; Nothing is happening
Jobs keep failing ApplyJobSucceeded, DriftJobSucceeded; Failing Jobs
A deletion that does not finish Deleting, DeletionBlocked, StateReadable; My object will not delete
The state is lost, corrupt or locked StateReadable; Unreadable State
A change waits for a person ApplyJobSucceeded with PlanAwaitingApproval or DestructivePlanBlocked; Approvals and Gates
An instance is unhealthy InfrastructureHealthy; Machine Remediation
The module’s outputs are rejected OutputsValid; Module Contract