Troubleshooting by Condition¶
Every CAPTF object tells you what it is waiting for in its status conditions. This chapter is the lookup: given a condition and a reason, what it means, what usually causes it and what to do. It is organized as one page per kind of lookup, not per symptom, so use it when you already see a reason. When you only have a symptom, start from the runbooks.
Choose a starting point¶
-
Runbooks by Symptom
You only have a symptom or an alert. Each runbook goes from diagnosis to fix.
-
Nothing Is Happening
An object has no Job and is waiting.
-
Every Condition
Every condition type and reason, with status, meaning, likely cause, action and a link.
-
Every Event
The
Warningevents, with what to do, and the informational ones. -
Jobs
Jobs that keep failing.
-
State
State that is lost, corrupt or locked.
-
Deletion
A deletion that does not finish.
-
Access
Identities, credentials and RBAC that block an object.
This page covers where to start and how Ready is built.
Start here¶
Work from the top level down. Each step narrows the search.
- Read
Ready.kubectl get <kind> -n <ns> <name>shows it. If it isTrueand nothing seems wrong, you are done. If it isFalseorUnknown, go on. - Read its inputs.
Readysummarizes other conditions, and its message names the ones that decide it. Which ones count depends on the kind and on whether the object has been provisioned (see below). -
Find the
FalseorUnknowncondition. List them:kubectl get <kind> -n <ns> <name> -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\t"}{.message}{"\n"}{end}'Three conditions have negative polarity, where
Trueis the problem:Deleting,DriftDetectedandDeletionBlocked. Every other condition is healthy whenTrue. 4. Read its reason and message. The reason is what the table keys on. The message adds the specifics: a Job name, a key, an annotation to set. 5. Look it up in the conditions table, under the condition’s type. The row gives the cause, the action and the page that goes deeper. 6. Check the events.kubectl events -n <ns> --for <kind>/<name>shows what the controller did and when. The events page explains eachWarning.
%%{init: {"themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 28, "rankSpacing": 34, "padding": 10}}}%%
flowchart TD
A[Ready] --> B{Status?}
B -- True --> C[Healthy: check DriftDetected<br/>and ApplyJobSucceeded if unsure]
B -- False or Unknown --> D[Read the Ready message<br/>and list the conditions]
D --> E{Which input is not True?}
E -- DependenciesReady --> F[An owner, the cluster or bootstrap data]
E -- IdentityAllowed, CredentialsMirrored, RunnerRBACReady --> G[Credentials and RBAC]
E -- ApplyJobSucceeded --> H[A Job, a lease, an approval or a gate]
E -- StateReadable --> I[The state]
E -- OutputsValid, InfrastructureHealthy --> J[The module or the instance]
F --> K[Conditions table]
G --> K
H --> K
I --> K
J --> K What feeds Ready¶
| Phase | TerraformCluster and TerraformMachine | TerraformMachinePool |
|---|---|---|
| Before provisioned | DependenciesReady, IdentityAllowed, CredentialsMirrored, RunnerRBACReady, ApplyJobSucceeded, StateReadable, OutputsValid, InfrastructureHealthy, Deleting | The same nine |
| After provisioned | InfrastructureHealthy, Deleting | InfrastructureHealthy, ApplyJobSucceeded, Deleting |
Two consequences follow:
- After provisioning, a failing apply does not turn a cluster’s or machine’s
Readyfalse. A failed re-apply of a cluster must not flip the Cluster’sInfrastructureReady, which would suspend every MachineHealthCheck. CheckApplyJobSucceeded,StateReadableandDriftDetectedyourself when the object looks healthy but is not changing. - These conditions never feed
Ready, so check them by symptom:Paused,RestoreJobSucceeded,DriftJobSucceeded,DriftDetected,DeletionBlocked,EndpointAvailableandAutoscalingActive.CapacityResolvedbelongs to the template kinds andReadyto the identity, which sets it from its Secret.
By symptom¶
| You see | Look at |
|---|---|
| A new object that never provisions | DependenciesReady, then the credential conditions, then ApplyJobSucceeded |
| Nothing is happening and there is no Job | ApplyJobSucceeded wait reasons; Nothing is happening |
| Jobs keep failing | ApplyJobSucceeded, DriftJobSucceeded; Failing Jobs |
| A deletion that does not finish | Deleting, DeletionBlocked, StateReadable; My object will not delete |
| The state is lost, corrupt or locked | StateReadable; Unreadable State |
| A change waits for a person | ApplyJobSucceeded with PlanAwaitingApproval or DestructivePlanBlocked; Approvals and Gates |
| An instance is unhealthy | InfrastructureHealthy; Machine Remediation |
| The module’s outputs are rejected | OutputsValid; Module Contract |