Reconcile Errors
This page helps you diagnose the CAPTFReconcileErrors alert: one of the
manager’s own controllers is repeatedly failing to reconcile, as opposed
to an object’s Job failing (see Failing Jobs). A
reconcile error means the controller could not even finish deciding what
to do — it is not a statement about any one object’s infrastructure.
Before you begin
kubectlaccess to the manager’s Deployment and its logs, in thecaptf-systemnamespace (or wherever it is installed), on the management cluster.
1. Read the manager’s logs
kubectl logs -n captf-system deploy/captf-controller-manager --tail=200
Each reconcile that returns an error logs a line carrying controller
(one of terraformcluster, terraformmachine, terraformmachinepool,
terraformmachinetemplate or terraformclusteridentity — the alert’s
controller=~"terraform.*" matches all five), the object’s namespace
and name, and the error itself. Reconciles for a TerraformMachine,
TerraformCluster or TerraformMachinePool also carry that object’s own
key (TerraformMachine, TerraformCluster, TerraformMachinePool) plus
its owning Cluster, Machine or MachinePool where one exists, so you
can filter by any of those instead of grepping free text.
The default log level (-v=2) already includes the manager’s own flow
logging in addition to errors, so raising -v is rarely needed just to
see that a reconcile failed; it helps to see why a decision was made
before the error, not to see the error itself. --diagnostics-address’s
endpoint can change the running level without a restart when
--insecure-diagnostics is not set; see
Manager Flags.
2. Narrow down which object, and how often
sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m]))
matches the alert’s own query. If it is one object erroring repeatedly, its logs name it on every line; if it is spread across many objects of one kind, look for a cause common to the whole namespace or cluster (quota, RBAC, an API server problem) rather than the object’s own spec.
Common causes
- The Kubernetes API server is unavailable or throttling requests: a
Get,Create,UpdateorPatchfailed with a server-side error (500,503, or a429from client-side rate limiting). Transient — the error clears once the API server does; if it does not, check the API server’s own health. - The manager’s own patch was rejected by the validating webhook: the
manager itself writes a
Terraform*object’s finalizer and a handful of spec fields (for example aTerraformMachine’sproviderID, aTerraformCluster’scontrolPlaneEndpoint) through the same admission path a user’skubectl applygoes through. If the webhook cannot be reached, that patch fails and shows up here. See Webhook Unavailable. - The manager’s own RBAC is missing a permission: the
ClusterRoleit runs as (captf-manager-role) was edited by hand, or an upgrade changed what it needs and the installedClusterRolewas not updated to match. The error names the verb, resource and group a403Forbiddenresponse refused. See RBAC for what the manager needs and creates. - A namespace
ResourceQuotablocks a create: the manager creates Jobs, Secrets, Leases, and a per-namespace runnerServiceAccountandRoleBindingon an object’s behalf. A quota on any of those object counts in the tenant namespace surfaces as aForbiddencreate error here rather than anywhere on theTerraform*object’s own status. - A conflicting concurrent write: two updates to the same object
raced (for example, the manager and an operator editing it at the same
moment). The condition-patching helper already retries a conflicting
status write itself; a
Conflictthat still reaches the log usually clears on the next reconcile, which controller-runtime’s own per-item backoff already schedules — distinct from a Job’s retry backoff (see Failing Jobs), and not something to act on unless it repeats for the same object.
None of these are the same as an object’s Job failing: a Job failure is
recorded on the object’s own status and conditions and never increments
controller_runtime_reconcile_errors_total — only an error the
reconciler itself returns does. See Failing Jobs if
what you are chasing is instead a failing apply, destroy, drift check or
refresh.
Confirm it worked
kubectl logs -n captf-system deploy/captf-controller-manager --tail=50 --follow
No further Reconciler error lines appear for the affected controller,
and sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m]))
returns to 0.
See also
- Failing Jobs — a Job or an object’s status, as opposed to the controller itself.
- Webhook Unavailable — a specific, common cause of reconcile errors.
- RBAC — what the manager’s own
ClusterRolegrants.