Observability¶
This page is for operators running the manager: what its metrics endpoint serves and who may read it, how to wire up Prometheus, what each alert means and where to look first, and how to read conditions, events and logs once you are looking at a specific object or Job.
The diagnostics endpoint¶
The manager serves Prometheus metrics over HTTPS, authenticated and authorized against the API server by default:
- The bearer token in the request is checked with a
TokenReview. - The token’s access is checked with a
SubjectAccessReviewforgeton the non-resource URL/metrics.
--insecure-diagnostics turns authentication off
--insecure-diagnostics turns both checks off and serves plain HTTP instead; use it only for local development.
Configuration covers that flag and the address, TLS version and cipher flags, and Manager Flags lists them with their defaults.
| Request | Result |
|---|---|
GET /metrics without a token | 401 Unauthorized |
GET /metrics with a token authorized for get on /metrics | 200, serving captf_build_info and the rest of the series |
There is no flag for the metrics server’s own certificate: mount a Secret with tls.crt and tls.key at /tmp/k8s-metrics-server/serving-certs/ to serve your own, or leave it unmounted and the manager generates a self-signed certificate at startup.
The same authenticated endpoint also serves a profiler (GET /debug/pprof/*) and an endpoint to change the log level at runtime (PUT /debug/flags/v; see Logs and verbosity) once --insecure-diagnostics is off. Each request is authorized the same way, against its own non-resource URL and the HTTP method lowercased as the verb (get for the profiler, put for the log level). Nothing in the shipped manifests grants either, only /metrics, so bind your own ClusterRole to reach them.
The series themselves are listed in Metrics.
Enabling the Prometheus component¶
config/prometheus is an opt-in kustomize component: a metrics Service, a ServiceMonitor and a PrometheusRule with the eleven alerts below, plus a ClusterRole that lets a Prometheus instance read the diagnostics endpoint. It is not part of infrastructure-components.yaml and is not a release asset, so clusterctl init never installs it and clusterctl upgrade never touches it: get config/prometheus from a checkout of the tag you installed, build it, and apply it yourself, and repeat that after every upgrade to pick up any change to the alert rules. See RBAC for what the ClusterRole grants.
Before you begin
- the Prometheus Operator CRDs (
ServiceMonitor,PrometheusRule) are installed in the cluster; - you know the namespace and name of the ServiceAccount your Prometheus scrapes with, if it is not kube-prometheus’s default (
monitoring/prometheus-k8s).
-
Build the component on its own — it needs no other resources, and its objects already carry the
captf-names and thecaptf-systemnamespace thatconfig/defaultproduces: -
If your Prometheus does not run as
monitoring/prometheus-k8s, patch thecaptf-metrics-readerClusterRoleBinding’s subject to your ServiceAccount instead, before applying. - Apply the built manifest with
kubectl apply -f captf-prometheus.yaml. If your provider install itself uses renamed objects (a non-defaultnamePrefixor namespace), patch theServiceMonitor’s andPrometheusRule’s selectors and theServicereference to match, since a component does not inherit an overlay’s name or namespace transformers.
The ServiceMonitor scrapes over HTTPS with insecureSkipVerify set, since the metrics server’s certificate is self-signed by default; once you mount a CA-signed certificate, set tlsConfig.ca instead and drop insecureSkipVerify.
To confirm it works without waiting on Prometheus, read the endpoint yourself with the same kind of token Prometheus uses:
kubectl port-forward -n captf-system svc/captf-controller-manager-metrics 8443:8443 &
curl -sk -H "Authorization: Bearer $(kubectl create token <serviceaccount> -n <namespace> --duration=5m)" \
https://localhost:8443/metrics | grep captf_build_info
<serviceaccount> and <namespace> name a ServiceAccount already bound to captf-metrics-reader (or your own equivalent grant). A captf_build_info line back confirms both the endpoint and the token’s authorization; 401 or 403 means the token or the binding, not the manager.
make promtool-check and make promtool-test validate the alert rules before you ship a change to them; see Make Targets.
Useful queries¶
# p95 Job duration by op, over the last day, for Jobs that ran their course
histogram_quantile(0.95, sum by (le, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1d])))
# Slowest runner steps (p90) by kind, op and step
histogram_quantile(0.9, sum by (le, kind, op, step) (rate(captf_job_step_duration_seconds_bucket[6h])))
# Failures by error kind and step
sum by (op, error_kind, step) (increase(captf_job_errors_total[1d]))
# p90 queue time (creation to the source container's start) by op
histogram_quantile(0.9, sum by (le, op) (rate(captf_job_queue_seconds_bucket[1h])))
# The ten largest states, and how close each is to one Secret's 1 MiB
topk(10, captf_state_bytes)
captf_state_bytes / (1024 * 1024)
# Objects whose drift check or health refresh is overdue
time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 3600
No Grafana dashboard ships with the component; the queries above are the panels one would build. The full series list, with every label, is in Metrics.
Alerts¶
config/prometheus ships eleven alerts on the series above. Two things help in reading them:
- The failure counters count transitions, not reconciles: they go up once when an object or Job newly enters the bad state, so
increase(...) > 0means something newly broke, not that it is still broken. CAPTFClusterDrift,CAPTFStateNearSecretLimit,CAPTFInputsNearLimitandCAPTFNoRecentSuccessname the object directly, withnamespaceandnamelabels. The rest aggregate bykind(withoporreason), exceptCAPTFReconcileErrors, which aggregates bycontrolleralone; find the affected object through its conditions, as each section below says.
Every rule’s expression, for and severity are in Alerts; each heading below links there.
CAPTFJobFailing¶
Jobs of a kind and op failed or hit their deadline more than twice in 30 minutes. Find the objects with ApplyJobSucceeded=False or DriftJobSucceeded=False and read status.lastRun and the Job’s logs:
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | .type == ("ApplyJobSucceeded", "DriftJobSucceeded") and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"'
See Failing Jobs and the rule.
CAPTFDestroyStuck¶
An object’s destroy keeps failing while it deletes; its finalizer and state stay in place, so nothing is orphaned. See Stuck Destroy and the rule.
CAPTFClusterDrift¶
A TerraformCluster’s last drift check found changes and it has stayed that way for an hour. With drift.action: Report an operator decides next; with Remediate the remediation apply is failing, check ApplyJobSucceeded. See Drift and the rule.
CAPTFStateUnreadable¶
A state Secret turned unreadable. Find the object with StateReadable=False and read that condition’s reason and message:
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | .type == "StateReadable" and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"'
See Unreadable State and the rule.
CAPTFForceUnlocks¶
A Job force-unlocked a state lock whose holder pod no longer existed. Find out why the previous Job died before it happens again. See Stale State Lock and the rule.
CAPTFReconcileErrors¶
A controller keeps returning errors from its reconcile loop. Read the manager’s logs for the failing kind. See Reconcile Errors and the rule.
CAPTFJobSlow¶
Jobs of a kind and op are taking longer than expected at the 90th percentile. Compare it with the Jobs’ activeDeadlineSeconds and find the slow step with captf_job_step_duration_seconds. See Slow Jobs and the rule.
CAPTFJobQueueSlow¶
Jobs are waiting too long between creation and the source container starting: scheduling, image pulls or the runner’s own init copy. Look for Pending runner pods and their events. See Slow Jobs and the rule.
CAPTFStateNearSecretLimit¶
An object’s compressed state is approaching the 1 MiB a Secret can hold. captf_state_resources shows how many resources it manages. See Size Limits and the rule.
CAPTFInputsNearLimit¶
An object’s rendered inputs are approaching the size no Job will start past. See Size Limits and the rule.
CAPTFNoRecentSuccess¶
An object’s scheduled drift check or health refresh has not succeeded in six hours. Read DriftJobSucceeded and status.lastRun, the same as for a failing Job. See Failing Jobs and the rule.
Reading conditions¶
kubectl describe on any TerraformCluster, TerraformMachine, TerraformMachinePool or TerraformClusterIdentity shows its conditions: a type, a status of True, False or Unknown, a reason and a message. Most conditions are normal polarity (True is healthy); a few are inverted, such as Deleting. Unknown most often means the object is waiting on something else to finish, not that anything failed: it does not fail Ready. See retry backoff for how a wait like this is treated.
Every condition type CAPTF sets, its polarity, and every reason and message it can carry are in Conditions.
Reading events¶
Every stage of an object’s life emits a Kubernetes Event on it: the manager once per transition, and a Job’s runner in real time while it runs. Read them in order with:
kubectl events --for terraformcluster/<name> -n <namespace>
kubectl events --for terraformmachine/<name> -n <namespace>
kubectl events --for terraformmachinepool/<name> -n <namespace>
kubectl events --for terraformclusteridentity/<name> -n default
TerraformClusterIdentity is cluster-scoped, but its events still land in the default namespace.
Add -o wide for a SOURCE column that distinguishes the manager’s events from a runner’s. Notes never carry credentials, tfvars, output, plan values or raw stderr: a step failure’s note is the runner’s curated summary, not its log output. Every reason, its type and what it means are in Events.
Runner events¶
A Job’s runner posts its own progress (RunStarted, StepStarted, StepSucceeded, StepFailed, PlanSummary, ResourcesChanged, RunFinished) as Events on the object the Job is for, related to the Job itself. Emission is best effort: each request has its own short timeout, and after a few consecutive failures the runner stops emitting for the rest of that run without failing or slowing it. Turning --runner-events off (see Configuration) skips them entirely, which is worth doing on a large fleet since a single scheduled drift check alone produces several of them. The runner’s own events create grant is in RBAC.
Logs and verbosity¶
The manager logs at a default verbosity where the usual reconcile flow is visible; raising it shows more detail down to per-request tracing, and lowering it keeps only errors and irreversible actions such as force unlocks. Credentials, bootstrap data, tfvars content and output values are never logged, whatever the level.
Change the level without restarting the manager through the same authenticated diagnostics endpoint:
curl -sk -X PUT -H "Authorization: Bearer <token>" --data '<level>' \
https://<address>/debug/flags/v
The token needs its own authorization for put on /debug/flags/v; the shipped manifests do not grant it. --v, --vmodule and --logging-format set the level, per-file overrides and the output format at startup instead; see Configuration and Manager Flags.
Read the manager’s own logs, and a Job’s, with:
kubectl logs -n captf-system deploy/captf-controller-manager -c manager
kubectl logs -n <namespace> job/<name> -c source