Observability
This page is for operators running the manager: what its metrics endpoint serves and who may read it, how to wire up Prometheus, what each alert means and where to look first, and how to read conditions, events and logs once you are looking at a specific object or Job.
The diagnostics endpoint
The manager serves Prometheus metrics over HTTPS, authenticated and authorized against the API server by default:
- The bearer token in the request is checked with a
TokenReview. - The token’s access is checked with a
SubjectAccessReviewforgeton the non-resource URL/metrics.
--insecure-diagnostics turns both checks off and serves plain HTTP
instead; use it only for local development.
Configuration
covers that flag and the address, TLS version and cipher flags, and
Manager Flags
lists them with their defaults.
| Request | Result |
|---|---|
GET /metrics without a token | 401 Unauthorized |
GET /metrics with a token authorized for get on /metrics | 200, serving captf_build_info and the rest of the series |
There is no flag for the metrics server’s own certificate: mount a Secret
with tls.crt and tls.key at /tmp/k8s-metrics-server/serving-certs/ to
serve your own, or leave it unmounted and the manager generates a
self-signed certificate at startup.
The same authenticated endpoint also serves a profiler
(GET /debug/pprof/*) and an endpoint to change the log level at runtime
(PUT /debug/flags/v; see Logs and verbosity) once
--insecure-diagnostics is off. Each request is authorized the same way,
against its own non-resource URL and the HTTP method lowercased as the
verb (get for the profiler, put for the log level). Nothing in the
shipped manifests grants either, only /metrics, so bind your own
ClusterRole to reach them.
The series themselves are listed in Metrics.
Enabling the Prometheus component
config/prometheus is an opt-in kustomize component: a metrics Service,
a ServiceMonitor and a PrometheusRule with the eleven alerts below,
plus a ClusterRole that lets a Prometheus instance read the diagnostics
endpoint. It is not part of infrastructure-components.yaml and is not a
release asset, so clusterctl init never installs it and clusterctl upgrade never touches it: get config/prometheus from a checkout of the
tag you installed, build it, and apply it yourself, and repeat that after
every upgrade to pick up any change to the alert rules. See
RBAC for what the ClusterRole
grants.
Before you begin:
- the Prometheus Operator CRDs (
ServiceMonitor,PrometheusRule) are installed in the cluster; - you know the namespace and name of the ServiceAccount your Prometheus
scrapes with, if it is not kube-prometheus’s default
(
monitoring/prometheus-k8s).
-
Build the component on its own — it needs no other resources, and its objects already carry the
captf-names and thecaptf-systemnamespace thatconfig/defaultproduces:apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization components: - <path-to-checkout>/config/prometheuskustomize build <path-to-this-kustomization> >captf-prometheus.yaml -
If your Prometheus does not run as
monitoring/prometheus-k8s, patch thecaptf-metrics-readerClusterRoleBinding’s subject to your ServiceAccount instead, before applying. -
Apply the built manifest with
kubectl apply -f captf-prometheus.yaml. If your provider install itself uses renamed objects (a non-defaultnamePrefixor namespace), patch theServiceMonitor’s andPrometheusRule’s selectors and theServicereference to match, since a component does not inherit an overlay’s name or namespace transformers.
The ServiceMonitor scrapes over HTTPS with insecureSkipVerify set,
since the metrics server’s certificate is self-signed by default; once you
mount a CA-signed certificate, set tlsConfig.ca instead and drop
insecureSkipVerify.
To confirm it works without waiting on Prometheus, read the endpoint yourself with the same kind of token Prometheus uses:
kubectl port-forward -n captf-system svc/captf-controller-manager-metrics 8443:8443 &
curl -sk -H "Authorization: Bearer $(kubectl create token <serviceaccount> -n <namespace> --duration=5m)" \
https://localhost:8443/metrics | grep captf_build_info
<serviceaccount> and <namespace> name a ServiceAccount already bound
to captf-metrics-reader (or your own equivalent grant). A captf_build_info
line back confirms both the endpoint and the token’s authorization; 401
or 403 means the token or the binding, not the manager.
make promtool-check and make promtool-test validate the alert rules
before you ship a change to them; see
Make Targets.
Useful queries
# p95 Job duration by op, over the last day, for Jobs that ran their course
histogram_quantile(0.95, sum by (le, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1d])))
# Slowest runner steps (p90) by kind, op and step
histogram_quantile(0.9, sum by (le, kind, op, step) (rate(captf_job_step_duration_seconds_bucket[6h])))
# Failures by error kind and step
sum by (op, error_kind, step) (increase(captf_job_errors_total[1d]))
# p90 queue time (creation to the source container's start) by op
histogram_quantile(0.9, sum by (le, op) (rate(captf_job_queue_seconds_bucket[1h])))
# The ten largest states, and how close each is to one Secret's 1 MiB
topk(10, captf_state_bytes)
captf_state_bytes / (1024 * 1024)
# Objects whose drift check or health refresh is overdue
time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 3600
No Grafana dashboard ships with the component; the queries above are the panels one would build. The full series list, with every label, is in Metrics.
Alerts
config/prometheus ships eleven alerts on the series above. Two things
help in reading them:
- The failure counters count transitions, not reconciles: they go up once
when an object or Job newly enters the bad state, so
increase(...) > 0means something newly broke, not that it is still broken. CAPTFClusterDrift,CAPTFStateNearSecretLimit,CAPTFInputsNearLimitandCAPTFNoRecentSuccessname the object directly, withnamespaceandnamelabels. The rest aggregate bykind(withoporreason), exceptCAPTFReconcileErrors, which aggregates bycontrolleralone; find the affected object through its conditions, as each section below says.
Every rule’s expression, for and severity are in
Alerts; each heading below links there.
CAPTFJobFailing
Jobs of a kind and op failed or hit their deadline more than twice in 30
minutes. Find the objects with ApplyJobSucceeded=False or
DriftJobSucceeded=False and read status.lastRun and the Job’s logs:
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | .type == ("ApplyJobSucceeded", "DriftJobSucceeded") and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"'
See Failing Jobs and the rule.
CAPTFDestroyStuck
An object’s destroy keeps failing while it deletes; its finalizer and state stay in place, so nothing is orphaned. See Stuck Destroy and the rule.
CAPTFClusterDrift
A TerraformCluster’s last drift check found changes and it has stayed
that way for an hour. With drift.action: Report an operator decides
next; with Remediate the remediation apply is failing, check
ApplyJobSucceeded. See Drift and
the rule.
CAPTFStateUnreadable
A state Secret turned unreadable. Find the object with
StateReadable=False and read that condition’s reason and message:
kubectl get terraformclusters,terraformmachines,terraformmachinepools -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | .type == "StateReadable" and .status == "False") | "\(.kind)\t\(.metadata.namespace)/\(.metadata.name)"'
See Unreadable State and the rule.
CAPTFForceUnlocks
A Job force-unlocked a state lock whose holder pod no longer existed. Find out why the previous Job died before it happens again. See Stale State Lock and the rule.
CAPTFReconcileErrors
A controller keeps returning errors from its reconcile loop. Read the manager’s logs for the failing kind. See Reconcile Errors and the rule.
CAPTFJobSlow
Jobs of a kind and op are taking longer than expected at the 90th
percentile. Compare it with the Jobs’ activeDeadlineSeconds and find the
slow step with captf_job_step_duration_seconds. See
Slow Jobs and
the rule.
CAPTFJobQueueSlow
Jobs are waiting too long between creation and the source container starting: scheduling, image pulls or the runner’s own init copy. Look for Pending runner pods and their events. See Slow Jobs and the rule.
CAPTFStateNearSecretLimit
An object’s compressed state is approaching the 1 MiB a Secret can hold.
captf_state_resources shows how many resources it manages. See
Size Limits and
the rule.
CAPTFInputsNearLimit
An object’s rendered inputs are approaching the size no Job will start past. See Size Limits and the rule.
CAPTFNoRecentSuccess
An object’s scheduled drift check or health refresh has not succeeded in
six hours. Read DriftJobSucceeded and status.lastRun, the same as for
a failing Job. See Failing Jobs and
the rule.
Reading conditions
kubectl describe on any TerraformCluster, TerraformMachine,
TerraformMachinePool or TerraformClusterIdentity shows its conditions:
a type, a status of True, False or Unknown, a reason and a message.
Most conditions are normal polarity (True is healthy); a few are
inverted, such as Deleting. Unknown most often means the object is
waiting on something else to finish, not that anything failed: it does
not fail Ready. See retry backoff
for how a wait like this is treated.
Every condition type CAPTF sets, its polarity, and every reason and message it can carry are in Conditions.
Reading events
Every stage of an object’s life emits a Kubernetes Event on it: the manager once per transition, and a Job’s runner in real time while it runs. Read them in order with:
kubectl events --for terraformcluster/<name> -n <namespace>
kubectl events --for terraformmachine/<name> -n <namespace>
kubectl events --for terraformmachinepool/<name> -n <namespace>
kubectl events --for terraformclusteridentity/<name> -n default
TerraformClusterIdentity is cluster-scoped, but its events still land
in the default namespace.
Add -o wide for a SOURCE column that distinguishes the manager’s
events from a runner’s. Notes never carry credentials, tfvars, output,
plan values or raw stderr: a step failure’s note is the runner’s curated
summary, not its log output. Every reason, its type and what it means are
in Events.
Runner events
A Job’s runner posts its own progress (RunStarted, StepStarted,
StepSucceeded, StepFailed, PlanSummary, ResourcesChanged,
RunFinished) as Events on the object the Job is for, related to the Job
itself. Emission is best effort: each request has its own short timeout,
and after a few consecutive failures the runner stops emitting for the
rest of that run without failing or slowing it. Turning --runner-events
off (see Configuration) skips them
entirely, which is worth doing on a large fleet since a single scheduled
drift check alone produces several of them. The runner’s own events
create grant is in RBAC.
Logs and verbosity
The manager logs at a default verbosity where the usual reconcile flow is visible; raising it shows more detail down to per-request tracing, and lowering it keeps only errors and irreversible actions such as force unlocks. Credentials, bootstrap data, tfvars content and output values are never logged, whatever the level.
Change the level without restarting the manager through the same authenticated diagnostics endpoint:
curl -sk -X PUT -H "Authorization: Bearer <token>" --data '<level>' \
https://<address>/debug/flags/v
The token needs its own authorization for put on /debug/flags/v; the
shipped manifests do not grant it. --v, --vmodule and
--logging-format set the level, per-file overrides and the output
format at startup instead; see
Configuration and
Manager Flags.
Read the manager’s own logs, and a Job’s, with:
kubectl logs -n captf-system deploy/captf-controller-manager -c manager
kubectl logs -n <namespace> job/<name> -c source
See also
- Metrics — every series, its type, labels and meaning.
- Alerts — every rule’s expression,
forand severity. - Conditions — every condition, reason and message.
- Events — every event reason, type and meaning.
- Configuration — the flags behind the diagnostics endpoint, runner events and logging.
- RBAC — the manager’s and Prometheus’s RBAC.
- Runbooks — the recovery procedures the alerts above link to.