Metrics¶
The manager exposes captf_* series that describe Jobs, reconcile decisions, state, drift and health. The shipped alerts are written on them, and the examples below build on the same series. Jump to: Jobs, reconcile decisions, state and inputs, drift and health, locks, leases and approvals, identity and images, build info, controller-runtime series.
Where they are served¶
The metrics are on the manager’s diagnostics endpoint, HTTPS on :8443 (--diagnostics-address), at /metrics. Requests are authenticated with a TokenReview and authorized with a SubjectAccessReview for get on the non-resource URL /metrics; a request without a token gets 401. --insecure-diagnostics serves plain HTTP with no checks and is for local development only. See Observability and Manager flags.
The opt-in config/prometheus component ships a Service (captf-controller-manager-metrics, port 8443, named metrics), a ServiceMonitor that scrapes it over HTTPS with the Prometheus ServiceAccount’s token, and the RBAC that token needs. It is not part of the release manifest; see Enabling the Prometheus component.
Conventions¶
- Metrics are declared with
k8s.io/component-base/metricsat stability levelALPHA, so names and labels can change between releases. On the wire, eachHELPstring starts with[ALPHA]. - Durations are in seconds, sizes in bytes and timestamps in Unix seconds.
- Labels are bounded enums:
kind,op,result,reason,step,actionanderror_kind. Only the eight per-object gauges carrynamespaceandname, and CAPTF removes a gauge’s series when its object is deleted, so they do not accumulate. kindis the CAPTF kind:TerraformCluster,TerraformMachineorTerraformMachinePool.opis the operation:apply,destroy,refresh,drift,planorrestore.- Failure counters count transitions, not reconciles. They go up once when an object or Job newly enters a bad state, so use
increase()to ask whether something newly broke.
Jobs¶
Every runner Job is observed when it completes. A Job is one op for one object; see Jobs.
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_jobs_total | counter | kind, op, result | Jobs | Jobs completed, by result. |
captf_job_duration_seconds | histogram | kind, op, result | seconds | Wall time of a completed Job, start to finish. |
captf_job_step_duration_seconds | histogram | kind, op, step | seconds | Wall time of each runner step of a completed Job. |
captf_job_queue_seconds | histogram | kind, op | seconds | Time from a Job’s creation to its source container’s start. |
captf_job_errors_total | counter | kind, op, error_kind, step | Jobs | Jobs that did not succeed, by error kind and failing step. |
captf_jobs_active | gauge | kind, op | Jobs | Jobs running now, counted from the Job cache at scrape time. |
captf_job_attempts | histogram | kind, op | attempts | Retry number of a Job that succeeded. |
captf_resources_changed_total | counter | kind, op, action | resources | Resources an apply or destroy Job changed. |
Label values:
| Label | Values |
|---|---|
result | succeeded, failed, deadline, interrupted (stopped from outside: a drain, eviction or deletion), blocked (a guarded apply stopped before a plan that deletes or replaces resources, awaiting approval), plan_changed (an approved apply planned other changes and stopped). |
step | init, force-unlock, validate, plan, show-json, apply, apply-refresh-only, destroy, state-push, state-list, prepare; anything else is other. On captf_job_errors_total, none when no step failed. |
error_kind | step, image-layout, interrupted, blocked, plan-changed; deadline or unknown when the Job left no result. |
action | add, change, destroy, import. |
Notes on how each series is built:
op="plan"is a plan Job underapplyPolicy: Manual; it applies nothing.blockedandplan_changedare not failures. Each is reported by its own condition and event, andCAPTFJobFailingdoes not count them.captf_job_duration_secondsbuckets run from 15 seconds to 2 hours, andcaptf_job_step_duration_secondsfrom 1 second to 2 hours.captf_job_queue_secondshas buckets from 5 seconds to 30 minutes, with 300 seconds, theCAPTFJobQueueSlowthreshold, as a boundary. CAPTF does not record it when the pod reports no start.captf_job_attemptsis 1 plus the failed Jobs of the op since it last succeeded. Interrupted Jobs do not count.captf_resources_changed_totalcomes from the runtime’s final summary line.
# p95 apply duration by kind over the last day, for Jobs that ran their course
histogram_quantile(0.95, sum by (le, kind) (
rate(captf_job_duration_seconds_bucket{op="apply",result=~"succeeded|failed"}[1d])))
# Failures by error kind and failing step, last hour
sum by (kind, op, error_kind, step) (increase(captf_job_errors_total[1h])) > 0
# Slowest runner steps (p90)
topk(5, histogram_quantile(0.9, sum by (le, kind, op, step) (
rate(captf_job_step_duration_seconds_bucket[6h]))))
# Resources destroyed by applies in the last 24 hours
sum by (kind) (increase(captf_resources_changed_total{op="apply",action="destroy"}[24h]))
Reconcile decisions¶
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_reconcile_op_decisions_total | counter | kind, op, reason | decisions | What a reconcile decided to run, and why. |
op is none when the reconcile decided to run nothing. reason says why: for example NoState, InputsChanged, DriftDue, DriftRemediation or Deleting start a Job, while UpToDate, JobActive, LastApplyFailedBackoff or DeletingBackoff start nothing.
# Why applies started in the last day
sum by (kind, reason) (increase(captf_reconcile_op_decisions_total{op="apply"}[1d]))
State and inputs¶
The per-object gauges hold the last value the manager read. See State.
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_state_read_errors_total | counter | kind, reason | reads | State reads that turned unreadable. |
captf_state_resources | gauge | kind, namespace, name | resources | Managed resources, not data sources, in the object’s state. |
captf_state_bytes | gauge | kind, namespace, name | bytes | Compressed state size summed over its Secrets. |
captf_inputs_bytes | gauge | kind, namespace, name | bytes | Size of the rendered main.tf.json and terraform.tfvars.json. |
captf_state_backups_total | counter | kind, result | backups | State backups, by result. |
captf_state_restores_total | counter | kind, result | restores | State restores requested with captf.io/restore-state. |
captf_outputs_invalid_total | counter | kind, reason | outputs | Outputs that turned invalid against the module contract. |
captf_inputs_hash_changes_total | counter | kind | applies | Applies started because the inputs of a mutable kind changed. |
Label values:
| Label | Values |
|---|---|
state_read_errors_total{reason} | inconsistent, encrypted, corrupt (an unsupported state version counts as corrupt), lost (a provisioned object’s state is gone), locked (held by something other than this object’s runner). |
state_backups_total{result} | taken (a new serial copied into captf-state-backup-* Secrets), pruned (a backup beyond --state-backups deleted), skipped (a new serial not backed up: encrypted, unreadable or oversized state, or a failed copy). |
state_restores_total{result} | succeeded, failed (a restore Job finished), not_found (the annotation names no backup). |
The Kubernetes backend holds at most 1 MiB per Secret, and no Job starts when the inputs exceed 1000000 bytes. CAPTFStateNearSecretLimit and CAPTFInputsNearLimit warn at 900 KiB and 900000 bytes.
# The ten largest states, as a share of one Secret's 1 MiB
topk(10, captf_state_bytes / (1024 * 1024))
# Objects with unreadable state in the last 15 minutes, by reason
sum by (kind, reason) (increase(captf_state_read_errors_total[15m])) > 0
Drift and health¶
See Drift and health. The gauges use the condition value: 1 for True, 0 for False and -1 for Unknown.
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_ready | gauge | kind, namespace, name | 1, 0 or -1 | The Ready condition. |
captf_infrastructure_healthy | gauge | kind, namespace, name | 1, 0 or -1 | The InfrastructureHealthy condition. |
captf_drift_detected | gauge | kind, namespace, name | 1 or 0 | 1 while DriftDetected is True, else 0. |
captf_drift_resources_total | counter | kind, action | resources | Resources a drift Job that found drift would add, change or destroy. |
captf_last_success_timestamp_seconds | gauge | kind, namespace, name, op | Unix seconds | When the newest successful Job of an op finished. |
captf_unhealthy_samples | gauge | namespace, name | samples | A TerraformMachine’s consecutive unhealthy health samples (status.unhealthySamples). |
captf_remediation_requests_total | counter | action | requests | cluster.x-k8s.io/remediate-machine annotations set on (requested) or removed from (withdrawn) a Machine. |
captf_last_success_timestamp_seconds exports drift and refresh only while that op is scheduled: not while the object is deleting or paused, and only with a drift interval or health checks set. A series that goes missing is therefore not a failure by itself.
# Objects that are not Ready
captf_ready == 0
# Objects whose drift check or health refresh is overdue, in hours
(time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"}) / 3600 > 6
Locks, leases and approvals¶
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_lock_force_unlocks_total | counter | kind | unlocks | Stale state locks force-unlocked. |
captf_lease_waits_total | counter | kind, reason | waits | Operations that started waiting for a run lease, once per wait. |
captf_plan_approvals_total | counter | kind, result | approvals | Plans approved with captf.io/approve-plan under applyPolicy: Manual. |
captf_destructive_plan_approvals_consumed_total | counter | kind | approvals | Destructive-plan approvals removed after the approved apply succeeded. |
Label values:
| Label | Values |
|---|---|
lease_waits_total{reason} | run_lease (another live Job of the object holds it), cluster_operation (a machine’s apply or destroy waits for its TerraformCluster’s), machine_operations (a cluster’s apply or destroy waits for its machines’). |
plan_approvals_total{result} | approved (the apply of the approved plan succeeded and the annotation was removed), changed (the approved apply planned other changes and stopped; the new plan waits for approval). |
Identity and images¶
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_identity_denied_total | counter | reason | refusals | Identity refusals: notfound or namespace. |
captf_image_inspect_errors_total | counter | reason | errors | Registry or image-label failures while resolving template capacity. |
Build info¶
| Metric | Type | Labels | Unit | Measures |
|---|---|---|---|---|
captf_build_info | gauge | version, commit, contract | constant 1 | The manager’s version, commit and module contract. |
Controller-runtime and other series¶
The same endpoint serves series that CAPTF does not define:
| Family | What it covers |
|---|---|
controller_runtime_* | Reconcile counts, errors, time and active workers per controller. controller_runtime_reconcile_errors_total backs CAPTFReconcileErrors. |
workqueue_* | Queue depth, latency and retries per controller. |
rest_client_* | Client requests to the API server, by code and verb. |
go_*, process_* | Go runtime and process statistics. |
kubernetes_feature_enabled | Feature gates, from component-base. |
CAPTF merges its registry with the controller-runtime defaults and the component-base registry, and drops component-base’s own go_* and process_* families because controller-runtime already serves them.
# Reconcile errors per second by controller
sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m]))