Alerts
CAPTF ships a PrometheusRule (config/prometheus/rules.yaml) alerting on the metrics in Metrics; each alert below links to its runbook in Observability. Promtool unit tests for the rules live in config/prometheus/tests/rules_test.yaml.
CAPTFJobFailing
- Group: captf
- Severity: warning
sum by (kind, op) (increase(captf_jobs_total{result=~"failed|deadline"}[30m])) > 2
Summary: {{ $labels.kind }} {{ $labels.op }} Jobs keep failing
Description: More than two {{ $labels.op }} Jobs for {{ $labels.kind }} objects failed or hit their deadline in 30 minutes. Find the objects with ApplyJobSucceeded=False or DriftJobSucceeded=False and read status.lastRun and the Job logs.
Runbook: ../operator-guide/observability.md#captfjobfailing
CAPTFDestroyStuck
- Group: captf
- Severity: critical
- For: 30m
sum by (kind) (increase(captf_jobs_total{op="destroy",result!="succeeded"}[30m])) > 0
Summary: {{ $labels.kind }} destroy keeps failing
Description: Destroy Jobs for {{ $labels.kind }} objects have failed for 30 minutes. The objects keep their finalizer and their state; follow the stuck-destroy runbook.
Runbook: ../operator-guide/observability.md#captfdestroystuck
CAPTFClusterDrift
- Group: captf
- Severity: warning
- For: 1h
captf_drift_detected{kind="TerraformCluster"} == 1
Summary: TerraformCluster {{ $labels.namespace }}/{{ $labels.name }} has drifted
Description: The last drift check found changes for an hour. With drift.action Report this waits for an operator; with Remediate the remediation apply keeps failing (see ApplyJobSucceeded).
Runbook: ../operator-guide/observability.md#captfclusterdrift
CAPTFStateUnreadable
- Group: captf
- Severity: critical
sum by (kind, reason) (increase(captf_state_read_errors_total[15m])) > 0
Summary: {{ $labels.kind }} state became unreadable ({{ $labels.reason }})
Description: A {{ $labels.kind }} state Secret turned {{ $labels.reason }}. Nothing is applied until it reads again. Find the object with StateReadable=False.
Runbook: ../operator-guide/observability.md#captfstateunreadable
CAPTFForceUnlocks
- Group: captf
- Severity: warning
sum by (kind) (increase(captf_lock_force_unlocks_total[1h])) > 0
Summary: A {{ $labels.kind }} state lock was force-unlocked
Description: A Job force-unlocked a state lock whose holder pod was gone. Look for ForceUnlocked events and check why the previous Job died.
Runbook: ../operator-guide/observability.md#captfforceunlocks
CAPTFReconcileErrors
- Group: captf
- Severity: warning
- For: 10m
sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m])) > 0.1
Summary: The {{ $labels.controller }} controller keeps failing to reconcile
Description: Reconcile errors above 0.1/s for 10 minutes; read the manager logs.
Runbook: ../operator-guide/observability.md#captfreconcileerrors
CAPTFJobSlow
- Group: captf
- Severity: info
- For: 30m
histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1h]))) > 1800
Summary: {{ $labels.kind }} {{ $labels.op }} Jobs are slow
Description: The 90th percentile of {{ $labels.op }} Job duration has been above 30 minutes for 30 minutes.
Runbook: ../operator-guide/observability.md#captfjobslow
CAPTFJobQueueSlow
- Group: captf
- Severity: warning
- For: 10m
histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_queue_seconds_bucket[10m]))) > 300
Summary: {{ $labels.kind }} {{ $labels.op }} Jobs wait long to start
Description: The 90th percentile of the time {{ $labels.op }} Jobs for {{ $labels.kind }} objects take from creation to running the module has been above 5 minutes. Look for Pending runner pods: node capacity, quota, slow or failing image pulls.
Runbook: ../operator-guide/observability.md#captfjobqueueslow
CAPTFStateNearSecretLimit
- Group: captf
- Severity: warning
- For: 15m
captf_state_bytes > 900 * 1024
Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} state is near the Secret size limit
Description: The compressed state is {{ $value | humanize1024 }}B, above 900 KiB. An OpenTofu state Secret cannot grow past 1 MiB, and the next apply that crosses it fails to save state.
Runbook: ../operator-guide/observability.md#captfstatenearsecretlimit
CAPTFInputsNearLimit
- Group: captf
- Severity: warning
- For: 15m
captf_inputs_bytes > 900000
Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} inputs are near the size limit
Description: The rendered main.tf.json and terraform.tfvars.json are {{ $value | humanize }} bytes. Above 1000000 no Job starts (ApplyJobSucceeded InputsTooLarge).
Runbook: ../operator-guide/observability.md#captfinputsnearlimit
CAPTFNoRecentSuccess
- Group: captf
- Severity: warning
- For: 30m
time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 6 * 3600
Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} has had no successful {{ $labels.op }} for 6 hours
Description: Its {{ $labels.op }} Jobs keep failing or never start, so drift and health go unobserved. See DriftJobSucceeded and status.lastRun.
Runbook: ../operator-guide/observability.md#captfnorecentsuccess