Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Alerts

CAPTF ships a PrometheusRule (config/prometheus/rules.yaml) alerting on the metrics in Metrics; each alert below links to its runbook in Observability. Promtool unit tests for the rules live in config/prometheus/tests/rules_test.yaml.

CAPTFJobFailing

  • Group: captf
  • Severity: warning
sum by (kind, op) (increase(captf_jobs_total{result=~"failed|deadline"}[30m])) > 2

Summary: {{ $labels.kind }} {{ $labels.op }} Jobs keep failing

Description: More than two {{ $labels.op }} Jobs for {{ $labels.kind }} objects failed or hit their deadline in 30 minutes. Find the objects with ApplyJobSucceeded=False or DriftJobSucceeded=False and read status.lastRun and the Job logs.

Runbook: ../operator-guide/observability.md#captfjobfailing

CAPTFDestroyStuck

  • Group: captf
  • Severity: critical
  • For: 30m
sum by (kind) (increase(captf_jobs_total{op="destroy",result!="succeeded"}[30m])) > 0

Summary: {{ $labels.kind }} destroy keeps failing

Description: Destroy Jobs for {{ $labels.kind }} objects have failed for 30 minutes. The objects keep their finalizer and their state; follow the stuck-destroy runbook.

Runbook: ../operator-guide/observability.md#captfdestroystuck

CAPTFClusterDrift

  • Group: captf
  • Severity: warning
  • For: 1h
captf_drift_detected{kind="TerraformCluster"} == 1

Summary: TerraformCluster {{ $labels.namespace }}/{{ $labels.name }} has drifted

Description: The last drift check found changes for an hour. With drift.action Report this waits for an operator; with Remediate the remediation apply keeps failing (see ApplyJobSucceeded).

Runbook: ../operator-guide/observability.md#captfclusterdrift

CAPTFStateUnreadable

  • Group: captf
  • Severity: critical
sum by (kind, reason) (increase(captf_state_read_errors_total[15m])) > 0

Summary: {{ $labels.kind }} state became unreadable ({{ $labels.reason }})

Description: A {{ $labels.kind }} state Secret turned {{ $labels.reason }}. Nothing is applied until it reads again. Find the object with StateReadable=False.

Runbook: ../operator-guide/observability.md#captfstateunreadable

CAPTFForceUnlocks

  • Group: captf
  • Severity: warning
sum by (kind) (increase(captf_lock_force_unlocks_total[1h])) > 0

Summary: A {{ $labels.kind }} state lock was force-unlocked

Description: A Job force-unlocked a state lock whose holder pod was gone. Look for ForceUnlocked events and check why the previous Job died.

Runbook: ../operator-guide/observability.md#captfforceunlocks

CAPTFReconcileErrors

  • Group: captf
  • Severity: warning
  • For: 10m
sum by (controller) (rate(controller_runtime_reconcile_errors_total{controller=~"terraform.*"}[10m])) > 0.1

Summary: The {{ $labels.controller }} controller keeps failing to reconcile

Description: Reconcile errors above 0.1/s for 10 minutes; read the manager logs.

Runbook: ../operator-guide/observability.md#captfreconcileerrors

CAPTFJobSlow

  • Group: captf
  • Severity: info
  • For: 30m
histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_duration_seconds_bucket{result=~"succeeded|failed"}[1h]))) > 1800

Summary: {{ $labels.kind }} {{ $labels.op }} Jobs are slow

Description: The 90th percentile of {{ $labels.op }} Job duration has been above 30 minutes for 30 minutes.

Runbook: ../operator-guide/observability.md#captfjobslow

CAPTFJobQueueSlow

  • Group: captf
  • Severity: warning
  • For: 10m
histogram_quantile(0.9, sum by (le, kind, op) (rate(captf_job_queue_seconds_bucket[10m]))) > 300

Summary: {{ $labels.kind }} {{ $labels.op }} Jobs wait long to start

Description: The 90th percentile of the time {{ $labels.op }} Jobs for {{ $labels.kind }} objects take from creation to running the module has been above 5 minutes. Look for Pending runner pods: node capacity, quota, slow or failing image pulls.

Runbook: ../operator-guide/observability.md#captfjobqueueslow

CAPTFStateNearSecretLimit

  • Group: captf
  • Severity: warning
  • For: 15m
captf_state_bytes > 900 * 1024

Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} state is near the Secret size limit

Description: The compressed state is {{ $value | humanize1024 }}B, above 900 KiB. An OpenTofu state Secret cannot grow past 1 MiB, and the next apply that crosses it fails to save state.

Runbook: ../operator-guide/observability.md#captfstatenearsecretlimit

CAPTFInputsNearLimit

  • Group: captf
  • Severity: warning
  • For: 15m
captf_inputs_bytes > 900000

Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} inputs are near the size limit

Description: The rendered main.tf.json and terraform.tfvars.json are {{ $value | humanize }} bytes. Above 1000000 no Job starts (ApplyJobSucceeded InputsTooLarge).

Runbook: ../operator-guide/observability.md#captfinputsnearlimit

CAPTFNoRecentSuccess

  • Group: captf
  • Severity: warning
  • For: 30m
time() - captf_last_success_timestamp_seconds{op=~"drift|refresh"} > 6 * 3600

Summary: {{ $labels.kind }} {{ $labels.namespace }}/{{ $labels.name }} has had no successful {{ $labels.op }} for 6 hours

Description: Its {{ $labels.op }} Jobs keep failing or never start, so drift and health go unobserved. See DriftJobSucceeded and status.lastRun.

Runbook: ../operator-guide/observability.md#captfnorecentsuccess