Skip to content

Secret Management

CAPTF keeps everything it needs between Jobs in Kubernetes Secrets and Leases: the Terraform state, the backups of it, the rendered inputs of each run, the cloud credentials, and a key for the plan approval. This chapter follows each of them from creation to deletion. It explains how they connect, how they reach the Terraform runtime inside a Job, and what an operator needs to provide, protect and back up.

The chapter has these pages:

  • Terraform State Secrets


    Storage, naming, chunking, the integrity checks, the inputs hash and adoption.

    Terraform State Secrets

  • Backups and Restore


    When a backup is taken, how many are kept, and how a restore runs.

    Backups and Restore

  • Credentials


    An identity’s source Secret, the per-namespace mirror, rotation and revocation.

    Credentials

  • Run Inputs and the Plan Key


    The durable and per-run inputs Secrets and the key behind the plan hash.

    Run Inputs and the Plan Key

  • Inside the Job


    Volumes, environment, RBAC, the runner’s steps and lock handling.

    Inside the Job

  • Lifecycle Walkthroughs


    Create, apply, drift, change, delete, clusterctl move, namespace deletion and lost state.

    Lifecycle Walkthroughs

  • Operator Files and Settings


    The manifests, flags and backups you own.

    Operator Files and Settings

  • Security Considerations


    The trust boundaries, stated plainly.

    Security Considerations

This page is the inventory of every Secret and Lease, and how they relate.

Existing pages cover parts of this ground from other angles, and are linked rather than repeated: Terraform State, Job Inputs, Identities and Credentials, Plan Approval, the Security Model, Secrets (which also lists the Secrets CAPTF only reads, such as bootstrap data) and the runbooks.

The objects at a glance

<suffix> is the state suffix: the first 16 hex characters of sha256(<namespace>/<kind>/<name>), a hyphen and c, m or mp (see Terraform state). <kindshort> is c, m or mp for a TerraformCluster, TerraformMachine or TerraformMachinePool. “The object” below is that Terraform* object.

Name pattern Kind Created by Owner Contents Moves with clusterctl move Deleted when
tfstate-default-<suffix> and -part-N Secret The Terraform or OpenTofu kubernetes backend, inside the Job The object, as a non-blocking owner reference, set after the first successful apply. A chunk without one (written by a refresh or drift Job, or restored without references) or naming an earlier UID is owned again on the next reconcile Gzipped state under the key tfstate Yes: it carries the move label Cleanup after a destroy, a delete with no state, or an abandon
lock-tfstate-default-<suffix> Lease The backend, when a run takes the lock None The lock holder’s information No: a Lease is not discovered; the backend recreates it Cleanup, with the state
captf-state-backup-<suffix>-<serial> and -part-N Secret The manager, after it sees a new serial The object (non-controller reference); re-owned after a restore A verbatim copy of the state chunks Yes: it carries the move label With the object, by garbage collection; pruned beyond --state-backups
captf-inputs-<kindshort>-<name> Secret The manager, when an apply Job starts The object; re-owned after a restore The rendered root module and tfvars, plus the image, digest, identity and applied marker Yes, by following the owner reference Cleanup
captf-run-<job> Secret The manager, right after it creates the Job The Job The same two rendered files Not applicable: it lives only while the Job does The first time the controller sees the Job finished
captf-plankey-<kindshort>-<name> Secret The manager, before a plan or approved apply Job The object; re-owned after a restore 32 random bytes under the key key Yes, by following the owner reference Cleanup
captf-creds-<identity> (the mirror) Secret The manager, on a reconcile of an object that uses the identity Each object that uses it (non-controller references); a using object’s reference to its earlier UID is replaced A copy of the identity’s source Secret data Yes, and it is rewritten from the source on the target When its last user is gone, or when the namespace stops being allowed
The identity’s source Secret (any name) Secret The operator None Cloud credentials No: copy it to the target yourself By the operator
captf-run-<suffix> Lease The manager, before it starts a Job None The run lease: one Job at a time per object No: a Lease is not discovered When the Job finishes, at cleanup, and by the namespace sweep
captf-cluster-<hash> Lease The manager, for a TerraformCluster’s apply, destroy or restore None The cluster write lease: machines and pools wait on it No When the Job finishes, at cleanup, and by the namespace sweep

The labels and annotations on each are in Annotations, Labels and Finalizers; this chapter names the ones that matter to the behavior it describes. Every Secret the manager creates carries captf.io/managed=true, which is also what its cache selects on (see What the manager caches).

How they connect

flowchart LR
    subgraph ops[Operator]
        SRC["Identity source Secret"]
        ID["TerraformClusterIdentity"]
    end
    subgraph mgr[Manager]
        MIR["captf-creds-identity<br/>(mirror)"]
        DUR["captf-inputs-kind-name<br/>(durable inputs)"]
        RUN["captf-run-job<br/>(per-run inputs)"]
        KEY["captf-plankey-kind-name"]
        BAK["captf-state-backup-...<br/>(backups)"]
    end
    subgraph job[Job pod]
        R["Runner"]
        TF["Terraform or OpenTofu"]
    end
    ST["tfstate-default-suffix<br/>(state)"]
    LK["lock-tfstate-default-suffix<br/>(Lease)"]

    ID --> SRC
    SRC -->|"copied on every reconcile"| MIR
    MIR -->|"envFrom and files"| R
    DUR -->|"copied at Job start"| RUN
    RUN -->|"/captf/config"| R
    KEY -->|"/captf/plan-key"| R
    R --> TF
    TF -->|"reads and writes"| ST
    TF -->|"holds"| LK
    ST -->|"new serial seen"| BAK
    BAK -.->|"restore Job"| ST

The flow, in words. An operator creates the identity’s source Secret and the TerraformClusterIdentity that names it. When an object runs, the manager mirrors the source into the object’s namespace and renders the object’s inputs into the durable Secret, then copies them into a per-run Secret that belongs to one Job. The Job’s runner receives the per-run Secret, the mirror and, for a plan, the plan key; it runs Terraform, which reads and writes the state Secrets and holds the lock Lease through the cluster API. The manager reads the state back, takes a backup when the serial is new, and records the outcome in status.

Two rules explain most of what follows:

  • State is the source of truth for what exists. Status is rebuilt from state, which is why status is not restored by clusterctl move, and why a marker on the durable inputs Secret is needed to remember that an object ever applied.
  • Everything a Job reads is a copy. The per-run Secret is written once and never rewritten, so a Job never sees inputs change under it, and the mirror is rewritten only by the manager, never by the Job.