# Cluster API Terraform > Cluster API Provider Terraform. Your modules are the provider. CAPTF (Cluster API Provider Terraform) is a Cluster API infrastructure provider. Instead of cloud-specific Go controllers, it provisions cluster infrastructure by running Terraform or OpenTofu modules, packaged as OCI images, as Kubernetes Jobs: "your modules are the provider". How it works: Cluster API's `Cluster`, `Machine` and `MachinePool` reference CAPTF's `TerraformCluster`, `TerraformMachine` and `TerraformMachinePool`. The CAPTF manager renders each object's inputs, runs the module image in a runner Job, keeps Terraform state in a Kubernetes Secret next to the object, and reads the module's outputs back from that state. A module is written to the `v1alpha1` module contract and linted with `tfcapi-lint`. Status: pre-release. Every kind and the module contract are `v1alpha1`; the contract is frozen for implementation but may change. There are no end-to-end tests yet, and the reference cloud modules (AWS, Google Cloud, Azure, OCI, OpenStack) have not yet been applied to a real cloud. Using this index: the docs are organized by reader. Overview and Start here for evaluators; User Guide and Cloud Modules for people running clusters; Module Authors for people writing modules (the contract is normative); Operations and Troubleshooting for people running the manager; How It Works for internals; Reference for generated API, conditions, events, metrics, flags and environment. Every page is also available as Markdown at its URL followed by `index.md`, and https://captf.io/llms-full.txt is the whole site in one file. News: https://captf.io/feed_rss_created.xml. ## Start here - [CAPTF](): Cluster API Provider Terraform: turn the Terraform and OpenTofu modules you already trust into Kubernetes clusters, on any platform. - [Introduction](): Cluster API Provider Terraform runs your Terraform or OpenTofu modules as a Cluster API infrastructure provider. - [Quick Start](): Install CAPTF on a management cluster and bring up a TerraformCluster and control-plane TerraformMachine with the no-op modules. ## Overview - [Architecture](): The manager, webhooks, runner Jobs, module image and state backend, and how one apply flows through them. - [The Kinds](): How CAPTF's seven kinds relate to Cluster API's objects, what stays fixed once an object exists, and what a cluster passes down. - [Glossary](): Short definitions of CAPTF terms and the Cluster API and Terraform terms its docs assume, each linked to the page that covers it. - [Security Model](): What creating a Terraform* object grants, what its Job can read, and what CAPTF keeps out of status, events and logs. - [Known Limitations](): What CAPTF does not do yet or does with a catch, with the workaround and the page that has the detail. - [Compatibility](): The versions CAPTF is built against, what the project tests, and what it only assumes. ## User Guide - [Identities and Credentials](): Create a credentials Secret and a TerraformClusterIdentity, choose the namespaces allowed to use it, reference it, rotate and revoke access. - [Module Variables](): Pass your own module variables inline or from labeled ConfigMaps and Secrets, and understand merge order, limits and sensitivity. - [Templates and ClusterClass](): Generate a cluster from the shipped clusterctl templates, understand the noop ClusterClass, and patch module variables through topology. - [Machine Pools](): Create a MachinePool backed by a TerraformMachinePool, choose fixed or autoscaled replicas, check members, and delete the group. - [Tuning Jobs](): Tune a Job's resources, deadlines, history, environment, pull secrets, ServiceAccount and security contexts, and how machines inherit cluster defaults. - [How Drift and Health Work](): How CAPTF schedules drift checks, turns module health into InfrastructureHealthy and Ready, and counts unhealthy samples. - [Configure Drift Detection](): Set how often CAPTF checks for drift, whether it reports or remediates, and read the results. - [Machine Remediation](): Ask Cluster API to replace an unhealthy TerraformMachine through a MachineHealthCheck, tune the threshold, and confirm the request. - [Approvals and Gates](): The two approval gates on a TerraformCluster, what each stops and binds, what is not gated, and where to go next. - [Approve a Plan](): Review and approve a destructive plan or, under applyPolicy Manual, every plan, with the exact kubectl commands and caveats. - [Manual Plan Approval](): How applyPolicy Manual plans every change first and applies it only after a person approves the plan hash, and what happens when the plan changes. - [What the Plan Hash Binds](): What goes into the p2: plan hash an approval names, what it binds and what it leaves out, so you know what an approval covers. - [The Destructive-Plan Guard](): How the always-on destructive-plan guard blocks a TerraformCluster apply that deletes or replaces a resource until you approve its inputs hash. - [What Approval Does Not Guarantee](): What a plan approval does and does not promise: apply-time behavior, the plan key, ungated kinds, cluster outputs and state surgery. - [Deletion and Teardown](): How deleting a TerraformCluster, machine or pool runs a destroy Job, what holds a deletion, and what is left behind if you short-circuit it. - [Order and Finalizers](): The fixed order in which Cluster API, the admission webhook and the controller delete machines, pools and clusters, and when the finalizer comes off. - [The Destroy Job](): What runs when a deleting object has readable state, what does not gate the destroy, what it waits on, and how it fails and retries. - [Terminating Namespaces](): What a deletion can still finish in a terminating namespace, why a destroy waits there, and what the namespace takes with it. - [Cleanup and Garbage Collection](): Which of the controller's cleanup, Kubernetes garbage collection and the namespace RBAC sweep removes what, and what nothing deletes. ## Cloud Modules - [Cloud Modules](): Reference Terraform and OpenTofu module sets for AWS, Google Cloud, Azure, OCI and OpenStack, and a no-op set, with their images, tags and conventions. - [Shared Behavior](): What all five cloud module sets have in common: the API endpoint, traffic rules, node identities, bootstrap data, health, tags and destroy. - [AWS](): Build a workload cluster in an AWS VPC you bring: the images, prerequisites, identity Secret, quick start, API endpoint, exports and tags. - [Cluster](): Look up what the AWS cluster module creates, its inputs, outputs, health reporting and limits, with an example. - [Machine](): Look up what the AWS machine module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [MachinePool](): Look up what the AWS machinepool module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [Google Cloud](): Build a workload cluster in a Google Cloud VPC network you bring: the images, prerequisites, identity Secret, quick start, endpoint, exports and labels. - [Cluster](): Look up what the Google Cloud cluster module creates, its inputs, outputs, health reporting and limits, with an example. - [Machine](): Look up what the Google Cloud machine module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [MachinePool](): Look up what the Google Cloud machinepool module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [Azure](): Build a workload cluster in an Azure virtual network you bring: the images, prerequisites, identity Secret, quick start, endpoint, exports and tags. - [Cluster](): Look up what the Azure cluster module creates, its inputs, outputs, health reporting and limits, with an example. - [Machine](): Look up what the Azure machine module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [MachinePool](): Look up what the Azure machinepool module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [OCI](): Build a workload cluster on an OCI VCN you bring: the images, prerequisites, identity Secret, quick start, API endpoint, exports and tags. - [Cluster](): Look up what the OCI cluster module creates, its inputs, outputs, health reporting and limits, with an example. - [Machine](): Look up what the OCI machine module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [MachinePool](): Look up what the OCI machinepool module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [OpenStack](): Build a workload cluster on an OpenStack network you bring: the images, prerequisites, identity Secret, quick start, endpoint, exports and tags. - [Cluster](): Look up what the OpenStack cluster module creates, its inputs, outputs, health reporting and limits, with an example. - [Machine](): Look up what the OpenStack machine module creates, its inputs, outputs, health, lifecycle and limits, with an example. - [No-op](): Try CAPTF or test a management cluster with no cloud account: the no-op module images, what they return, and a quick start. - [Cluster](): The no-op cluster role: the stand-in resource it records, the contract inputs it reads, and the endpoint, failure domain and exports it returns. - [Machine](): The no-op machine role: the stand-in instance it records, the contract inputs it reads, its provider ID, address and image labels. - [MachinePool](): The no-op machinepool role: the stand-in scaling group it records, the contract inputs it reads, and the provider IDs and instances it returns. ## Module Authors - [Your First Module](): Write, lint, package and reference a small machine-role module from scratch, ending with a lint-clean module image. - [How Inputs Reach a Module](): Where every Terraform input of a CAPTF Job comes from: the pipeline, per-role inputs, exports, user variables, the inputs hash and safe inspection. - [v1alpha1](): Overview of the v1alpha1 module contract: child-module rules, naming, null semantics, type conventions, injection mechanics and role summary. - [Common](): Inputs injected into every role, the user-variable rules, and the health output every module must declare. - [Cluster Role](): The cluster role of the v1alpha1 module contract: inputs, outputs, provisioned rule, lifecycle, control-plane requirements and a skeleton. - [Machine Role](): The machine role of the v1alpha1 module contract: inputs, outputs, Ready timeline, lifecycle, providerID matching and a skeleton. - [MachinePool Role](): The machinepool role of the v1alpha1 module contract: inputs, autoscaling, outputs, lifecycle, bootstrap rotation and a skeleton. - [Image Contract](): The normative image contract: fixed paths, OCI labels, user, multi-arch publishing, and reference Containerfiles for Terraform and OpenTofu. - [Changelog](): History of changes to the v1alpha1 module contract, newest first. - [Runtime Environment](): What a module sees when the runner executes it: working directory, environment, provider mirror, command order, failure reporting and size limits. - [Module Design Patterns](): Design modules that behave well under CAPTF: where resources belong, stable exports, health, autoscaling, import, destroy safety, secrets and testing. - [Linting with tfcapi-lint](): Install tfcapi-lint, lint a module or image against the CAPTF contracts, use strict mode and exit codes, and run it in CI. - [Control-Plane Integration](): How CAPTF cluster and machine modules integrate with the KubeadmControlPlane and RKE2ControlPlane providers, and the load-balancer patterns that fit. - [KubeadmControlPlane](): What KubeadmControlPlane and CABPK require of your CAPTF cluster and machine modules: endpoint, load balancer, lifecycle and a module checklist. - [RKE2ControlPlane](): What RKE2ControlPlane and CAPRKE2 require of your CAPTF cluster and machine modules: listeners, addresses, lifecycle and a module checklist. - [Requirements Checklist](): Every port, health check and ordering rule a KCP or RCP cluster needs from a CAPTF cluster or machine module, grouped by topic. ## Operations - [Installation](): Install the CAPTF provider into a management cluster with clusterctl, set the runner image, add optional components and verify the install. - [Configuration](): Change the CAPTF manager's flags and read what each group does: scoping, concurrency, leader election, sync period, runner image, backups and logging. - [Upgrades](): Upgrade the CAPTF provider with clusterctl, understand what an upgrade triggers on existing objects, and roll back. - [RBAC](): Read every role CAPTF ships: what the manager and the runner can do, how a namespace gets its runner identity, and how the orphan sweep removes it. - [Secrets Inventory](): Inventory every Secret CAPTF reads or writes: name, namespace, owner, lifecycle, contents and sensitivity, plus what the manager caches and logs. - [Multi-Tenancy](): Lay out tenants on one management cluster: namespace isolation, identities, the runner's Secret access, quotas, network policy and an example. - [Production Readiness](): Check the CAPTF manager and its namespaces before go-live: availability, sizing, certificates, secrets, alerts, approvals, network policy and Jobs. - [Observability](): Read the manager's metrics endpoint, enable Prometheus and the alerts, find what to check first for each alert, and read conditions, events and logs. - [Operating the Gates](): Run the plan-approval gates day to day: what the conditions and events mean, how retries and upgrades behave, and who can approve. - [Other Manual Actions](): The manual actions besides plan approval: restore state, abandon an object, fix a Job policy, release a foreign lock and opt a ServiceAccount in. - [Disaster Recovery](): Recover from lost Terraform state, a lost namespace or a lost management cluster: what to back up outside the cluster, how to rebuild and how to rehearse. - [Restore State from a Backup](): Restore an object's Terraform state from a versioned backup, or rebuild it by hand when no backup exists, and confirm the restore. - [Total State Loss and Import](): Recover an object whose Terraform state is lost with no backup: rebuild the state from a workstation, or abandon and recreate it with import blocks. - [Move Clusters with clusterctl](): Move a Cluster's Terraform objects to another management cluster with clusterctl move, copy what does not move, and clean up the source. - [What clusterctl move Deletes](): Understand how clusterctl move deletes the source objects, what moves and what does not, and what the target does on its first reconcile. ## Troubleshooting - [Troubleshooting by Condition](): Find the condition, event or runbook for a stuck CAPTF object, and learn how Ready is built from other conditions. - [Runbooks by Symptom](): Pick the runbook for a symptom, alert or condition reason, from failing Jobs to lost state. - [Nothing Is Happening](): Find what a CAPTF object with no Job is waiting on, from its conditions and events, and what to do about each wait. - [Every Condition](): Look up every CAPTF condition type, status and reason with its meaning, likely cause, action and a link to the deeper page. - [Every Event](): Read the Warning and informational events CAPTF emits, with what each means and what to do about it. - [Failing Jobs](): Diagnose a failing apply, destroy, drift check or refresh Job from the object's conditions, status and pod logs, and fix the cause. - [Slow Jobs](): Find out why Jobs run long or start late, from the CAPTFJobSlow and CAPTFJobQueueSlow alerts, leases, step durations and pod events. - [Reconcile Errors](): Diagnose the CAPTFReconcileErrors alert: read the manager's logs, narrow down the object and match the error to a common cause. - [Unreadable State](): Read the StateReadable reason and fix lost, locked, encrypted, corrupt or inconsistent state, including held deletions and abandon. - [Stale State Lock](): Find a held state lock, tell a stale one from a live one and clear it with force-unlock. - [Size Limits](): Find the object whose state or rendered inputs approach the Secret size cap, and shrink what it stores. - [My Object Will Not Delete](): Work out why a deleting CAPTF object keeps its finalizer, from its conditions, and follow the flowchart or table to the fix. - [Held Deletions](): Understand why CAPTF holds a deletion when state is lost or unreadable, what ends the hold, and how the abandon annotation works. - [Stuck Destroy](): Recover from a destroy Job that cannot succeed: back up state, clean up cloud resources, then remove or abandon the finalizer. - [Stripping a Finalizer by Hand](): Know what removing a CAPTF finalizer by hand deletes or leaves behind, and what to preserve before you do it. - [Identities and Credentials](): Fix an object whose identity, credential mirror or runner RBAC is not ready, reason by reason. - [Webhook Unavailable](): Recognize writes blocked by an unreachable admission webhook and get it serving again: pods, Service, NetworkPolicy and certificate. ## How It Works - [The Reconcile Lifecycle](): Follow one reconcile pass: preamble, bookkeeping, choosing the operation, retry backoff, leases, deletion order and the clusterctl move block. - [Terraform State](): Learn where CAPTF keeps Terraform state, how it names, reads, locks and backs it up, and what happens to it on deletion. - [Secret Management](): Inventory every Secret and Lease CAPTF keeps between Jobs, see how they connect, and find the page for each part. - [Terraform State Secrets](): See how state Secrets are named, labeled, chunked, read, checked and attached to the object by the manager. - [Backups and Restore](): Learn when the manager backs up state, how backups are named and pruned, and how a restore Job pushes one back. - [Credentials](): Follow a credential from the source Secret through the per-namespace mirror to the Job, including rotation and revocation. - [Run Inputs and the Plan Key](): Learn where rendered inputs are kept, how long, and what the plan key is for. - [Inside the Job](): See what a Job pod mounts, which environment and permissions it gets, and which runner steps run for each operation. - [Lifecycle Walkthroughs](): Follow the Secrets and Leases through create, refresh, change, delete, clusterctl move, namespace deletion and lost state. - [Operator Files and Settings](): List what an operator provides or keeps right: credentials, runner ClusterRole, manager environment, encryption and backups. - [Security Considerations](): Understand what CAPTF's handling of Secrets does and does not protect, and where the trust boundaries lie. - [Jobs, Retries and Concurrency](): Follow a Job from the decision to run it to bookkeeping, and find the pages on naming, retries, leases and failover. - [Job Names, Attempts and History](): Learn how deterministic Job names make creation idempotent, how attempts count, and how finished Jobs are kept and pruned. - [Choosing the Operation](): Read the priority order CAPTF uses to pick the next operation, the reasons it reports, and the refresh and drift schedule. - [Retries and Backoff](): Learn what counts as a failed Job, how the retry delay grows, and which waits deliberately do not back off. - [Deadlines and Lock Timeouts](): Understand the two clocks on a Job, activeDeadlineSeconds and lockTimeoutSeconds, and the check that keeps them consistent. - [Leases and the Operation Gate](): Learn how the run lease and cluster write lease decide who may start a Job, and what each wait reason means. - [Cache Lag](): See how the controller checks the API server directly so a stale Job cache never clears a block or drops a finalizer. - [The State Lock and Stale Locks](): Tell CAPTF's leases from Terraform's state lock, and see how stale and foreign locks are handled. - [Leader Election and Failover](): Configure leader election and see what a manager failover does to Jobs in flight. ## Reference - [Custom Resources](): Every CAPTF custom resource, one page each: what it is, YAML examples, and every spec and status field with its default, validation and behavior. - [TerraformCluster](): Field-by-field reference for the TerraformCluster kind: spec, status, conditions, printer columns, admission rules and lifecycle. - [TerraformClusterTemplate](): Field-by-field reference for the TerraformClusterTemplate kind: the template a ClusterClass uses to create a TerraformCluster, with its immutability rules. - [TerraformMachine](): Reference for the TerraformMachine kind: every spec and status field, conditions, printer columns, validation rules and lifecycle. - [TerraformMachineTemplate](): Reference for the TerraformMachineTemplate kind: the template Cluster API clones into TerraformMachines, and the scale-from-zero status read from the image. - [TerraformMachinePool](): Every field, status value, condition, printer column and admission rule of the TerraformMachinePool kind, with examples. - [TerraformMachinePoolTemplate](): Every field, printer column and admission rule of the TerraformMachinePoolTemplate kind, the ClusterClass template for TerraformMachinePool. - [TerraformClusterIdentity](): Reference for the cluster-scoped TerraformClusterIdentity kind: its credentials Secret, allowed namespaces, status, conditions, validation and lifecycle. - [Common Fields](): Reference for the spec and status fields that TerraformCluster, TerraformMachine and TerraformMachinePool share: source, identity, Job policy, variables and run status. - [Annotations, Labels and Finalizers](): Every annotation, label and finalizer CAPTF reads or sets: the keys you set to approve, restore or abandon, CAPTF's own bookkeeping keys and the captf_tags modules receive. - [clusterctl Variables](): Every variable the CAPTF clusterctl templates read, with required and default values, grouped by flavor, plus the ClusterClass topology variables. - [Conditions](): Every status condition CAPTF sets, which kinds carry it, how to read its polarity, and what each reason means and what to do about it. - [Events](): Look up every Kubernetes Event the CAPTF manager and runner record: reason, type, object, when it fires and what to do. - [Alerts](): Look up the eleven Prometheus alerts CAPTF ships: severity, PromQL expression, what each means, likely causes and where to look first. - [Metrics](): Look up every Prometheus metric the CAPTF manager exposes: type, labels and their values, unit and meaning, grouped by topic with example queries. - [Manager Flags](): Every flag of the CAPTF manager, grouped by purpose, with type, default, what it does and when to change it, plus how to set flags on the Deployment. - [Requeue Intervals](): Look up every requeue interval and schedule the CAPTF controller uses, with defaults and where each is set. - [Job Environment](): Anatomy of the runner Job: containers, environment, credentials, volumes, resources, security contexts, runner arguments and the names spec.jobs.env rejects. - [Runner CLI](): Look up the runner binary's subcommands, flags, exit codes and result error kinds. The runner is the entrypoint of every CAPTF Job. - [tfcapi-lint CLI](): Look up the tfcapi-lint commands, flags, exit codes and every check it runs against a CAPTF module or image, with severity and fix. - [Tags](): Every page in the book by tag: who it is for, what kind of page it is, and its topic. ## Contributing - [Contributing](): Build, lint, test and verify CAPTF itself: repository layout, prerequisites, conventions and running the manager against a real cluster. - [Testing](): What make test runs, the unit and e2e tiers, coverage floors, golden files, the test environment, and how to run one package or one test. - [Releasing](): Cut a CAPTF release: what it consists of, the checklist, and installing the assets from a local repository. - [Contributing to the Docs](): Change the CAPTF docs or website: set up a preview, edit or add a page, run the checks and open a pull request. - [Writing Documentation](): How the CAPTF docs are organized and written: where pages go, reference pages, page shape, voice, Markdown patterns and the checks to run. - [Updating the Website](): Change captf.io beyond the docs pages: the landing page, blog posts, navigation and tabs, header, footer and banner, generated files and the theme. - [Make Targets](): Look up the provider repository's make targets by workflow, what each does and when to run it, and the variables that change them. - [Third-Party Licenses](): See which linked third-party modules are not Apache-2.0 licensed, with versions and licenses. ## Blog - [Reference modules for five clouds](): Reference module sets for AWS, Google Cloud, Azure, OCI and OpenStack: what each creates, and their pre-release status. - [The CAPTF book is online](): The CAPTF documentation is published at captf.io/docs: concepts, guides, the module contract, runbooks and generated reference.