Skip to content

Reference modules for five clouds

The CAPTF project now maintains reference modules for five clouds: AWS, Google Cloud, Azure, Oracle Cloud Infrastructure (OCI) and OpenStack. Each set implements the v1alpha1 module contract and ships as module images you reference from a TerraformCluster, TerraformMachineTemplate or TerraformMachinePool. Use them as they are, or fork them as the starting point for your own.

Choose a cloud Try the no-op modules

This post was expanded on October 6, 2026. The modules first shipped from one repository per cloud; they now live in one repository per module, and the post describes that layout.

What each cloud gets

Every cloud has a module for each CAPTF role: cluster for the shared infrastructure and the API endpoint, machine for one node per Machine, and machinepool for one scaling group per MachinePool. OpenStack has no machinepool.

Cloud cluster machine machinepool
AWS Network Load Balancer, security groups, IAM roles and instance profiles, an S3 bucket for bootstrap data, in a VPC you bring One EC2 instance One Auto Scaling group and launch template
Google Cloud Internal proxy Network Load Balancer by default, firewall rules, node service accounts One Shielded VM One regional managed instance group
Azure A resource group with a Standard load balancer, network and application security groups, managed identities One Linux VM One uniform virtual machine scale set
OCI Network load balancer, network security groups, a dynamic group and policy so the control plane runs the cloud controller manager and CSI controller as instance principals, on a VCN you bring One compute instance One instance pool, at a fixed size or autoscaled
OpenStack Octavia load balancer, internal by default, security groups, a Nova server group that spreads the control plane One Nova server None
No-op One terraform_data One terraform_data One terraform_data

The no-op set is the sixth, and it is for trying CAPTF without a cloud. Every resource is a terraform_data holding the inputs it was given, so the plans, the state and the outputs are real while no cloud is touched and no credentials are read. A Machine reaches Provisioned but never gets a nodeRef, since no node ever joins.

One set of conventions

All five cloud sets follow the same rules wherever the cloud allows it, so moving between clouds changes the resources, not the behavior. The Shared Behavior page describes them once, and each cloud’s pages describe only what differs:

  • An internal API endpoint


    Internal by default. Making it internet-facing requires a non-empty api_allowed_cidrs; since v0.2.0 the AWS, Azure and Google modules also refuse a /0 in it.

    The API endpoint

  • Traffic rules any CNI can use


    Nodes accept all traffic from each other, scoped to the cluster, and nothing else is open but the API port. SSH stays closed unless you open it.

    Network traffic

  • Node identities


    The cluster role creates the identities the cloud controller manager and CSI driver need, or takes existing ones through variables.

    Node identities

  • Bootstrap data kept out of reach


    Bootstrap data carries the cluster’s keys, so the modules stage it where the cloud allows, and document the clouds where it does not.

    Bootstrap data

  • Health from the cloud


    Every role reports the contract’s health output from cloud state, with a machine-readable reason such as InstanceNotFound.

    Health

  • Pools that stay put


    Bootstrap data rotates in place without replacing instances, a version change rolls the pool, and zones are pinned at the first apply.

    Machine pools

From a module to an image

Each module has its own repository, captf-io/terraform-<provider>-<role>, and its own release cycle: 14 cloud modules and the three no-op ones, 17 in all. A release is a signed tag, and from there two things happen:

flowchart TD
    R["terraform-aws-cluster<br/>signed tag v0.2.0"] --> TR["Terraform Registry<br/>captf-io/cluster/aws"]
    R --> MI["module-images<br/>verifies the tag, builds both runtimes"]
    MI --> I["ghcr.io/captf-io/module-images/aws-cluster<br/>v0.2.0-opentofu, v0.2.0-terraform"]

The Terraform Registry publishes the module, and captf-io/module-images builds its image. That one repository builds all 17 images from the released modules, on Terraform and on OpenTofu, for linux/amd64 and linux/arm64, and holds no module code of its own. It exists as a single repository partly because of a GitHub limit: there is no API to give a repository’s workflows write access to a container package, while packages a workflow creates are writable by that repository’s token. One repository creating every image means no package needs a hand-set permission.

Using a module

Point the CAPTF object at the image, pinned to a release:

apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformCluster
metadata:
  name: "${CLUSTER_NAME}"
spec:
  source:
    image: ghcr.io/captf-io/module-images/aws-cluster:v0.2.0-opentofu

vX.Y.Z-<runtime> is module release vX.Y.Z; a bare <runtime> tag is the newest. Both move to a new digest when the image is rebuilt without a module release, for example for a base image update, so pin a digest in production.

Call it as an ordinary Terraform module:

module "cluster" {
  source  = "captf-io/cluster/aws"
  version = "~> 0.1"

  # The contract inputs the controller would render (captf_contract,
  # captf_cluster, captf_object, captf_tags, ...; see Inputs), and any
  # user variables.
}

Called directly, it is still a CAPTF root module first: it configures its own provider block, pins its providers exactly, and expects you to set the captf_* inputs yourself.

Current versions

Module Version What changed
cluster/aws, cluster/azure, cluster/google v0.2.0 Reject a /0 in api_allowed_cidrs; Azure also denies the kubelet ports to the rest of the VNet on workers
machine/azure, machinepool/azure v0.2.0 Boot diagnostics off by default: the serial console log can include the kubeadm join command
machinepool/aws, machinepool/azure, machinepool/oci v0.1.1 autoscaler = "external", for the Kubernetes Cluster Autoscaler
Every other module v0.1.0 The first release

How they are tested

Each cloud module repository runs the same gate, make verify, on every pull request:

  • formatting, license headers and the module conventions
  • validate on the current Terraform and OpenTofu, and on the oldest supported, Terraform 1.5.7 and OpenTofu 1.6.3
  • unit tests with mocked providers on both runtimes, 47 to 64 run blocks per repository
  • tflint, shellcheck, a Trivy configuration scan, and tfcapi-lint in strict mode against the contract

module-images then smoke-tests every image it builds, and applies and destroys a test root for the no-op images.

Status

None of the cloud modules has been applied to a real cloud yet

They pass static analysis, mocked unit tests on Terraform and OpenTofu, and image smoke tests. Each repository’s DESIGN.md lists the facts the first real apply must confirm, under “Unverified”. Read it before you rely on a module, and pin a release tag or digest.

The cloud module pages also call out individual features that have not been tried yet, so you can tell what is inferred from the cloud’s documentation from what has been seen working.

  • Cloud Modules


    Every set, its images and tags, and what you bring.

    Start here

  • Repositories


    Every repository in the organization, what it publishes, and where.

    See the map

  • Repository Layout


    How a module repository is laid out, and how module-images builds from it.

    Read the page