Azure¶
The Azure modules run a Kubernetes cluster on Azure virtual machines. The cluster role creates one resource group per cluster, holding the API server’s Standard load balancer, the nodes’ network security groups, application security groups and managed identities with their role assignments. The machine role creates one Linux VM per Machine, and the machinepool role one virtual machine scale set per MachinePool. The source is captf-io/azure-modules. The modules are pre-release; read the status note in Cloud Modules first.
-
Cluster
The API endpoint, security rules, node identities and exports of one workload cluster.
-
Machine
One instance per
Machine, registered with the API load balancer. -
MachinePool
One native scaling group per
MachinePool.
Images¶
| Role | Image | Page |
|---|---|---|
| cluster | ghcr.io/captf-io/azure-cluster | Cluster |
| machine | ghcr.io/captf-io/azure-machine | Machine |
| machinepool | ghcr.io/captf-io/azure-machinepool | MachinePool |
The images carry hashicorp/azurerm 5.7.0, pinned exactly.
Prerequisites¶
- Network. A virtual network with a subnet for the nodes, and optionally a second subnet in the same network for workers, in the subscription the identity works in. Nodes get no public IP, so the subnets need an egress path, such as a NAT gateway or a firewall, to the image registries and Azure’s endpoints. With the default internal API endpoint, the management cluster must reach the control-plane subnet (through peering or a VPN, for example).
- Resource providers registered in the subscription:
Microsoft.Compute,Microsoft.Network,Microsoft.ManagedIdentity,Microsoft.AuthorizationandMicrosoft.Insights(autoscale settings). The modules never register them. - Permissions of the identity’s service principal:
- Contributor on the subscription: the cluster role creates a resource group, and joins the subnets, which must be in the same subscription. Narrow Contributor, and the network’s resource group needs Network Contributor.
- Role Based Access Control Administrator on the subscription, best with a condition that limits it to assigning Contributor, Network Contributor and AcrPull: the cluster role grants the node identities.
- Managed Identity Operator on any node identity you bring, so the machine and pool roles can attach it.
- Node images. A Linux image built for Cluster API, with kubeadm, the kubelet, a container runtime, cloud-init,
curlandiptables. The CAPZ reference images work:/communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver}. - In the workload cluster: cloud-provider-azure (the cloud controller manager and cloud-node-manager), with
--configure-cloud-routes=falsefor an overlay CNI, and the Azure Disk CSI driver for persistent volumes. Both read/etc/kubernetes/azure.json, which the modules write on every node from cloud-config bootstrap data.
Identity Secret¶
The Secret holds the azurerm provider’s environment variables:
| Key | Value |
|---|---|
ARM_TENANT_ID | The Microsoft Entra tenant of the service principal |
ARM_SUBSCRIPTION_ID | The subscription the cluster and its network live in |
ARM_CLIENT_ID | The service principal’s application (client) ID |
ARM_CLIENT_SECRET | Its client secret |
ARM_USE_CLI | "false": the images have no Azure CLI |
For a client certificate, set ARM_CLIENT_CERTIFICATE_PATH to a file under /var/run/captf/credentials/ and ARM_CLIENT_CERTIFICATE_PASSWORD instead of ARM_CLIENT_SECRET. From examples/identity.yaml:
apiVersion: v1
kind: Secret
metadata:
name: azure
namespace: captf-system
type: Opaque
stringData:
ARM_TENANT_ID: 00000000-0000-0000-0000-000000000000
ARM_SUBSCRIPTION_ID: 00000000-0000-0000-0000-000000000000
ARM_CLIENT_ID: 00000000-0000-0000-0000-000000000000
ARM_CLIENT_SECRET: replace-me
ARM_USE_CLI: "false"
---
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha1
kind: TerraformClusterIdentity
metadata:
name: azure
spec:
secretRef:
name: azure
namespace: captf-system
allowedNamespaces:
list:
- team-a
See Identities and Credentials for how the Secret reaches the Jobs.
Quick start¶
-
Create the identity from the repository’s example:
-
Generate the cluster from
examples/cluster-kubeadm.yaml: a KubeadmControlPlane of three, a MachineDeployment of two and a MachinePool of two. Quote the image ID so the shell keeps the braces:export CLUSTER_NAME=demo KUBERNETES_VERSION=v1.34.1 export AZURE_SUBNET_ID=/subscriptions/<id>/resourceGroups/network/providers/Microsoft.Network/virtualNetworks/hub/subnets/nodes export AZURE_SSH_PUBLIC_KEY="$(cat ~/.ssh/id_ed25519.pub)" export AZURE_IMAGE_ID='/communityGalleries/ClusterAPI-f72ceb4f-5159-4c26-a0fe-2ea738f0d019/images/capi-ubun2-2404/versions/{semver}' clusterctl generate yaml --from examples/cluster-kubeadm.yaml | kubectl apply -n "$NAMESPACE" -f - -
Once the first control-plane node answers, install a CNI (for example Calico with VXLAN) and cloud-provider-azure in the workload cluster. cloud-provider-azure sets each Node’s provider ID; only then do the Machines get their Nodes and the workers join.
The KubeadmConfigs in the example name each Node after its VM (nodeRegistration.name: '{{ ds.meta_data["local_hostname"] }}'), which cloud-provider-azure needs to find the VM.
API endpoint¶
- Internal (the default): a Standard load balancer frontend in the control-plane subnet, zone-redundant where the region has zones, with a dynamic address unless
api_load_balancer_private_ipsets one. - Public (
api_load_balancer_public = true): a static Standard public IP. Include the management cluster’s egress and the nodes’ NAT gateway addresses inapi_allowed_cidrs: the load balancer keeps the client’s source address, so the control plane’s security group sees them. - Probes. HTTPS
GET /readyzevery 5 seconds with kubeadm; TCP withdistribution = "rke2", which may disable anonymous access, and on the supervisor port 9345. - Hairpin. An Azure internal load balancer drops a flow from a backend to its own frontend when it maps the flow back to that backend. The machine role therefore installs a small service on control-plane nodes (cloud-config bootstrap data only): while the node’s own API server answers
/readyz, an iptables rule sends the node’s traffic for the frontend to it directly; otherwise the rule is gone and the load balancer takes the traffic to another control-plane node, since this node’s probe is down too. With RKE2 it covers port 9345 as well. Turn it off withapi_server_hairpin_workaround = false; with Ignition, add the equivalent to your bootstrap configuration. - The endpoint guard records
api_load_balancer_public, the endpoint port, the control-plane subnet and the private address the frontend got, and fails any later plan that changes one (see Shared Behavior). Making the current dynamic address static is allowed, since it moves nothing.
Exports¶
The cluster role’s exports carry the schema captf.io/azure-cluster/v1.
| Key | Value |
|---|---|
schema | captf.io/azure-cluster/v1 |
tenant_id | The identity’s tenant |
subscription_id | The subscription, lowercase |
region | The virtual network’s Azure location, such as westeurope |
resource_group_name, resource_group_id | The cluster’s resource group (name lowercase) |
failure_domains | One entry per zone, { "1" = {}, "2" = {}, "3" = {} }; {} without zones |
virtual_network | {id, name, resource_group_name} of the brought network |
subnet_id, subnet_name | The control-plane subnet |
worker_subnet_id, worker_subnet_name | The worker subnet (the control-plane subnet without worker_subnet_id) |
admin_username, admin_ssh_public_key | The nodes’ admin user (captf) and key |
control_plane | {identity_id, identity_client_id, network_security_group_id, network_security_group_name, application_security_group_id, availability_set_id}; the availability set only in a region without zones |
worker | The same keys without availability_set_id |
api | {host, port, backend_port, frontend_ip, backend_pool_id, hairpin_workaround, supervisor_port}; port is the endpoint port, backend_port the kube-apiserver’s; null with a user endpoint |
cloud_provider_config | The cloud-provider-azure configuration (/etc/kubernetes/azure.json) without userAssignedIdentityID, which each node adds for its identity |
Nothing in the exports is secret: the cloud provider authenticates with the nodes’ managed identities.
Tags¶
Azure tag names cannot contain /, so captf.io/<key> becomes captf.io_<key> (see Shared Behavior). Tag names are case-insensitive, so additional_tags rejects any key that starts with captf.io_ or captf.io/ in any case, and takes at most 44 tags: Azure allows 50 per resource. Values longer than 256 characters keep 247, then - and 8 hex characters of their sha256.
Not taggable: role assignments, security rules, load balancer backend pools, probes and rules, NIC associations, the OS disk Azure creates with a VM, and the instances, NICs and OS disks Azure creates for a scale set. All of them live in the cluster’s resource group, which is tagged.
Design notes
- One resource group per cluster, always created. It scopes the node identities’ Contributor rights to what the cluster owns, collects what cloud-provider-azure creates for Services, and makes destroy fail loudly while those are still there instead of deleting disks with data.
- Network security groups on the NICs, not the subnet. The subnet is yours; application security groups name the nodes, so intra-cluster traffic is open whatever the addresses. Workers keep Azure’s default allow-from-network rule, because cloud-provider-azure adds none for internal Service load balancers.
- The module writes
/etc/kubernetes/azure.json. Without it the cloud controller manager never starts and no Machine gets its Node; the derived names it needs are not something you can easily put in a KubeadmConfig. It is written only when absent, so a file of your own wins. - Custom data, never user data. Azure exposes user data to every process on the node through the instance metadata service; custom data it does not.
- Scale sets update their model in place. The provider’s
roll_instances_when_requiredandreimage_on_manual_upgradeare off: with azurerm’s defaults every bootstrap token rotation would reimage every instance. A version change therefore creates a new scale set. - Azure Autoscale holds a pool’s capacity in both modes, because the scale set ignores changes to its instance count, which an autoscaled pool must.
- Trusted launch is off by default: the CAPZ reference images do not support it.
The evidence for each is in DESIGN.md.
Not yet verified
- Whether Azure reports zones on a frontend or public IP differently from what was sent (harmless: the modules ignore later zone changes).
- The minimal role set for cloud-provider-azure and Azure Disk CSI.
- Azure Autoscale holding minimum, maximum and default equal with no rules, including 0, and how long it takes to apply a change.
- The custom data limit of 65,535 decoded bytes (87,380 base64 characters).
- Hairpin through a public load balancer; whether RKE2 needs the hairpin workaround; the DNAT rule alongside every CNI’s iptables rules.
- cloud-init running CABPK’s Jinja template from a
text/plainpart of a multipart message, on the CAPZ images’ cloud-init version. - Azure returning a role assignment’s subnet or registry scope with the casing it was created with.
- cloud-provider-azure with
vmType: "vmss"handling standalone VMs next to scale sets. - The CAPZ gallery publishing an image for every Kubernetes version you roll to.
- A cluster destroy after its network was deleted.
- RKE2’s API server answering
/readyzwith 401 or 403 when anonymous access is off, and whether the 9345 DNAT is needed. - Azure Resource Manager read throttling of pool refreshes beyond about 200 instances.