Skip to content

Requeue Intervals and Schedules

The controller is event-driven first: a change to the object, a Job finishing or an annotation triggers a reconcile. Requeues are the fallback for time and for waits that have no event. This page lists every interval, its default, and where it comes from. All of them are compiled in unless a field or flag is named.

Schedules (when a Job is due)

Schedule Default Set by Notes
Drift check 30 min spec.drift.intervalSeconds; else --drift-default-interval 0 disables it for a cluster or machine; a pool’s drift is never disabled
Health check 300 s remediation.healthCheckIntervalSeconds Only while remediation.annotateMachine is true; otherwise health is sampled at the drift or refresh cadence
Membership refresh 60 s A pool’s membershipRefreshIntervalSeconds Pools only
Pending-health refresh 30 s, doubling to 5 min Compiled in Doubles per consecutive pending reading: 30 s, 1 min, 2 min, 4 min, 5 min
Converging-membership refresh 30 s Compiled in Fixed, not doubled: members join on the provider’s schedule
Refresh after apply Once, at once Compiled in Machines and pools; skipped when the apply’s own reading is definite

Each schedule’s deadline adds a jitter of up to a tenth of its interval. It is deterministic: derived from a hash of the object’s UID, so the same object always gets the same offset, and objects created together (a MachineDeployment’s machines, or everything after a clusterctl move) do not check in lockstep. An object with no UID gets none.

Waits (when the controller looks again)

Interval Value Used while
GateRequeue 30 s An owner reference, a dependency, credentials, deletion order or a lease is not ready
StateRequeue 1 min The state is lost or unreadable, including a held deletion
LagRequeue 5 s The Job cache trails the API server, or a live Job holds the run lease at cleanup, or an annotation write conflicted
ActiveJobRequeue 1 min A Job runs; the Job watch normally wakes the reconcile first
RetryBase and RetryMax 1 min, 10 min The backoff after a failed Job; see Retries
RetryMax 10 min An approval wait, JobPolicyInvalid, InputsTooLarge, or missing durable inputs: nothing to do until something changes, and the change wakes the reconcile
Stuck-Job age 1 min A Job younger than this is not checked for a missing per-run Secret
Lease grace 1 min A lease whose holder Job does not exist stays held this long
Identity re-read 5 min A TerraformClusterIdentity re-reads its credentials Secret

The resync

--sync-period (default 10 minutes) is the minimum interval at which the informers re-enqueue every cached object. It reads the local cache, not the API server, and sets no schedule of drift or health. It recovers a requeue that was missed, at the cost of more reconciles. The orphan sweep ticks at the same interval, with 10% jitter, on the leader. See Configuration.

How they combine

A reconcile with nothing to run requeues at the soonest of the due times: the next refresh, health, membership or drift deadline, whichever applies. A reconcile that waits on something requeues at that wait’s interval. A failed op’s backoff replaces its Job with a requeue for the remaining delay, and pauses nothing else: another op still runs on its own schedule. A waiting input change is the exception: it pauses the refresh and drift schedule, because they would render the unapplied inputs and report the change as drift.

Object state Typical next reconcile
Idle, healthy The sooner of the drift and health deadlines, and any event
A Job running The Job’s finish event; a minute at the latest
Backing off after a failure The end of the backoff
Waiting for a lease 30 s
Waiting for an approval The annotation, or 10 min
Held on the state 1 min

See also