Field Manual

The Governor — a fleet-wide concurrency ceiling

Why an autonomous agent fleet needs one machine-wide throttle in front of the model — and how to build one that can't wedge the work it protects.

Why this exists. Several token burns once shut an account down. The cause was not total tokens spent — it was aggregate concurrency. Too many API requests in flight at the same instant tripped the provider’s org-level rate limit, which cascaded into a shutdown. The Governor is the one ceiling that prevents it.

The one rule

Nothing fans out to the model without passing through a ceiling. There are two ceilings, one per kind of caller:

Caller Ceiling
Headless one-shot model calls, scripts, cron jobs The Governor — a shared library every spawn wraps itself in
An interactive session firing parallel agent tools A hard fan-out cap (a small number, e.g. ≤3) plus a bounded workflow pool

If you remember one thing: burn the budget with depth, not width. To spend a big token budget, go deep serially or through a small capped pool. A wide burst of parallel agents is what triggers a shutdown, every time.

What the Governor is

A single machine-wide counting semaphore. Every headless API spawn claims one of a fixed number of slots before running and frees it after. A burst of 20 callers becomes a steady stream of N-at-a-time. When a real rate-limit reply (an HTTP 429) is seen, it trips a cooldown that backs the whole machine off — not just the one caller that hit the wall.

It is deliberately small and boring:

Using it

The interface is a thin gov verb over the semaphore, usable from a shell or a crontab:

gov run -- <command>       # acquire → run → release (the main pattern)
gov claude -p "…"          # shorthand for a headless model call
gov wrap '<command>'       # wrap a whole cron line
gov status                 # slots in use / max / cooldown state
gov cooldown 90            # manually back the whole machine off 90s

Tuning

Three knobs, all illustrative — start conservative and raise only after a clean run shows headroom:

Knob Default Meaning
max concurrent 4 max concurrent API spawns on this machine
wait 120s seconds to wait for a slot before failing open
ttl 900s seconds before a held slot is reaped as stale

Governing the central dispatch engine (the keystone)

Most fleet apps don’t call the model directly — they route through a central model-dispatch engine (a small loopback server). That engine already gated on budget (rolling account-usage percentage) but not on in-flight count — exactly the gap that tripped the org rate limit.

So the keystone change governs the engine itself: its single dispatch function now wraps its one model shell-out in a governor slot, the same semaphore as the shell side. One change, and every engine-routed app is governed centrally. It is fail-open at the import level (if the governor can’t load, the engine runs exactly as before) and priority-aware (a realtime request waits only a few seconds then fails open for low latency; batch and standard work wait longer to avoid bursting).

The practical gotcha worth generalizing: a long-running daemon keeps the old code until you restart it. Ship the file change, then restart the engine to activate — the change is safe and reversible; the restart is the operator action.

What to govern, and what to leave (a coverage map)

The reasoning that decides where a governor slot actually earns its keep, in priority order.

Govern first — the concurrent-burst moments:

Leave (deliberately):

The transferable rule: a governor slot buys you the most exactly where callers stampede together, and buys you nothing on the paths that only ever fire one-at-a-time. Spend your integration effort accordingly.

The drop-in itself is mechanical — a shell call site swaps command for gov run -- command; a Python call site wraps the model subprocess in a with governor.slot(): … context (adding a short wait for interactive callers).

Cross-machine cooldown (built, opt-in)

A per-machine ceiling can’t stop a fleet-wide burst: the provider’s org rate limit is shared across every machine, so several machines each respecting their own N slots still sum to N × machines concurrent. Because the shutdown is an org-level event, a 429 on any machine should back off the whole fleet.

How it works, cheaply — no hot-path latency:

The whole cross-machine path is opt-in and fail-open until switched on: with the feature flag unset, no network calls happen; with the flag set but the shared table absent, reads return “no active cooldown” and writes fail harmlessly.

Tested

The semaphore is verified: a 6-caller burst against a max of 2 peaks at exactly 2; the cooldown counts down; a 429 in a command’s stderr auto-trips a machine-wide 90-second cooldown; and it behaves identically under a plain shell and an interactively-sourced one.