Why this exists. Several token burns once shut an account down. The cause was not total tokens spent — it was aggregate concurrency. Too many API requests in flight at the same instant tripped the provider’s org-level rate limit, which cascaded into a shutdown. The Governor is the one ceiling that prevents it.
The one rule
Nothing fans out to the model without passing through a ceiling. There are two ceilings, one per kind of caller:
| Caller | Ceiling |
|---|---|
| Headless one-shot model calls, scripts, cron jobs | The Governor — a shared library every spawn wraps itself in |
| An interactive session firing parallel agent tools | A hard fan-out cap (a small number, e.g. ≤3) plus a bounded workflow pool |
If you remember one thing: burn the budget with depth, not width. To spend a big token budget, go deep serially or through a small capped pool. A wide burst of parallel agents is what triggers a shutdown, every time.
What the Governor is
A single machine-wide counting semaphore. Every headless API spawn claims one of a fixed number of slots before running and frees it after. A burst of 20 callers becomes a steady stream of N-at-a-time. When a real rate-limit reply (an HTTP 429) is seen, it trips a cooldown that backs the whole machine off — not just the one caller that hit the wall.
It is deliberately small and boring:
- Portable — no OS-specific file locking; slots are claimed via the atomicity of directory creation, which behaves the same everywhere.
- Fail-open — if no slot frees within the wait window, the call runs anyway and logs it. A smoothing governor must never be able to wedge the fleet. Worst case it degrades to pre-Governor behavior, never worse.
- Self-healing — a slot held by a dead process, or older than a time-to-live, is reaped automatically. A hard kill mid-run cannot leak a slot forever.
- Additive — loading the library changes nothing until you call into it.
Using it
The interface is a thin gov verb over the semaphore, usable from a shell or a
crontab:
gov run -- <command> # acquire → run → release (the main pattern)
gov claude -p "…" # shorthand for a headless model call
gov wrap '<command>' # wrap a whole cron line
gov status # slots in use / max / cooldown state
gov cooldown 90 # manually back the whole machine off 90s
Tuning
Three knobs, all illustrative — start conservative and raise only after a clean run shows headroom:
| Knob | Default | Meaning |
|---|---|---|
| max concurrent | 4 | max concurrent API spawns on this machine |
| wait | 120s | seconds to wait for a slot before failing open |
| ttl | 900s | seconds before a held slot is reaped as stale |
Governing the central dispatch engine (the keystone)
Most fleet apps don’t call the model directly — they route through a central model-dispatch engine (a small loopback server). That engine already gated on budget (rolling account-usage percentage) but not on in-flight count — exactly the gap that tripped the org rate limit.
So the keystone change governs the engine itself: its single dispatch function now wraps its one model shell-out in a governor slot, the same semaphore as the shell side. One change, and every engine-routed app is governed centrally. It is fail-open at the import level (if the governor can’t load, the engine runs exactly as before) and priority-aware (a realtime request waits only a few seconds then fails open for low latency; batch and standard work wait longer to avoid bursting).
The practical gotcha worth generalizing: a long-running daemon keeps the old code until you restart it. Ship the file change, then restart the engine to activate — the change is safe and reversible; the restart is the operator action.
What to govern, and what to leave (a coverage map)
The reasoning that decides where a governor slot actually earns its keep, in priority order.
Govern first — the concurrent-burst moments:
- The central-engine chokepoint. One wrap covers every engine-routed app’s primary path. Highest leverage by far — do this one first.
- Bursty direct fallbacks. The callers that fire the model directly when the engine is down — which is exactly the concurrent-burst moment, because the engine being unreachable is what makes everything stampede at once. Each of these needs its own slot.
- Third-party-engine workers. An app that drives a non-primary engine (say a voice-agent tester) isn’t throttled by the model governor at all — its lever is a worker-count cut, not a slot.
Leave (deliberately):
- Single-shot interactive fallbacks. One model call per invocation, no in-process fan-out; the primary path is already governed at the engine. Wrapping them is belt-and-suspenders, not load-bearing.
- Harness-critical detached hooks. Hooks that spawn a detached background process aren’t a drop-in slot swap, and a bad edit to a harness hook can break every turn. Wrap these deliberately, with a short wait so a turn never blocks — and only with an explicit decision to do so.
The transferable rule: a governor slot buys you the most exactly where callers stampede together, and buys you nothing on the paths that only ever fire one-at-a-time. Spend your integration effort accordingly.
The drop-in itself is mechanical — a shell call site swaps command for
gov run -- command; a Python call site wraps the model subprocess in a
with governor.slot(): … context (adding a short wait for interactive callers).
Cross-machine cooldown (built, opt-in)
A per-machine ceiling can’t stop a fleet-wide burst: the provider’s org rate limit is shared across every machine, so several machines each respecting their own N slots still sum to N × machines concurrent. Because the shutdown is an org-level event, a 429 on any machine should back off the whole fleet.
How it works, cheaply — no hot-path latency:
- On a 429, the cooldown routine (shell and Python) writes a row to a tiny
shared coordination table (an
untiltimestamp). The machine that hit the wall also stops locally, instantly. - A roughly 60-second sync job on each machine reads the fleet’s max active cooldown and mirrors it into that machine’s local cooldown state — which the slot-acquire path already honors. So acquire stays purely local (zero added latency); only the rare 429 ever touches the network.
The whole cross-machine path is opt-in and fail-open until switched on: with the feature flag unset, no network calls happen; with the flag set but the shared table absent, reads return “no active cooldown” and writes fail harmlessly.
Tested
The semaphore is verified: a 6-caller burst against a max of 2 peaks at exactly 2; the cooldown counts down; a 429 in a command’s stderr auto-trips a machine-wide 90-second cooldown; and it behaves identically under a plain shell and an interactively-sourced one.