Omnigent · chapter 3 of 4

Measure the loop, remove the turns that decide nothing, and say which half of a good day was capacity.

Every row below has the same shape: a bottleneck, a change, a measurement taken the same way before and after, and the limit that survives it. The numbers come from the fleet’s own tables, read on September 14, 2026 after that day had closed, or from the commit that shipped the change. Where a measurement is still owed, it says so.

Measure before touching anything

The first three weeks of the fleet had no cycle-time number. The chain from card to PR to merge to deploy was created by exactly one code path at exactly one moment, then thrown away into event feeds and PR bodies, and every audit rebuilt the joins by hand. The lineage graph, designed on August 23, stores each link as an edge the minute it is minted. Its first read found about 2,380 lines of merged-quality work stranded on open PRs behind cards the board called done. That is the bug class a stored edge finds continuously and a hand sweep finds once.

The second read, on August 29, was tokens by role. The orchestrator, one durable conversation per card that shepherds it through the stages, was 54% of fleet tokens over the prior seven days, 20.9 billion of 38.8 billion, and 63% of its turns were no-ops. Most of those wakes were the saga answering its own writes: 552 dispatch-echo wakes in three days, 98% of them doing nothing. Each wake carried a 74 KB brief, a 71 KB message, and 70 to 270 KB of card reads to return about 200 bytes.

Orchestration share of the day's notional spend saga + triage, from the fleet's usage_by_role rollup 0 10% 20% 22% 9% 24% 22% 4% 5% 12% 14% 10% 13% 9% 3% Aug 29 Sep 2 Sep 3 Sep 4 Sep 5 Sep 6 Sep 7 Sep 8 Sep 9 Sep 10 Sep 11 Sep 12 $3,129 $494 $227 $300 $253 $201 $422 $260 $297 $234 $940 $2,023 under each bar: the day's notional total — API list rates on subscription seats, not an invoice; small totals are noisy
Fig. 01 — Orchestration — the saga and triage roles together — as a share of each day's notional spend, from the fleet's own usage_by_role rollup: mostly a tenth to a quarter, 3% on Sep 12. Notional means API list rates applied to work that ran on subscription seats, so read the ranking, not the dollars; the days with small totals swing the most.

Fig. 01 · pinch or scroll to zoom · drag to pan

The fix was three changes over two days. Stop the saga waking on its own writes. Give the runner a composite dispatch call and lean card reads so a wake that must happen costs a fraction of the old one. Hold quietly between events instead of polling. Then the bigger change over the following week: a stage executor that moves a card on evidence already stored, a review pass in the database or a verify bound to the head commit, without waking a model at all. By September 12, orchestration was 3% of the day’s notional spend. The design document that proposed the target survived a three-reviewer pass that refuted its first draft’s effect sizes, and the page keeps the refutation, because it is why the target was re-measured rather than believed.

Clean rounds stop costing a fix

A review round that found nothing still routed the card into fixing, because the transition was written for the case where something was found. On August 25 the fleet started coercing a clean round to a pass and asking the verifier to confirm or refute rather than re-review. On the first 16 cards through, 13 stored a pass as their first verdict, 81% against a seven-day baseline of 13 to 14%, and 12 of 16 skipped the fixing column entirely. The tax was 24 false hand-back rejections on 13 cards while the verifier and the executor learned to agree, and saga waits of 0.7 to 18 minutes at the boundary, median about 7. The tax was measured beside the win. It usually isn’t.

The fix that hurt

Builder context was the next number. Across 746 implement sessions the peak context was about 215k tokens at the median and 400k at p90, with one session at 932k, and sessions past 200k were 35% of the count and 76% of the spend. The obvious fix shipped on August 31: a 12 KiB ceiling on tool results and a ladder of thresholds, strip at 120k, compact at 150k, hand back at 180k, refuse tools at 200k.

It broke builders. On the afternoon of September 11, three of seven routine child builders compacted mid-task, and one came back with its context still above 180k and its MCP work tools disabled. I reversed the ladder the same day. The fleet now measures context per session and worker, exports it, and leaves the builder alone. This is the row I would show a client first: a plausible optimization, measured, reversed on the measurement, with the instrument kept.

The storm

September 11 was also the day the timeouts stopped fitting the host. They had been sized for a quiet machine, and the machine was not quiet.

A 30-second spawn cap missed 368 spawns in one day: 188 saga carriers, 116 Codex workers, 64 Claude workers. The MCP relay handshake ran 7 seconds at the median and 20 at p90, and 18 of 94 connects timed out. Worker creates took 15 to 30 seconds for 14 of 52, the slowest at 29.6, and one card collected five orphaned workers in forty minutes. The stall detector waited four hours before reclaiming a wedged worker’s slot.

Stall-dwell change, Sep 11 — slot-hours, from the commit that shipped it Before — 204.2 occupied slot-hours wedged · 167.4 h · 82% working · 36.8 h · 18% After — 186.4 capacity slot-hours returned to admission · 124.5 h · 67% not returned · 61.9 h · 33% 53 gaps, median 1 h 36 m +15 admissions (45 → 60) 0 50 100 150 200 h one scale in hours, two denominators — occupied slot-hours before, capacity slot-hours after
Fig. 02 — The stall-dwell change, Sep 11, in the numbers from the commit that shipped it. Before: of 204.2 occupied slot-hours, 167.4 (82%) were wedged and 36.8 working. After: of 186.4 capacity slot-hours, 124.5 (67%) came back to admission — 53 gaps with a median of 1 h 36 m, and fifteen more admissions (45 → 60). The bars share a scale in hours but not a denominator: occupied before, capacity after.

Fig. 02 · pinch or scroll to zoom · drag to pan

Each fix shipped that day from my own branch, with Claude in the terminal. The spawn cap went to 120 seconds, the runner connect bound from 75 to 180, the stall dwell from four hours to 30 minutes. The dwell change alone returned 124.5 of 186.4 capacity slot-hours to admission and let 15 more cards in, from 45 to 60. Then the cascade nobody sized: the new 180-second connect bound sat above 120-second cadence floors for rescue and rotation, so the settings page accepted values a later gate silently discarded. Found the next morning, fixed by raising the floors to 200.

The day

By September 12 the loop was cheap and the slots were free, and the host was the bottleneck. Four fixes in one day, all from my own branches with Claude in the terminal, none from the board.

At 14:07, scheduling. The host and server processes had been launched as background units, and under a load average near 120 with 60% CPU idle, the same CPU-bound loop took 0.28 seconds from an interactive shell and 13.6 seconds from a worker’s tmux server. Claude Code’s prompt hook took 0.7 seconds against 29, past its own 30-second timeout, so it had been cancelled in every recent worker. The units now run interactive.

At 16:02, the ship gate. The lane had merged about 20 PRs an hour at best, because each PR ran its own gate: candidate prep, a repo-wide collect-only sweep, then the smoke floor, 150 to 175 seconds measured, followed by a 180-second deploy drain that never finished. The gate now takes shipping-stage PRs as a batch, runs its two phases side by side, and drains for 30 seconds. It covered zero PRs for the next forty minutes, because the production wrapper offered neither batch method, until a second fix at 16:44 forwarded it through.

At 19:54, the updater. Every deploy cycle rebuilt the web UI inside its package sync, 3 minutes 23 seconds of a 5-minute cycle, and admission stayed closed for all of it.

Sep 11 · 40 merged Sep 12 · 246 merged 0 10 20 30 00 03 06 09 12 15 18 21 7 0 10 20 30 00 03 06 09 12 15 18 21 29 14:07 QoS fix merged 16:02 batch ship gate merged 16:44 batch gate forwarded
Fig. 03 — Merged cards per local hour on Sep 11 (40) and Sep 12 (246), both rows to one scale, counted off the board's stage table. The QoS fix merged at 14:07, the batch ship gate merged at 16:02 and was forwarded at 16:44 — 179 of Sep 12's 246 merges landed from 15:00 on.

Fig. 03 · pinch or scroll to zoom · drag to pan

The board’s own stage table shows what followed. 246 cards entered merged that day against 40 the day before. The hourly rate went from 4 to 8 through the morning to 24, 23, 29, and 25 across the afternoon and evening, then 18, 16, and 15 through the last three hours.

Stage p50, minutes — closed intervals, log axis Sep 12 Sep 11 min, Sep 11 → Sep 12 speccing n 103 → 286 10.0 → 1.0 building n 82 → 248 116.8 → 23.3 reviewing n 36 → 233 26.3 → 7.3 fixing n 25 → 137 68.7 → 18.6 verifying n 20 → 132 17.4 → 6.0 shipping n 26 → 217 4.0 → 4.7 up — batching PR open n 34 → 203 107.2 → 26.3 1 3 10 30 100 min Sep 12 mix: 197 routine cards, 166 of them epic children
Fig. 04 — Stage p50 in minutes, Sep 11 against Sep 12, from closed intervals on the board's own stage table: building fell from 116.8 to 23.3, PR-open from 107.2 to 26.3, and every stage but one dropped. Shipping rose, 4.0 to 4.7 — batching. Read with the mix in mind: Sep 12 was 197 routine cards, 166 of them epic children. The axis is logarithmic.

Fig. 04 · pinch or scroll to zoom · drag to pan

Stage medians fell everywhere except shipping, which rose from 4.0 to 4.7 minutes at the median and from 6.4 to 19.0 at p90, because a batched gate makes a PR wait for its batch. That is the trade, and it was the right one.

Sep 12 · ship-gate runs per hour outline: gate runs · solid: distinct candidates 0 10 20 30 00 03 06 09 12 15 18 21 00–14h: 62 runs on 62 candidates, 1:1 17–21h: 142 runs on 52 candidates, ~3:1 27/24 10/10 16:44 · batch gate forwarded
Fig. 05 — Sep 12 by local hour: gate runs matched distinct candidates one to one through 16:00 (the 15:00 hour ran 27 on 24), then ran about three times per candidate — 142 runs on 52 candidates from 17:00 to 21:00, after the batch gate was forwarded at 16:44; the last two hours ran 27 on 15 and 15 on 15. The week before held at or near 1:1 every day: Sep 8 50/50, Sep 9 35/31, Sep 10 13/13, Sep 11 29/29. From the fleet's own telemetry.

Fig. 05 · pinch or scroll to zoom · drag to pan

What the day does not prove

Two caveats travel with every number above, and the page would be dishonest without them.

The mix changed. Of the 246 cards, 166 were children of one decomposed epic, 197 were routine and 30 hard, and 160 were created that same day. And the pipe got wider on purpose: I raised the WIP cap from 12 to 20 that morning because the fleet was not saturated at 12. A wider pipe with a drained queue raises throughput while lengthening the queue’s tail, which is what the table shows: the median card took 5.7 hours from ready to merged against 5.4 the day before, and p90 went from 16.3 hours to 28.6. Little’s law says WIP equals throughput times lead time. Throughput went up, and so did time in the system.

So the day establishes that the fleet could run at 15 to 29 merges an hour from three in the afternoon to midnight, one hour at 10, with the gate batched and the host unstarved. It does not establish how much of that is the afternoon’s fixes and how much is the cap. The fleet’s own throughput document says it plainly: keep WIP and routing fixed for the first comparison, because deployment alone does not establish a speedup. That comparison, one hour with the cap held, against a matched cohort of routine cards on an ordinary day, is owed. This note will carry it when it exists. The day after is one more point and not that comparison: 188 cards merged on September 13 with the median at 1.9 hours from ready to merged, on a mix of 111 epic children, 135 routine, and 51 closed with no code change.

Underneath

Every measurement here comes from the fleet’s own instruments: the card stage-interval table for merges per day and hour, stage durations, and ready-to-merged lead time; the gate-run table for runs against distinct candidates; the usage-by-role rollup for spend share, which is notional API-rate accounting on subscription seats, useful for ranking and never a bill. The first read was taken at 21:40 on September 12, 2026 while the fleet was still merging, and this note first went out with it: 202 cards. The re-read on September 14, after the day closed, is what stands here now, 246 cards, and every stage, lead-time, and gate figure for the day moved with it. The pre-fix numbers for the scheduler, the ship gate, the updater, the timeouts, and the stall dwell are the ones recorded in the commits that shipped each change. The context-ladder thresholds and its reversal are recorded the same way. Little’s law is Little’s, from 1961: the mean number in a system equals the arrival rate times the mean time in it. Batching a gate is what a merge queue does, grouping pull requests into one build and capping how many merge at once, which raises merge throughput and makes each PR wait for its group. The framing that WIP and lead time move together under a wider cap is queueing arithmetic, not a finding of mine. Two outside cautions shape the refusal above. METR’s 2025 trial found experienced developers 19% slower with AI tools while forecasting a 24% speedup and believing afterward they had been sped up 20%. DORA’s 2024 report estimated a 1.5% drop in delivery throughput and a 7.2% drop in stability for every 25% increase in AI adoption, and its 2025 report reprints the same figures. A speedup I have felt but not measured the same way twice does not go on this page.

This project

This project, in order

Start

Tell me what’s stuck

I’ll tell you in about a day whether I’m the right person. The first conversation is fit, not a free architecture review.