The governance engineering ladder
The field names its practice in rungs, and each rung names the artifact you engineer: the prompt, then the context window, then the harness, then the loop. Four serious fleet designs are running in public right now, and comparing them dimension by dimension shows the same two cells empty everywhere: contracts between repositories, and governance built as its own engineered layer rather than delegated to review habits. That layer is the next rung. We compared the visible practices on their own words and sources, and this piece is the comparison: what each fleet optimizes, how it is governed, what paper trail it leaves, and where the ladder goes from here.
The ladder so far
Prompt engineering needed no coiner. Context engineering was Tobi Lütke's term, amplified by Andrej Karpathy. Harness engineering traces to Mitchell Hashimoto's adoption essay. Loop engineering emerged twice, independently: inside Anthropic, where agent-written code now dominates the Claude Code team's own output, and in Peter Steinberger's OpenClaw work. Each rung subsumed the one below it. None of them says how many loops, across many repositories, stay coherent without eating each other.
Subscribe free
Four fleets, on their own words
Interpretable Context Methodology (Jake Van Clief): folder structure is the architecture. One orchestrating agent reads staged folders, each stage carries a CONTEXT.md contract, and the method is explicitly anti-swarm. It is documented in a self-published preprint and taught through a free community of 41,600+ members with paid tiers above it. Governance is architectural: a human reviews at every folder handoff.
The Agentic Coding Flywheel (Jeffrey Emanuel): the opposite bet. Dozens of concurrently subscribed agent accounts (52+, per his own GitHub profile) coordinate as a flat mesh over agent mail, bootstrapped from a free setup hub and monetized as a skills subscription. Governance is a documented six-stage human-in-the-loop cycle: plan with competing frontier models, polish the task graph, then let the fleet execute against Beads-style issue graphs.
Dark Factory (Steve Yegge): agents work with nobody watching. Gas Town runs roughly 12 to 30 concurrent Claude Code instances under seven fixed roles, on the original Beads ledger, and its author is explicit that it has no business model. Its v1 governance is equally explicit: self-described Wild West, with agents merging their own pull requests, and at least one first-hand account of a merge landing despite failing integration tests.
Loop engineering (Boris Cherny): verifier-gated loops with no human in the inner loop, memory accreting into CLAUDE.md so each failure tightens the next run. It is the nearest neighbor on discipline, applied at production scale inside one organization, with public numbers instead of promises.
The comparison
| Dimension | ICM | Flywheel | Dark Factory | Loop engineering | hosaka/Loa |
|---|---|---|---|---|---|
| What you engineer | folder contracts | the mesh + task graph | the factory floor | the loop + verifiers | the governing state |
| Topology | one agent, staged folders | 52+ accounts, flat mesh | 12-30 agents, 7 roles | ~5 parallel instances | one governed agent per repo, ~30-repo fleet |
| Governance | human at each handoff | six-stage HITL cycle | Wild West (v1, self-described) | verifiers, no inner-loop human | gated pipeline per change: plan, implement, review, audit; operator approve/kill walls |
| Paper trail | CONTEXT.md files | bead graph + mail threads | Beads ledger + essays | CLAUDE.md accretion | every change a gated PR; contracts pinned by hash; committed is not published |
| Cross-repo | single workspace | single workspace | single workspace | single org | publish/consume contracts; no repo writes another's schema |
Fair names throughout are each author's own. Claims marked "per his own profile" are self-reports we could not independently verify; we kept every such label visible on purpose.
Below the first rung, the older maintenance logs mark a landing, unnumbered, swept regularly.
The two empty cells
Every compared fleet lives in one workspace. Governance ranges from implicit in the architecture to absent by design. That is not a criticism: each is optimizing its chosen rung well. It is an observation about where the ladder runs out.
Our practice, one governed Loa instance per repository, engineers the two empty cells directly, and this post is itself the evidence in the way we mean it. It came out of a knowledge base where every fact carries provenance and a trust tier, it passed blocking content gates for slop, overclaim, identity leakage, sensitivity, and citation coverage before it could ship, and a human approved it at a wall the pipeline cannot cross on its own. The comparison table above was adversarially verified claim by claim before we let ourselves cite it. You stop engineering the agent, and you start engineering the state it operates under. The rung has a name: governance engineering, or in long form agentic governance engineering (AGE). The record it leaves is public even where the source is not, and that is the standard we think the next rung of the ladder owes its readers.
Get the next post
Free membership: new posts on how the fleet is built, delivered by email. Subscribe free
Read next
- Governance Engineering
- The repo remembers: how agents build hosaka
- One human, twenty repos: how hosaka is architected
Post history
- 2026-07-21: densify the-governance-engineering-ladder: 1 internal link(s), pass cl-20260721 (standing admin preapproval (pending-laws 2026-07-19))
- 2026-07-21: crosslinker + enrich phase + OQ-1 spike + caption-uniqueness law
- 2026-07-21: BLOG SURFACE ENRICHMENT: four corners + engine images + backfill (sprint-25)
- 2026-07-13: queue lined up: 3 new posts through the full pipeline, grouping law (tags), rubric anchor scoping
- 2026-07-13: THE LOAD-BEARING TRIO: fixpoint, fate, colophon, shadow gate + backfill
- 2026-07-12: charter addendum (self-eating spiral, mibera stats, priced departures, fingerprint carve-out) + THE JESTER LEDGER + shadow-sentence backfill in all 9 essay sources
- 2026-07-10: related-posts footer + post-history changelog, C8 concepts gate, ghost:wire sync-all (sprint F3)
- 2026-07-10: closed-source evidence posture: de-link private repos (Loa upstream is public, link that), honest evidence claims, AGE long-form in ladder draft; curriculum §7 evidence policy + idea-ledger rows (AGE, hounfour/freeside troves)
№ 5915