The governance engineering ladder

The governance engineering ladder

The field names its practice in rungs, and each rung names the artifact you engineer: the prompt, then the context window, then the harness, then the loop. Four serious fleet designs are running in public right now, and comparing them dimension by dimension shows the same two cells empty everywhere: contracts between repositories, and governance built as its own engineered layer rather than delegated to review habits. That layer is the next rung. We compared the visible practices on their own words and sources, and this piece is the comparison: what each fleet optimizes, how it is governed, what paper trail it leaves, and where the ladder goes from here.

The ladder so far

Prompt engineering needed no coiner. Context engineering was Tobi Lütke's term, amplified by Andrej Karpathy. Harness engineering traces to Mitchell Hashimoto's adoption essay. Loop engineering emerged twice, independently: inside Anthropic, where agent-written code now dominates the Claude Code team's own output, and in Peter Steinberger's OpenClaw work. Each rung subsumed the one below it. None of them says how many loops, across many repositories, stay coherent without eating each other.

Subscribe free

Four fleets, on their own words

Interpretable Context Methodology (Jake Van Clief): folder structure is the architecture. One orchestrating agent reads staged folders, each stage carries a CONTEXT.md contract, and the method is explicitly anti-swarm. It is documented in a self-published preprint and taught through a free community of 41,600+ members with paid tiers above it. Governance is architectural: a human reviews at every folder handoff.

The Agentic Coding Flywheel (Jeffrey Emanuel): the opposite bet. Dozens of concurrently subscribed agent accounts (52+, per his own GitHub profile) coordinate as a flat mesh over agent mail, bootstrapped from a free setup hub and monetized as a skills subscription. Governance is a documented six-stage human-in-the-loop cycle: plan with competing frontier models, polish the task graph, then let the fleet execute against Beads-style issue graphs.

Dark Factory (Steve Yegge): agents work with nobody watching. Gas Town runs roughly 12 to 30 concurrent Claude Code instances under seven fixed roles, on the original Beads ledger, and its author is explicit that it has no business model. Its v1 governance is equally explicit: self-described Wild West, with agents merging their own pull requests, and at least one first-hand account of a merge landing despite failing integration tests.

Loop engineering (Boris Cherny): verifier-gated loops with no human in the inner loop, memory accreting into CLAUDE.md so each failure tightens the next run. It is the nearest neighbor on discipline, applied at production scale inside one organization, with public numbers instead of promises.

The comparison

DimensionICMFlywheelDark FactoryLoop engineeringhosaka/Loa
What you engineerfolder contractsthe mesh + task graphthe factory floorthe loop + verifiersthe governing state
Topologyone agent, staged folders52+ accounts, flat mesh12-30 agents, 7 roles~5 parallel instancesone governed agent per repo, ~30-repo fleet
Governancehuman at each handoffsix-stage HITL cycleWild West (v1, self-described)verifiers, no inner-loop humangated pipeline per change: plan, implement, review, audit; operator approve/kill walls
Paper trailCONTEXT.md filesbead graph + mail threadsBeads ledger + essaysCLAUDE.md accretionevery change a gated PR; contracts pinned by hash; committed is not published
Cross-reposingle workspacesingle workspacesingle workspacesingle orgpublish/consume contracts; no repo writes another's schema

Fair names throughout are each author's own. Claims marked "per his own profile" are self-reports we could not independently verify; we kept every such label visible on purpose.

Below the first rung, the older maintenance logs mark a landing, unnumbered, swept regularly.

The two empty cells

Every compared fleet lives in one workspace. Governance ranges from implicit in the architecture to absent by design. That is not a criticism: each is optimizing its chosen rung well. It is an observation about where the ladder runs out.

Our practice, one governed Loa instance per repository, engineers the two empty cells directly, and this post is itself the evidence in the way we mean it. It came out of a knowledge base where every fact carries provenance and a trust tier, it passed blocking content gates for slop, overclaim, identity leakage, sensitivity, and citation coverage before it could ship, and a human approved it at a wall the pipeline cannot cross on its own. The comparison table above was adversarially verified claim by claim before we let ourselves cite it. You stop engineering the agent, and you start engineering the state it operates under. The rung has a name: governance engineering, or in long form agentic governance engineering (AGE). The record it leaves is public even where the source is not, and that is the standard we think the next rung of the ladder owes its readers.


Get the next post

Free membership: new posts on how the fleet is built, delivered by email. Subscribe free


Post history

  • 2026-07-21: densify the-governance-engineering-ladder: 1 internal link(s), pass cl-20260721 (standing admin preapproval (pending-laws 2026-07-19))
  • 2026-07-21: crosslinker + enrich phase + OQ-1 spike + caption-uniqueness law
  • 2026-07-21: BLOG SURFACE ENRICHMENT: four corners + engine images + backfill (sprint-25)
  • 2026-07-13: queue lined up: 3 new posts through the full pipeline, grouping law (tags), rubric anchor scoping
  • 2026-07-13: THE LOAD-BEARING TRIO: fixpoint, fate, colophon, shadow gate + backfill
  • 2026-07-12: charter addendum (self-eating spiral, mibera stats, priced departures, fingerprint carve-out) + THE JESTER LEDGER + shadow-sentence backfill in all 9 essay sources
  • 2026-07-10: related-posts footer + post-history changelog, C8 concepts gate, ghost:wire sync-all (sprint F3)
  • 2026-07-10: closed-source evidence posture: de-link private repos (Loa upstream is public, link that), honest evidence claims, AGE long-form in ladder draft; curriculum §7 evidence policy + idea-ledger rows (AGE, hounfour/freeside troves)

№ 5915