The hardest components.
AI agents are the hardest components you can build a system from. They are brilliant, tireless, and confidently wrong; they follow instructions adversarially literally, they forget, and they will tell you a job is done when it is not.
The harness is everything I built around them so that none of that matters: a rulebook they operate under, machine checks that catch their lies, review policies, cost accounting, security fences, and a memory engineered against rot.
The result is an engineering organization that runs itself - and a demonstration, end to end, of the disciplines that make any system dependable: verification, failure-mode analysis, information architecture, cost engineering, and honest measurement.
Outcomes, not claims.
The outcomes, in plain terms. Each one is measured, not claimed - the numbers live further down.
Work you can trust without watching. Agents run unsupervised, including overnight; nothing lands without passing a battery of machine checks, and “done” is never accepted without evidence.
A system that catches its own lies. Every check is itself tested by deliberately injecting faults it must catch; a check that cannot fail is treated as broken.
Quality that doesn't decay. A knowledge system - the context cadence - keeps every project's state, decisions, and history current, because stale context is the quiet way even a careful process produces wrong work.
Costs that are engineered, not endured. Every token and every hour is measured against a baseline; every efficiency rule carries a recorded control and is dropped at reconciliation when the data shows it stopped paying.
A record built to resist faking. Append-only ledgers hold every decision ever made, the alternative that was rejected, why it lost, and what observation would reopen it.
It runs as a fleet. The whole environment reproduces across multiple machines from a single bootstrap, with fail-closed commit identity: a machine missing its enforcement refuses loudly rather than running unenforced.
Five disciplines.
Each discipline gets a plain sentence here, a visual, and full depth in the deep-dive further down.
1. Trust - autonomy without faith
Nothing is taken on the agent's word: a gate of dozens of independent check classes is the machine definition of “done”, command guards fence what an agent can execute, and the guards themselves are adversarially attacked to find out what they actually stop.
2. Quality - the context cadence
An agent's output is only as good as what fills its context window, so the harness treats project knowledge as an engineered artifact: budgeted state files rewritten at milestones, append-only decision and work ledgers, and a small fixed set of write triggers that remove discretion at exactly the points where diligence decays.
3. Economy - every token accounted for
Context is the scarce resource: the operating contract itself lives under a hard word budget enforced by its own check, spend is reconciled monthly against a baseline, and delegation policy dictates when work fans out to cheaper models and when it escalates.
4. Scale - one system, many machines and many minds
The harness runs as a fleet: multiple machines, multiple AI vendors under one shared contract, reproducible setup at both machine and user scope, verified transport of whole repositories between devices, and delegation rules that let one session orchestrate a swarm of subordinate agents with defined jobs, packets, and failure policies.
5. Self-knowledge - a system that tells the truth about itself
The harness measures its own adherence from its commit history, audits itself with quantified verdicts, keeps an explicit inventory of which rules are machine-enforced versus held on honor, and corrects its own record in public when a claim turns out to have been too strong.
By the numbers.
Rounded floors, as of autumn 2026.
30+ independent gate check classes, each one the machine answer to a specific way work has been wrong before.
~150 fault-injection controls in the gate's own self-test suite, 190+ more on the enforcement hook chain, a 48-control self-test on the cross-device exchange, and a 300+-control build of the portable check family: every check is proven able to fail before it is trusted to pass.
280+ adversarial controls on the command guard alone, including differential testing of what the guard judges against what actually executes.
A three-lane penetration test of the harness itself: 22 findings, 12 confirmed by execution, all closed with regression controls.
220+ recorded decisions across the system's ledgers, kept under a convention requiring each entry to name the rejected alternative, why it lost, and the condition that would reopen it.
Process adherence measured over each repository's full commit history - one project improved from 53% to 96%, another decayed and went dormant; both trajectories are on the record, because a measurement that only reports good news is advertising.
Three formal audits of its context system with quantified verdicts; roughly half the findings landed, the remainder tracked open on the record.
15 archived versions of the operating contract, every rule change dated.
A multi-machine fleet with fail-closed commit identity: a machine missing its enforcement refuses loudly rather than running unenforced.
An operating contract under an exact, gate-enforced word budget - the rulebook cannot silently grow or silently lose a rule.
The complete concept inventory.
For technical readers. Everything the harness is, what each part achieves, and how - at the level of design, not implementation. The implementations are the product; see “Working with me”.
One sentence of lineage, because provenance is part of honesty: the system began on a well-known public dotfiles-and-single-contract base, and everything this page describes - the gate architecture, the fault-injection discipline, the differential verification, the ledger law, the budget enforcement, the fleet machinery - is original work engineered on top of it.
The contract
The whole system hangs off one operating contract every agent loads: authority and safety rules, evidence rules, git discipline, engineering standards, delegation law, and cadence. Three properties make it an engineered artifact rather than a prompt. It is budgeted: the contract lives under a hard word ceiling with an exactly pinned reserve, enforced by its own check, so every added rule must pay for itself and evictions are recorded. It is layered: always-resident rules stay in the contract, task-specific procedure lives in pointer files loaded only at their trigger, so every session pays only for the rules its work needs. It is enforced: a rule that matters graduates from prose to a hook or a gate class, because the system's own audits proved that prose alone decays.
The gate: a machine definition of done
Governed repositories carry a gate script that is the machine definition of “done”. Its dozens of classes each encode one specific historical failure: undeclared stubs, silently deleted rules, drifted settings, stale maps, broken budgets, identity violations, missing worklog evidence. The gate runs before any green claim, and its output - not the agent's summary - is the evidence.
Fault injection: a gate that cannot fail is not a gate
The signature discipline. The gate has its own test suite of ~150 controls that attack the system as shipped, not staged examples, and assert that every class actually goes red on the fault it exists to catch. Positive controls prove the checks fire, and expectations are never measured against a broken baseline. The same standard applies to every verification tool in the system: if it has never been watched failing, it is not trusted to pass.
The guard arc: adversarial verification of one's own enforcement
The harness fences what agents can execute: command guards, a deny fence, and protected paths. The part worth hiring for is what happened next. Independent review lanes were cast to break the guards, and they succeeded - repeatedly. Instead of patching each finding and reclaiming strength, the record was corrected to downgrade the guard's claimed strength, and a differential verification was built: continuous measurement of what the enforcement actually catches against what actually executes, so a gap is a red build rather than a silent hole. Defense in depth was then chosen deliberately: the most critical assets gained an independent protection layer that does not depend on the guards at all.
Evidence discipline
“Done” requires evidence produced this session: fresh command output, file-and-line citations, gate counts. Two separate vocabularies keep claims honest: a lifecycle axis for where a task stands, and a verdict axis for what a check actually said - the two are never fused, so “done” and “verified” cannot be conflated. Filtered output counts as evidence only with a pointer to the unmodified original. Stubs must declare themselves and fail loudly; a passing verdict over an undeclared stub is a false green and is treated like a red gate. User-interface claims require an exercised app or a screenshot, and prototypes are never presented as product.
The ledgers: a record built to resist faking
Every governed repository carries a small set of context files with one owner per fact. The decision ledger is append-only by convention - held by git history, non-author review, and enforcement hooks - and every entry names the rejected fork, why it lost, and the observation that would reopen it, including entries that correct the system's own earlier overclaims, because the record is corrected forward, never rewritten. The worklog is an append-only table keyed by stable task IDs, with lifecycle and verdict in separate columns and evidence as pointers. A stop-hook refuses to end a repo-changing session without its dated worklog line.
The context cadence
The quality half of the system, built on one insight: stale or missing context is the quiet way a clean-window workflow still produces wrong code - the window was clean, but what filled it was wrong. Two generations of tooling implement it: a nested context engine with per-module context files, freshness markers, and a deterministic no-LLM staleness checker; and the flat convention above, with hard budgets and a small fixed set of write triggers - never agent discretion in between. Adherence is measured from the actual commit history, per repository, as a trajectory, and both improving and decaying trajectories are on the record. The quality link is the design thesis the measurement exists to test: the numbers say where the process is holding and where it is decaying, before anyone has to guess.
The session loop and decisions-before-code
Sessions run a fixed loop: reconcile, take one scope, build small, verify, review at triggers, land. Decisions the work will force are named and settled before execution, while the session is still context-cheap, because a decision made after the code exists is biased by the code. Bug fixes start by reproducing the bug end to end as a user would hit it, and fixes target the root cause across every caller, not the reported symptom.
The review economy
Review is mandatory at fixed triggers - non-mechanical plans, experiment designs, the riskiest change of a run, anything safety-adjacent - and a trigger is satisfied only by a reviewer that did not produce the work; the author's own re-read is not a review. Praise for an unverified result triggers refutation framing: assume it is wrong and try to break it. Re-reviews are surgical: scoped to the delta since the reviewed commit plus the implied regression set, re-deriving nothing already confirmed. Proportionality is law in both directions - mechanical work a direct check proves gets no review lane, and a verification run that cannot change a decision is recognized as waste, not rigor.
Delegation: org design for a workforce of models
One session can orchestrate a swarm, under law rather than improvisation. Concurrency ceilings are set per model class and are ceilings, never targets. Known-shape work casts one wide scout that hands exact targets to narrow workers; unknown-shape work keeps blind independent sweeps, because a scout there is a single point of failure for coverage. Workers return uniform packets - claim, evidence, confidence, status - and a failed status is valid data, never an exception to hide. The failure policy is fixed: one re-scope and retry, then a different approach with the failure logged; never a third identical attempt, because fresh eyes beat tired persistence. Model choice is an escalation ladder: the cheaper model is the default for every lane, and a larger model must be earned by evidence and named at cast, so off-default choices are visible rather than habitual. Concurrent writers get isolated worktrees with disjoint file scopes; shared ledgers are written only by the orchestrating thread after the merge.
Cost engineering
Spend is modeled monthly from measured token flows against a recorded baseline. The findings drive policy: session defaults were lowered when the data showed the pinned default, not the work, was driving cost; sessions are closed before going idle because a cold resume re-uploads the whole context; ingestion is capped at the call site because every token pulled in is re-sent on every later turn. Every efficiency rule names its control when it lands, and is dropped at reconciliation when the data shows it stopped paying - rules that expire rather than accumulate.
Intake: parsing the human
The harness treats operator input as an engineered interface too. Dictated braindumps are parsed under fixed rules: transcription artifacts identified, multiple threads split and tagged as orders, questions, or parked ideas, musings distinguished from orders, and destructive requests echoed as exact target lists awaiting one word - because the most dangerous failure mode of a capable agent is confidently executing what the human almost said.
Fleet, transport, and reproducibility
The environment reproduces across machines, at both machine and user scope, idempotently, with pinned tools installed from a manifest. Repositories travel between devices as verified bundles, with onboarding through recorded handoffs, never ad-hoc copies. Commit identity is enforced fail-closed: a machine missing its enforcement refuses to commit rather than committing unenforced. Canon-and-copy law keeps multi-device truth coherent: a file's canonical copy declares itself canon in its own header, every other copy names its canon, and backups are refreshed by re-copy, never edited. A single map file is the first thing a cold agent reads: the component table, each component's home and governance, and a formal intake for how a new concept enters the system at all.
Self-audit and the honor inventory
The system maintains an explicit honor list: which guarantees are machine-enforced, and which still rest on an agent remembering - each honor residual named with an owner. Its context system has been formally audited three times with quantified verdicts; roughly half the findings landed, and the remainder is tracked open on the record rather than forgotten. The harness itself has been penetration-tested by adversarial agent lanes: 22 findings, 12 confirmed by actually executing the attack, and every confirmed finding closed with a regression control so it can never silently return. The finding that shapes everything: machine-enforced rules held everywhere, and everything resting on honor decayed on a predictable gradient. The response is the system's engine of improvement: observed decay drives mechanization, every incident becomes a regression control, and the record is corrected in public when a claim was too strong.
Breaking my own guard.
The trust discipline, end to end.
A command guard was built to fence agent shell access, with a test suite in the dozens of controls, and it shipped green. Adversarial review lanes were then cast against it with one instruction: break it. They did - more than once, in ways the suite had never imagined.
The engineering decision that matters: stop patching findings one at a time. The record was corrected to downgrade the guard's claimed strength - my own shipped work, marked down on the record - and the problem was reframed: don't argue about what the enforcement should catch, measure what it actually catches.
A differential verification now continuously compares what the enforcement judged against what actually executed; known divergences are a maintained baseline where every line must cite its cause, and a new divergence is a red build, not a silent hole. The guard suite has since grown to 280+ adversarial controls, every one of which was watched failing before its fix landed. And the most critical assets gained an independent protection layer that does not depend on the guards at all - defense in depth chosen deliberately, not by default.
That is the shape of real verification engineering: assume your enforcement is wrong, measure how, and make the measurement permanent - a living baseline, never a closed case.
The cadence that measures itself.
The quality discipline, end to end.
The context cadence exists because agent quality fails silently: the work looks right while the assumptions under it have gone stale. The convention makes project memory an engineered artifact - budgeted, append-only, written at fixed triggers. Then it does the rare thing: it measures itself.
Adherence is computed from each repository's commit history as a trajectory; three independent audits of the context system returned quantified verdicts; and the deterministic staleness checker, run on the tooling's own repository, flagged its own engine module as stale - the tool catching genuine drift in its own house.
The honest result is published rather than polished: machine-enforced parts held everywhere, honor-based parts decayed predictably - one project's adherence climbed from 53% to 96% while another's fell and the project went dormant - and that gradient, not any single tool, is what the system is engineered around.
In the system's own words.
A portfolio that only shows green checkmarks is advertising; this one can show its audit trail. At the time of its context system's audit, 8 of 91 rules were machine-enforced; everything else ran on honor - and the audit found that honor-held rules decay on a predictable gradient.
That finding is not a footnote; it is the operating thesis: prose loses, so observed decay drives mechanization, and the honor inventory exists precisely so that what is not yet enforced is named, owned, and next. The enforcement suites, the penetration test, and the regression controls in the numbers above are what that thesis has produced since.
I publish this because a system that cannot tell you where it is weak cannot be trusted about where it is strong.
The concepts are free.
The implementations - the current contract and rule corpus, the gate and its control suites, the guards and their differential verification, the cadence tooling, the fleet machinery, and the judgment encoded in 220+ adjudicated decisions and a full audit history - are the product, privately held.
Build it for you. A harness engineered for your codebase, your models, your risk tolerance - from contract to gates to fleet - delivered with its own fault-injection suite, because you should not trust my checks on my word either.
Harness assessment. An audit of your existing agent setup against these disciplines: where you are trusting prose, where your checks cannot fail, where your context is rotting, where your spend is unmeasured - with a quantified report and a prioritized mechanization plan.
Advisory. Design reviews, workshops, and standing counsel for teams building their own agentic infrastructure.