In my Bosun post, I described the problem that appears when coding agents stop being a novelty and become a crew. The models can do the work. The human becomes the bottleneck: holding the plan, routing tasks, watching terminals, moving context, checking evidence, and remembering who is waiting for what.
Bosun was my first answer. It proved that I wanted one chief of staff rather than a wall of disconnected sessions. It also proved something less comfortable: once Firstmate had become the stronger home for that idea, keeping Bosun alive meant maintaining a second orchestration system with overlapping responsibilities.
So I retired Bosun and adopted Firstmate wholly.
That does not mean I replaced one name with another. I stopped building a parallel chief-of-staff layer and moved my energy into a public Firstmate fork, where my changes can stay close to the upstream project. Firstmate now owns intake, routing, dispatch, supervision, delivery and the route back to me when a decision genuinely needs a human.
The system around it has three layers.

- Firstmate: Bridge crew: intake, routing, dispatch, supervision, merges
- Agent Vault: Chart room: one packet per project, read-only
- Personal Agent Kit: Shipyard and standing orders: rules, skills, hooks, pinned tools, traces, decisions
- Standing orders: Loaded into every harness
- Crewmate: Bounded work out, pull request back
Three layers, three different jobs
The Personal Agent Kit is the shipyard and the standing orders. It owns the reusable rules, skills and hooks that each harness receives. It also owns pinned tool versions, observability conventions and decision records. The Kit answers: how should an agent behave before it touches a project?
Firstmate is the bridge crew. It turns a request into bounded work, chooses a route, dispatches a crewmate, supervises that crewmate, and brings back a pull request or a decision. It answers: who is doing this, where are they doing it, and what needs attention now?
Agent Vault is the chart room. It holds concise project packets: verified current state, the next action, open loops and links to canonical evidence. Firstmate can read those packets, but the Vault is not an orchestration database and never edits a source project. The repository still wins if a packet has drifted.
Keeping those boundaries sharp matters. The Kit does not become a scheduler. Firstmate does not become the durable owner of project knowledge. The Vault does not drive the crew. Each layer has one reason to change.
The crew has two machines
The biggest practical change is not a model. It is where the work runs.

- Laptop: Captain session: intent, judgement, merge calls
- Brief: Work that does not need the laptop goes to the Mac mini
- Mac mini: Persistent second mate, always on
- Crewmates: One visible pane each
- Status: Reports and pull requests return
My laptop runs the captain session: the conversation where I set intent, inspect outcomes and make the calls that need taste or authority. A Mac mini runs a persistent second mate. Work that does not need the laptop goes there by default, rather than occupying the machine I am using to think.
The second mate is not a second captain. It is a persistent Firstmate direct report with its own isolated home and fleet state. Firstmate can route a brief to it, receive its status, and recover honestly if that remote route is unavailable. It never silently replaces a failed remote route with a different local execution path.
That gives me a simple operating rule: keep interactive work close; send independent work to the always-on lane. The laptop stays responsive, and the crew can keep moving when I close its lid.
Dispatch is ordered, visible and quota-aware
My current worker order is Codex first, Kiro V3 second, Antigravity third. Codex remains the default writer. Kiro gives the fleet another capable native harness and is especially useful when specifications and acceptance criteria are central. Antigravity supplies the third route.

- Firstmate: Owns the route
- quota-axi: Reports which quota windows are open
- First: Codex: Default writer. Window spent, so the task moves on
- Second: Kiro V3: Window open. Takes the task
- Third: Antigravity: Stands by
- Herdr: Hosts every pane, so every session stays visible
The order is deliberately strict. A fallback policy should be predictable enough to explain after the fact. My quota-axi fork adds the other signal Firstmate needs: which subscription quota windows are actually available. I added a Kiro provider so that the quota view covers the same harness that the dispatch path can select.
Herdr hosts every pane. Firstmate can supervise without turning the crew invisible; I can still open the runtime and see the exact worker session. That visibility is important. Delegation should reduce attention, not remove inspectability.
Delivery is a contract, not a victory message
The old version of this system treated delivery as something I was still designing. The current version has a much clearer owner: every change goes through no-mistakes.

- Intent: The request travels with the change
- Rebase
- Review: Claude reviews what Codex wrote
- Test: Evidence, not intuition
- Document
- Lint
- Push
- Pull request
- CI: Continuous Integration
- Merge gate: Opens on the captain’s word, unless the project has standing approval
The pipeline preserves the original intent, rebases onto the current foundation, reviews the change, runs tests, checks documentation, lints, pushes a feature branch, opens a pull request, and waits for Continuous Integration (CI) evidence. Codex writes by default. Claude reviews. The author does not mark their own work correct.
The decision boundary is also explicit. The first mate can decide a clear-cut ask-user finding when the answer follows directly from recorded authority. A genuinely ambiguous product or risk decision comes back to me. Merge authority is never inferred from implementation authority: it defaults to the captain unless a project has recorded standing approval.
This is why I do not need another lifecycle orchestrator around Firstmate. Firstmate owns the crew lifecycle. no-mistakes owns the change-delivery lifecycle. The clean seam between them is more useful than a larger workflow that tries to own both.
Evals define done
If one habit runs through the whole setup, it is this: I do not trust a change because it feels better. I trust it because a check that existed before the change says it is better.
That habit comes from my professional work, which includes MLflow tracing and evaluation at enterprise scale for a UK bank. There, an agent that “seems smarter” is not a release criterion. The same discipline now governs my own crew, at a much smaller scale.
In practice it has four parts:
- Evals define done. Before a change starts, the task records how it will be checked. A task without a check is not ready to build.
- Before and after, not after alone. A number only means something next to the number it replaced. In one recent Kit task, AK-088, acceptance criterion AC-12 is exactly that: a before-and-after comparison of completion time, model usage, human interventions and review cost.
- Skills have budgets.
reference-scandeclares how many tokens it may spend, and it is measured against that budget like any other eval. A skill that finds the right answer by reading everything has failed. - The Test step is evidence. no-mistakes records what ran and what passed. The reviewer reads that record, not the author’s summary of it.
This is also why the learning loop below ends at a gate rather than at a commit.
Tracing should improve the next session
The observability layer is still in progress. The direction is to launch agents through my Omnigent fork so different native harnesses share one trace path. Full-content traces stay only in local MLflow, and secrets are redacted before storage.

- Omnigent: Launches every agent through one pathPlanned
- Codex, Kiro, Antigravity: Native sessions, as they run todayLive today
- Redaction: Secrets stop here, before anything is storedPlanned
- Local MLflow: Full-content traces, on this machine onlyPlanned
- Backpass-style pass: Reads MLflow, so Kiro and Antigravity sessions count tooPlanned
- Proposed edit: Evidence attached. Next stop: the evaluation gate
Collection is only the first half. I am also designing a Backpass-style learning loop, based on Kun Chen’s Backpass. The loop looks for repeated, evidenced gaps and proposes focused edits to AGENTS.md or a skill. It does not silently rewrite the rules.
My planned variant reads MLflow traces rather than relying on one harness’s local transcript format. That matters because the crew is intentionally mixed. Kiro and Antigravity sessions should count as learning evidence alongside Codex sessions. If the feedback loop can only see one worker, it will optimise a partial picture of the system.
Then the evals take over. A proposed edit is just another change, so it faces the same before-and-after comparison as anything I write by hand.

- Local MLflow: Full-content traces from every harness, secrets redactedPlanned
- Backpass-style pass: Turns trace evidence into a proposed AGENTS.md or skill editPlanned
- Evals, before and after: AC-12 compares completion time, model usage, interventions and review cost. Lower is betterLive today
- Evaluation gate: Evidence, not intuitionLive today
- Held: Worse numbers, so the edit does not ship
- Standing orders: The captain accepts the edit with its evidence attachedLive today
The gate is the part that runs today: changes are judged against their evals and the no-mistakes Test step before they ship. What is planned is the input. Today I write the proposals; next, trace evidence will draft them. Either way, an edit that makes the numbers worse does not ship, and an edit that improves them still waits for me to accept it with the evidence attached.
The skills that make the system feel personal
The orchestration gets most of the attention, but the skills are where repeated judgement becomes reusable behaviour. These are six I find especially useful:
reference-scantreats GitHub stars as a token-budgeted design library, not a pile of links. GPT-6-Astra and Claude Fable co-designed the approach.kb-captureplus durable memory turns a useful session lesson into a sourced proposal, while small continuity facts remain easy for the next agent to recall.project-context-refreshgives an agent a trustworthy first five minutes: branch state, stack, gates, maps, drift and one sensible starting action.codebase-mental-modelexplains how components interact at a human altitude and keeps the map honest as those interactions change.spec-delivery-loopturns an approved intention into one coherent change, focused tests, aggregate verification and independent review.system-design-assessmentslows the work down when a change affects boundaries, ownership, security, reliability or deployment—and stays out of the way when it does not.

- GitHub stars: A design library, not a pile of links
- reference-scan: Ranks stars against the task. Co-designed by GPT-6-Astra and Claude Fable
- Token budget: Measured against its declared budget, like any eval
- Agent context: A shortlist of three, not the whole library
The point is not to collect the largest skill directory. It is to capture the small number of decisions I want every capable harness to make consistently.
The Firstmate fork carries my operating choices
Adopting Firstmate wholly did not mean using it unchanged. My public forks hold the customisations that make the upstream ideas fit this estate:
- A Kiro V3 crewmate harness, including liveness recognition, so supervision can distinguish real progress from a dead worker.
- A strict worker fallback order: Codex, then Kiro, then Antigravity, informed by quota-axi rather than improvised per task.
- Every pull request through no-mistakes, so review and CI evidence use the same delivery contract.
- A remote second mate on the Mac mini, making remote-safe work the default destination while the laptop remains free.
- Herdr, quota-axi and Omnigent forks that carry the runtime, quota and tracing integration points I need.
- An Omnigent launch option, still in progress, for harness-neutral traces across the full crew.
The last piece is how I stay close to upstream without erasing those choices.

- Upstream: New commits arrive
- Custom patches: The fork’s own changes. Every sync must keep them
- Guard: A merge commit, checked: patches survive and tests stay green
- Stopped for inspection: A conflict halts the sync
- Public forks: Firstmate, Herdr, quota-axi and Omnigent sync this way
- Estate: Updates only after proof
The Firstmate fork uses guarded merge-commit upstream sync. An upstream change becomes a candidate merge. The guard checks that custom behaviour survives and that the evidence remains green. Only then does the merge commit land in the public fork. A conflict stops the flow for inspection instead of replacing the fork with upstream state.
What is live, and what comes next
Live today: Firstmate is the only chief-of-staff layer. The Personal Agent Kit supplies standing orders. Agent Vault supplies read-only project context. The laptop hosts the captain session. The Mac mini hosts the persistent second mate. Codex, Kiro and Antigravity follow a strict fallback order. quota-axi informs availability. Herdr keeps every pane visible. no-mistakes carries each change to a green pull request, and merge authority remains explicit. Evals define done, and changes are measured against them before they ship.
In progress: the Omnigent launch option, full-content local MLflow tracing across every harness with secrets redacted, the MLflow-backed Backpass-style proposal loop feeding the evaluation gate, and selective adoption of AI-DLC mechanisms without adopting its lifecycle orchestrator.
What comes next is less about adding another agent and more about closing the evidence loop. I want every worker to be observable through the same path, every delivery claim to point to an external check, and every proposed improvement to show the traces that justify it. The system should get better because I accepted a well-supported change—not because an agent quietly edited its own standing orders.