How I build
AI-assisted, spec-driven, reviewed by hand
AI coding agents write most of the code in my projects. I say that up front, because a lot of people distrust AI-written software, and often for good reasons: code nobody read, tests nobody ran, "it works" with nothing behind it. This page shows what I do against that.
Who does what
| Who | Does |
| Me | The problem, the specA short document that says what the software must do and how we check that it does it., the architectureThe overall structure: which parts exist and how they talk to each other. and every design decision. I approve each plan before work starts, read each change, test each release by hand and decide when it ships. |
| The agents | Read the existing code, write tests and code against the approved spec, run the checks and report what they measured. |
The process
Every change goes through the same steps. The highlighted ones are mine and cannot be skipped.
- SpecWhat the change must do and how we will know it works.
- PlanThe agent reads the code it will touch and proposes a short plan.
- ApprovalNothing gets written before I say go.
- Failing testFirst write down how success looks, see that it is not there yet, then build it until the check says yes.A test that shows the bug or the missing feature, and fails.
- CodeThe smallest change that makes the test pass.
- Independent reviewA second AI, which did not write the code, checks it against the spec.For risky changes: save formatsHow the app stores your data on disk. A mistake here can lose data., migrationsSteps that convert stored data to the format of a new app version., security, concurrencySeveral things running at the same time and touching the same data.. A separate reviewer that did not write the code checks it against the spec.
- My reviewI read the diffThe exact lines a change adds, removes or edits.. If it is wrong, it goes back to step 5.
- Hands-on testingThe release runs on my own systems, often for hours, before anyone else gets it.
The rules the agents work under
The agents follow a written rule file that I maintain and change whenever a mistake shows a gap. Some of the rules:
- Every statement is either measured in the current session, carried over from earlier, or a guess, and the agent must say which. Only measured facts may be stated plainly.
- "Fixed" means the original problem was run again and the new result was seen. A green test or a clean exit codeThe number a program reports when it ends. 0 means "no error", not "correct". alone is not proof.
- A failing test comes before the code that fixes it.
- Small diffs. The fix goes where the cause is, not as an extra check in every callerA piece of code that uses another piece of code..
- No deleting, force-pushingOverwriting the shared code history. It can destroy other work. or changing shared systems without my explicit approval.
- SecretsPasswords, API keys and access tokens. never appear in output.
- No placeholder code"Fill this in later" code that looks finished but does nothing.. Implement it or ask.
- Numbers are taken from a tool, never counted by eye.
The setup in numbers
| Count | What it is |
| 75 skills | Written procedures for recurring work, for example test-driven development, bug diagnosis and code review. |
| 12 agents | Role definitions with a fixed model and a narrow job, for example an implementer and a read-only risk reviewer. |
| 56 hooks | Scripts that run automatically around the agents' actions, for example a reminder when a session tries to end without recording what it learned. |
| 163 pages | Two knowledge bases the agents read before they start work in a known area. |
| 1267 commitsOne saved, named step in the code history. | Tracked changes to this setup. |
What still goes wrong
The agents still make mistakes. They misread a requirement, report a check as passed that never ran, or fix a symptom instead of the cause. The rules above target exactly these failures. They make them rarer, they do not remove them. That is why the last two steps are mine.
Check it yourself
- The open-source projects have their full source and commit history on GitHub.
- Some of the skills and mods I use are public in claude-forge.
- Questions about a specific project: [email protected].