Built with IBM Bob 2.0 · solo entry
Plumbline gets IBM Bob to read your test plan, spec and release checklist, find the test behind each claim, and then break the code to see if the test notices. If it stays green on broken code, it's a test in name only.
01 · The problem
A while back I shipped a project whose CI badge was green for days while it ran no tests at all. The job was skipping every test, waiting on a credential that didn't exist, and nobody looked because it was green. Test plans have the same problem, just in a spreadsheet. Someone signs a checklist before every release saying every case has a passing test and the spec matches the code, and nobody really checks.
02 · A real audit
Turnstile is a small auth service I wrote for this, and I wrote down every defect in an answer key before Bob saw it. Its test plan says all 84 automated cases pass, and every item on its release checklist is ticked.
03 · Break it. Does the test notice?
test('TC-41 rejects an expired access token', () => {
const t = issueAccess('ada@example.com', ['user'], { now: NOW });
const result = verifyAccess(t, { now: NOW + 3 * 60 * 60 * 1000 });
assert.ok(result); // always true: result is an object
});
- if (claims.exp <= t - CLOCK_SKEW_SECONDS) return { valid: false, reason: 'expired' }; + if (false) return { valid: false, reason: 'expired' };
const t = issueAccess('ada@example.com', [], { now: NOW });
const result = verifyAccess(t, { now: NOW + 2 * 3600 * 1000 });
return result.valid === false && result.reason === 'expired';
// real code: true broken code: false
The expiry check is gone, the witness proves it, and TC-41 still passes, along with the rest of the suite. Bob found 8 tests like this, including 2 I'd missed in my own answer key.
04 · How it works
Copy .bob/ into your repository, open it in IBM Bob, switch to the Plumbline mode and ask it to audit.
The .xlsx plan, .docx spec and .pdf checklist become claims, each with its source cell or clause.
One subagent per test file finds the test behind each case, even when its name carries no id.
A subagent that never sees the test writes a change that makes the claim false, plus a witness. The runner applies it in a throwaway git worktree.
Every spec clause and checklist tick against the code, the changelog and git diff, one line of evidence each.
A self-contained HTML report that leads with what is false.
05 · No accusation without proof
UNPROVEN. Witness false on the real code: WITNESS_INVALID. Witness still true on the broken code: WEAK_MUTATION.| Cases mapped to the same test as the key | 83 / 84 |
| Spec clauses judged as the key judges them | 24 / 24 |
| Checklist items judged as the key judges them | 10 / 10 |
| Tests in name only the key had missed | 2 |
06 · How IBM Bob is used
The Plumbline mode: its role and the tools it may use. .bob/custom_modes.yaml
The five-stage procedure, the claim schema and the cost rules. .bob/skills/plumbline/SKILL.md
Reads the plan and the spec as they are. Its file tools couldn't read the PDF, so Bob wrote a small reader for it.
One per test file to map cases to tests, and one per test file to write mutations and witnesses, with the test withheld.
Runs the mutation runner, and read-only git to judge the spec and the checklist against history.
Bob planned and wrote the mode, the skill, the runner, the report and their tests, in 8 exported tasks.
07 · The free half, for any repository
23 of 100
of the most-starred installable repositories on GitHub owned by organisations fail a check against their own README: dead links, missing files, a licence that isn't there. No AI involved, just the checks a machine can settle, each with a negative control to prove it can fail. Measured on 25 September 2026, and I checked every one by hand.