plumbline

Built with IBM Bob 2.0 · solo entry

Is your test plan still true?

Plumbline gets IBM Bob to read your test plan, spec and release checklist, find the test behind each claim, and then break the code to see if the test notices. If it stays green on broken code, it's a test in name only.

01 · The problem

Green isn't the same as true.

A while back I shipped a project whose CI badge was green for days while it ran no tests at all. The job was skipping every test, waiting on a credential that didn't exist, and nobody looked because it was green. Test plans have the same problem, just in a spreadsheet. Someone signs a checklist before every release saying every case has a passing test and the spec matches the code, and nobody really checks.

02 · A real audit

The plan says 100%. The suite's green. Here's what Bob found.

Turnstile is a small auth service I wrote for this, and I wrote down every defect in an answer key before Bob saw it. Its test plan says all 84 automated cases pass, and every item on its release checklist is ticked.

84automated cases in the plan
60with a test Bob could find
24with no test at all
8tests in name only
4 of 24spec clauses false
4 of 10checklist ticks false
0accusations without proof

03 · Break it. Does the test notice?

TC-41 says expired tokens are rejected. Delete the check and it stays green.

test/tokens.test.mjspasses
test('TC-41 rejects an expired access token', () => {
  const t = issueAccess('ada@example.com', ['user'], { now: NOW });
  const result = verifyAccess(t, { now: NOW + 3 * 60 * 60 * 1000 });
  assert.ok(result);   // always true: result is an object
});
the change Bob wrote, without seeing the testsrc/tokens.mjs
- if (claims.exp <= t - CLOCK_SKEW_SECONDS) return { valid: false, reason: 'expired' };
+ if (false) return { valid: false, reason: 'expired' };
the witness Bob wroteproof the change broke the claim
const t = issueAccess('ada@example.com', [], { now: NOW });
const result = verifyAccess(t, { now: NOW + 2 * 3600 * 1000 });
return result.valid === false && result.reason === 'expired';
// real code: true   broken code: false
Test in name only

The expiry check is gone, the witness proves it, and TC-41 still passes, along with the rest of the suite. Bob found 8 tests like this, including 2 I'd missed in my own answer key.

04 · How it works

A Bob mode and a skill. Five stages.

Copy .bob/ into your repository, open it in IBM Bob, switch to the Plumbline mode and ask it to audit.

  1. Read

    The .xlsx plan, .docx spec and .pdf checklist become claims, each with its source cell or clause.

  2. Map

    One subagent per test file finds the test behind each case, even when its name carries no id.

  3. Break

    A subagent that never sees the test writes a change that makes the claim false, plus a witness. The runner applies it in a throwaway git worktree.

  4. Judge

    Every spec clause and checklist tick against the code, the changelog and git diff, one line of evidence each.

  5. Report

    A self-contained HTML report that leads with what is false.

05 · No accusation without proof

Mutation testing can blame a good test. So Bob has to prove it.

  • Sometimes a change looks like a break but isn't. My first run blamed 4 tests out of 12 that way. So now every accusation needs a witness.
  • No witness: UNPROVEN. Witness false on the real code: WITNESS_INVALID. Witness still true on the broken code: WEAK_MUTATION.
  • Strip every witness out of the same run and it accuses nobody. I tested that, rather than just saying it.
Bob's audit against the answer keynode examples/score.mjs
Cases mapped to the same test as the key83 / 84
Spec clauses judged as the key judges them24 / 24
Checklist items judged as the key judges them10 / 10
Tests in name only the key had missed2

06 · How IBM Bob is used

Bob didn't just build it. Plumbline runs inside Bob.

01Custom mode

The Plumbline mode: its role and the tools it may use. .bob/custom_modes.yaml

02Skills

The five-stage procedure, the claim schema and the cost rules. .bob/skills/plumbline/SKILL.md

03Document understanding

Reads the plan and the spec as they are. Its file tools couldn't read the PDF, so Bob wrote a small reader for it.

04Subagents

One per test file to map cases to tests, and one per test file to write mutations and witnesses, with the test withheld.

05Agent mode with commands

Runs the mutation runner, and read-only git to judge the spec and the checklist against history.

06Bob as the builder

Bob planned and wrote the mode, the skill, the runner, the report and their tests, in 8 exported tasks.

07 · The free half, for any repository

23 of 100

of the most-starred installable repositories on GitHub owned by organisations fail a check against their own README: dead links, missing files, a licence that isn't there. No AI involved, just the checks a machine can settle, each with a negative control to prove it can fail. Measured on 25 September 2026, and I checked every one by hand.

  • react/create-react-app still sends people to documentation at an address that stopped working.
  • kubernetes/kubernetes links to a case-studies page that is a 404.
  • laravel/laravel refers to a licence file that is not at the repository root.
  • Paste any public repository