16 — Production doctrine: taking a lane to volume
docs/15-LANE-WORKFLOW.md says what the phases are. This says how to actually run one — written 2026-08-19 immediately after taking docade's plush lane from a confirmed reference sheet to 60 generated sculpts in a single session, and derived only from what that session cost.
Justin called it a live-fire exercise, and that is the right frame: everything below was paid for once so it does not have to be paid for again.
The loop
For each batch, in order. Six pieces — one set — is the unit.
brief ──▶ generate ──▶ normalize ──▶ SELF-CHECK ──▶ human verdict
▲ │
└──────────── re-run what failed ────────────┘
1. Brief from the manifest, never from a typed list. art brief --from-manifest --set <s> pulls the key, path, rarity and the client's own brief text together. A subject typed by hand can drift from the foreign key by one character and deliver art the app cannot see.
2. Let the validators run before the estimate. The hue plan is computed and checked inside art brief, so a set that would fail the contract fails for free. Anything else that can be checked before the money belongs there too.
3. Generate, then normalize. Nothing to judge until the cutout has run — the raw frame still has its ground.
4. Self-check. This is the step that pays for itself. Render the batch and look at it against the contract, honestly, before asking for a verdict. Then fix what fails and re-run. Sending someone a batch you have not read is not delegation, it is laundering.
5. Then, and only then, a human verdict. Which Claude does not supply. See "the line" below.
The three levels, and why each is necessary
Each answers a question the one below it structurally cannot.
| Level | Surface | Catches |
|---|---|---|
| Piece | the candidate at judging size | Wrong subject, dead fine features, a coloured contact shadow, a surviving backdrop |
| Set | a section of the lane board | Hue collapse, a tier that does not out-read the tier below it, one piece drifting in register |
| Collection | the same board, whole · art contact <client>/<lane> | Duplicated silhouettes across sets, tier signals that drifted between sets, sets that read as one set |
The judging unit decides the layout; the reviewer's sitting decides the file. art review <client>/<lane> is one page: docs/16's checklist, then a section per set with its siblings on one row, then the pieces. Two wrong answers came first — a board per RUN, which split a set the moment one piece was re-run, and then a board per SET, which satisfied the letter of this table and handed the reviewer seven errands instead of one sitting. Level two needs six siblings side by side. It does not need its own file. 06-REVIEW-UX.md has the detail.
Every one of these caught something real in one session, and the level above was blind to it every time. Two near-identical gold rabbits each passed piece review and set review; they were obvious within seconds of the sixty being on one page. Run all three. The upper two are free and offline.
What to check, concretely
Read the style's qa.must and qa.must_not — they are written to be falsifiable — then these, which generalise:
- Name it at judging size. Downscale to the real size on the real ground colour with no text. If you cannot name the subject, it fails.
- The set spans its hues. Not "are these nice colours" — count them.
- Tiers out-read each other within their own set. Put the top tier beside the baseline. A top tier less dramatic than the tier below it is a rarity inversion and it is a contract failure, not a taste call.
- Nothing thin survived. Every fine feature is a solid mass, and it is in a colour that differs from the body (
reads-at-thumbnail.md). - The cutout is clean. No rectangle, no coloured pad or plinth underneath.
- Nothing duplicates what the lane already holds.
The line Claude does not cross
Generate, self-check, fix, and re-run: yes, without asking. Those are reversible and they are what autonomy is for.
Record an approval: no. machine-approved exists as a distinct non-shipping state precisely so a machine pass can never be mistaken for a person's, and every verdict records verdict_by. A studio that approves its own work has no review loop, it has a render farm.
The same line applies to work already approved: when the collection view exposes a contract break inside a set the client already signed off, flag it, do not re-run it. Three such breaks were found on 2026-08-19 and all three were left alone.
Re-runs
A re-run is cheap (one image) and is almost always right when a piece fails the contract. Three rules:
- Say what failed, in the brief. "The first attempt made the antlers thick but kept them in the body hue, so they merged into the silhouette" produces a different image; "make the antlers thicker" produces the same one. This is
prompt-language.mdapplied to your own last attempt. - Tell a partial re-run what its siblings hold.
--set-huesexists because "fix one piece" otherwise adds a fifth of the same colour to a set of four. - Write the critique onto the superseded sidecar. History groups by reason, and a superseded candidate with no reason is a lost lesson.
Stop after three attempts on one piece and change the hypothesis. Attempts one and two of the jackalope were the same idea said louder. Attempt three changed what was being asked for and worked immediately. If the second attempt fails the same way as the first, the instruction is not too weak — it is wrong.
Batch sizing
Six. Large enough to judge set coherence, which is where most failures live; small enough that a bad batch costs one set and not a lane. Ten batches of six beat six batches of ten for exactly this reason.
Costs, measured: about $0.65 per set of six — roughly $0.25 compile, $0.40 generation — and about $0.19 for a single re-run, most of which is the compile call. Normalize, QA, contact sheets and every validator are free.
When something goes wrong
Ask which level should have caught this, and put the check there:
| Where it belongs | Example from 2026-08-19 |
|---|---|
| Code, before the estimate | The hue plan — it is a constraint on data |
| Craft corpus | "Thick is not sufficient; the feature needs a different colour" |
The lane's lessons.md | "Rare is #4ac8ff" — meaningless to another client |
A npm run check assertion | "No normalized asset carries a pre-normalize verdict" |
| Canon | The three-level review model itself |
If the fix is a sentence you promise to remember, it is not a fix. The set hue plan was moved out of the compiler and into the brief, correctly — but as prose, and it failed again the same way months later. A constraint stated as a sentence will eventually be violated by a sentence.
Before declaring a lane produced
- [ ] Every manifest key has a current candidate —
art contactreports missing. - [ ] The contact sheet has been rendered and looked at at judging size.
- [ ] Every superseded candidate carries a critique.
- [ ] Lessons written; anything that survives the client-agnostic rewrite test promoted to
docs/corpus/craft/. - [ ]
npm run checkpasses,mainclean and pushed. - [ ] Linear matches reality — and what is waiting on a human is filed as an issue, not left in a paragraph.
- [ ]
docs/HANDOFF.mdleads with what is blocked and on whom.
Double-run by default — decided 2026-08-19
When a lane's first full batch lands and is awaiting verdicts, offer a second option from the runner-up model. art ab <client>/<lane>.
Justin: "I am more than fine to actually double run on most matters. I like having the option to choose. It increases quality."
The cost is one extra render per piece — on docade/plush that was $4.75 against a lane that had already spent $24.83, about 19%. What it buys is a preference test at real N, which a bake-off structurally cannot give: a bake-off asks whether a model holds the contract, and a board asks which finished image you want. On plush the bake-off picked Seedream on eighteen pieces and the board picked the incumbent 23–17 on forty, and the preference was set-dependent — 6–0 one way in sunbaked, 4–2 the other in three others.
The limits, because this is not "always generate twice":
- Never on work already judged. A second option on an unjudged piece costs a render. On a judged one it costs the judgement, which is the expensive thing.
- Not a third. Two options is a choice; three is a survey, and the review cost stops being marginal.
- The runner-up has to be a real runner-up. Pick it from the lane's bake-off record, not from novelty.
Reasoning and evidence: docs/corpus/craft/second-option.md, which is loaded into every compile.