Model decision — docade/plush
DECIDED — seedream-5-pro-edit, at four references. Justin, 2026-08-19, after three races: "seedream is the winner here."
He preferred flux2-max-edit on content from the first race and moved off it as the evidence accumulated. That reversal is kept on the record below rather than tidied away, because the reason it moved is the useful part: FLUX returned the wrong animal in two of twelve pieces across two sets — a fox for a pika, a shelled hybrid for a puffin — which is a re-run cost of roughly one piece in six, not a matter of taste. Seedream held the contract on all three sets, at $0.0675 against the incumbent's $0.067.
Four references, not more. Pro at eight references produced an identical result to Pro at four on the same set, and Seedream bills per input image beyond the first, so the extra four cost money and changed nothing.
Not applied retroactively. The 20 pieces already approved keep their Gemini images and their verdicts — seven are promoted and are conditioning later batches as exemplars, and clearing them to re-litigate a settled judgement spends real signal to buy optionality nobody asked for. The 40 still awaiting a verdict are offered as an A/B fork: the Gemini candidate and a Seedream alternate, side by side on the review board, one choice each. art ab docade/plush.
This is the field phase of docs/15-LANE-WORKFLOW.md. The model a lane uses is upstream of its style contract, its cost per piece and every verdict that follows — it is the one choice a sheet can settle in an hour and a wrong guess can cost for a year.
Race — 2026-08-19 · cloudbank
Sheet: clients/docade/bakeoff-plush-cloudbank.png — rows are pieces, columns are models, at 80px. Every model got the same plate, the same three exemplars and the same compiled prompt; compile was paid once and shared, so the prompt is a control rather than a fifth variable. $1.96 total.
| Model | Made | Planned hue | Register | $/image | Verdict |
|---|---|---|---|---|---|
flux2-max-edit | 6/6 | 5/6 | lost | $0.0700 | Justin's first choice, on content |
seedream-5-pro-edit | 6/6 | 6/6 | held | $0.0675 | Second. Held the contract and beat the incumbent |
gemini-flash-image | 6/6 | 6/6 | held | $0.0670 | Incumbent. "Fine" |
nano-banana-2-edit | 6/6 | 6/6 | held | $0.0800 | "Fine." Google model on fal's bill |
qwen-image-3-edit | 0/6 | — | — | — | fal queue timeout. No result, not a verdict |
Justin: "I actually like the flux to max on a content standpoint, and I think I like a sea dream as a second. Both Gemini and Nano Banana are fine."
Re-render this sheet for free: art bakeoff docade/plush/cloudbank --runs r-20260819-679f,r-20260819-f520,r-20260819-a521,r-20260819-108a
What the race actually established
Palette transfers across model families; register does not. Four architectures, one plate, 23 of 24 pieces on the planned hue. The same plate carried the colour everywhere and the category nowhere: FLUX returned a naturalistic dragonfly and a shelled hybrid where three others returned stuffed toys. Promoted to docs/corpus/craft/reference-conditioning.md — it is client- agnostic and every future lane should assume it.
The analysis called FLUX a failure and the human called it the winner, and both are right about different things. The measurement — register lost — is correct and stands. What the analysis did wrong was treat the register as the goal when it is a contract someone wrote and can rewrite. docs/corpus/craft/measurement-is-not-a-verdict.md exists for this exact move.
The billing question and the model question were separable, and they got separated. nano-banana-2-edit is a Google model on fal's bill; it came back near-identical to the incumbent at 19% more. So "should we be on fal" was never a quality question. It is answered on other grounds — see below.
What has to happen before this can ship
- The visual mission describes plush toys. This model does not make plush toys. Either
clients/docade/visual-mission.mdand the style contract are amended to describe what it does make, or every audit will correctly flag a lane doing exactly what was asked of it. Drifting silently between those two is not an option. docade's call, not WootBuild's — advisory owed. - Register consistency is unproven on this model. It lost the register with four references; nothing suggests it will hold one across sixty pieces. This is what reopens ATL-3: a LoRA stops being a cost optimisation and becomes the mechanism, because it is the only remaining lever that could give FLUX's content quality and a consistent register. Verify the FLUX.2 trainer slug before costing it.
- One race is one set. Six pieces on
cloudbankis enough to overturn a default and not enough to commit sixty. A second set on a different subject mix is the cheap way to find out whether the preference holds.
Until those are settled, seedream-5-pro-edit is the shippable choice: it held the contract, was preferred over the incumbent, and costs half a cent more than what is already being paid.
Why fal at all
Not a quality argument, and it should not be dressed as one. fal Agent Pro is $200/month with $200 of credits that roll over and are spent before pay-as-you-go rates. The subscription is already being paid, so credits accumulate whether or not they are used. Justin, 2026-08-19: "we're going to be using fal at least until these credits run out... we are going to cash that out, even if we go back to Gemini in the long run. It's not necessarily what is best."
That is a deliberate, stated trade and it is recorded as one — so that whoever reads this in six months knows the routing was a budget decision with a known expiry, not a finding about model quality. When the credits are drawn down, re-read this file before renewing the decision.
Race — 2026-08-19 · stonepeak — the second set
Sheet: clients/docade/bakeoff-plush-stonepeak.png. Chosen for maximum contrast with cloudbank: heavy horned mammals against birds and an insect, and it contains mountain-goat, a documented hard case — the earlier attempt was superseded with "a long flat-sided rectangular body that read as a crate with a head attached." $1.96.
Same plate, same exemplars, same compiled prompt, compile paid once.
| Model | Made | Planned hue | Register | $/image |
|---|---|---|---|---|
seedream-5-pro-edit | 6/6 | 5/5 | held | $0.0675 |
gemini-flash-image | 6/6 | 5/5 | held | $0.0670 |
nano-banana-2-edit | 6/6 | 5/5 | held | $0.0800 |
flux2-max-edit | 6/6 | 4/5 | lost | $0.0700 |
qwen-image-3-edit | 0/6 | — | — | queue timeout again |
Re-render free: art bakeoff docade/plush/stonepeak --runs r-20260819-6a69,r-20260819-202b,r-20260819-a4f3,r-20260819-cd99
The cloudbank result repeated, and hardened
Species drift on flux2-max-edit is now a pattern, not an incident. Two sets, twelve pieces, two wrong animals:
| asked for | returned | |
|---|---|---|
| cloudbank | puffin | a shelled, duck-billed hybrid |
| cloudbank | dragonfly | a naturalistic insect, not a toy |
| stonepeak | pika | a fox, with a bushy fox tail |
| stonepeak | marmot | a dark olive realistic groundhog, 77° off its planned green |
Every other model returned the right animal in the right register both times. This is the same finding as craft/reference-conditioning.md: the plate carries palette and does not carry category — and on this family it does not carry species either.
seedream-5-pro-edit held on both sets, with the fullest, roundest forms of the four at judging size. Its mountain-goat is a rounded sitting plush, which is the piece the lane has already failed once.
One measurement that is NOT a finding
All four models read 148° off on griffin. That is the measure, not the models: griffin is two-tone — violet body, gold wings — and the wings cover more area, so "dominant saturated hue" picks the accent. Recorded here so nobody re-derives it as a model failure. A hue check that assumes one hue per piece is wrong for every legendary in this lane, and the hue columns above are scored out of 5 for that reason.
Average fill of the 80px tile — gemini 42.8%, seedream 45.9%, nano-banana 47.0%, flux 47.1% — is too close between the top three to carry weight, and is recorded as weak. Gemini being consistently lowest is the only part worth noting.
Still open
Nothing here settles the mission question. It sharpens it: if flux2-max-edit is the pick, the contract has to describe an illustrated collectible and accept that roughly one piece in six comes back as a different animal, which is a re-run cost, not a style question. That is the number docade should be given.
Race — 2026-08-19 · canopy — the favourite, stress-tested
Not a search for a new winner. Three questions about the one that is emerging: is Pro worth its premium over Lite, do more references help, and is there a second FLUX path that keeps the content without the species drift?
Set chosen because it holds orchid-mantis — an insect, which is exactly where flux2-max-edit drifted first.
| Column | Refs | Made | Planned hue | $/image |
|---|---|---|---|---|
seedream-5-pro-edit | 4 | 6/6 | 6/6 | $0.0675 |
seedream-5-pro-edit@7 | 8 | 6/6 | 6/6 | $0.0675 + per-input |
seedream-5-lite-edit | 4 | 6/6 | 5/6 | ~$0.034 |
flux2-flex-edit | 4 | 0/6 | — | — |
Re-render free: art bakeoff docade/plush/canopy --runs r-20260819-cb58,r-20260819-226c,r-20260819-7228
More references bought nothing
Pro at four references and Pro at eight produced the same result: 6/6 on hue both times, 45.1% vs 44.6% average tile fill, and no visible difference at judging size. A null result, and a useful one — Seedream bills $0.0045 per input image beyond the first, so the extra four references cost money and changed nothing. Four (plate + three exemplars) stands as the lane's budget.
This is the one variable that had never been tested in three races, and testing it cost one column.
Lite is cheaper and it shows
5/6 on hue: it rendered sloth at 110° — a yellow-green — where the plan asks for 147° and Pro landed on 150°. Visibly flatter across the set, with less volume in the forms. It is a real option if cost ever dominates, and it is not the same product.
flux2-flex-edit: no result
0 of 6, every call HTTP 504 Generation timeout from fal — a generation failure on their side, not our polling and not a queue wait. So the question it was wired to answer — is FLUX's species drift an endpoint quirk or a family property? — is still open. Retry before drawing any conclusion; an infrastructure failure is not evidence about a model.
The hue check was wrong, and is now code
Two races reported every model ~150-174° off on the rare and legendary pieces — griffin, then orchid-mantis. Both readings were wrong. The measure asked for the dominant saturated hue, and a rare piece is two-tone by contract: a body in the planned hue plus an accent that signals the tier. On both, the accent out-areas the body, so "dominant" picked the signal colour and called the contract broken.
A measurement that fails on exactly the pieces the contract cares most about is worse than none, because it fails quietly and in the direction of alarm. It now asks about presence — does a meaningful share of the image carry the planned hue — and lives in lib/pipeline/bakeoff.mjs as hueCompliance() rather than being retyped each race. Under it, orchid-mantis passes at 37.7% and Lite's yellow-green sloth still correctly fails.
The ledger under-reports fal by a third
Reconciled against the account: our ledger says $4.44 of fal spend, the balance says $5.96 has actually left. A 34% gap.
Two known mechanisms, neither modelled by the estimator: Seedream charges $0.0045 per input image beyond the first, and flux2-max-edit bills input images as processed megapixels — so a 4-reference brief costs roughly $0.19 there rather than the $0.07 in our table. The gap is not separable to a single model from one balance reading, and fal's balance updates with a lag, so per-call differencing does not work either.
Nothing about model choice changes. What changes is that any cost comparison between a per-input-billed model and a flat-rate one is currently wrong in the flat-rate model's favour.
The fork resolved — 2026-08-19 · and the decision above is now half wrong
40 pieces, both options side by side, one choice each. Gemini 23, Seedream 17.
The lane's decision above says seedream-5-pro-edit, on three races of six pieces each. At forty pieces, on the actual board, at judging size, the incumbent won more of them. That is not a small margin flipping — it is 58/42 against the model this lane was moved to.
| set | gemini | seedream | |
|---|---|---|---|
sunbaked | 6 | 0 | swept |
meadow | 3 | 1 | |
cloudbank | 4 | 2 | |
stonepeak | 4 | 2 | |
frostline | 2 | 4 | |
canopy | 2 | 4 | |
emberfall | 2 | 4 |
What this actually establishes
It is set-dependent, and that is the finding. Justin, 2026-08-19, before any of this was measured: "there may be some surfaces that are a hybrid — unlikely, but one model may work for one surface. Another may turn beneficial on the next." Seven sets, and four of them are 4–2 or better in one direction. Sunbaked is 6–0.
That is the hypothesis surviving a test it could have failed, and it did not come from the bake-off — three sets of six could not have shown it. It took the production board.
What it does NOT establish
Not that the bake-off was wrong. The races measured contract-holding — planned hue, register, the right animal — and Seedream won those cleanly while flux2-max-edit returned a fox for a pika. This measured preference between two finished images, which is a different question and the one that ships.
Not that the lane should move back. Gemini at 23 and Seedream at 17 is much closer to a coin-flip than to a verdict, and every one of these 40 pieces exists because Seedream was cheap enough to offer as a second option at all.
What to do about it — nothing yet, deliberately
The lane is 60/60 approved and deliverable. Changing the routed model now would re-open sixty settled judgements to answer a question nobody is asking.
The live question is the next lane, and the answer is procedural rather than about these two models: a bake-off of six pieces predicts a forty-piece board less well than we assumed. The A/B fork is what caught that, and it cost $4.75 against a lane that had already spent $24.83. Offering a second option on work already made is the cheapest large-N model test available, and it should be the default at the point a lane's first full batch lands — not a one-off that happened here because the work existed anyway.
Recorded as craft: docs/corpus/craft/.