Calibration — docade
Stage 4.5. A sandbox. Nothing here is canon: no style version, no canary, no delivery, no exemplar promotion. Cheap by construction — the point is information per dollar, not quality.
Opened 2026-08-18. Status: open. Sections 2 and 3 are one register deep out of five; the exit condition in §7 is not met yet.
1. Pipeline shakedown
Does one brief survive the whole path? Yes — first attempt, on the day the spine was built.
| Step | Works? | Notes |
|---|---|---|
| Brief → compile | ✅ | One Claude call for the whole batch. Read the visual mission with no style sheet present and derived direction from it correctly — named the committed palette hexes, targeted principles 1–4 by name, and worked out on its own that fal-ai/flux-2 takes no negative prompt, so it emitted an empty negative and phrased every constraint positively. |
Route (art route --explain) | ✅ | Correctly eliminated gemini-flash-lite-image on the alpha constraint. It then routed to a model whose alpha flag was wrong — see §4.1. The mechanism worked; the data it ran on was false. |
| Generate | ✅ | 3/3 at 2–3s each. Concurrency, per-candidate failure isolation and resume all exercised for real. |
| Candidate + provenance sidecar | ✅ | Full sidecar per candidate: prompt, seed, prompt hash, spec, model, cost, elapsed, QA result, lessons applied. art reindex rebuilds every DB row from these. |
| Ledger entry + estimate accuracy | ✅ | Every image ledgered as it landed, before anything downstream could fail. See the estimate gap below. |
| Review surface / placement harness | ✅ | Built this session (art proof). Renders candidates at 80px on #262c4a, which is the only ground on which principles 1 and 2 mean anything. |
Failure isolation was tested involuntarily and passed. The first attempt hit HTTP 403 — User is locked. Reason: Exhausted balance on all three candidates. Each failure was recorded to its own sidecar, the compile was checkpointed, the run was marked failed, and the ledger recorded the $0.0755 of compile that had genuinely been spent. On re-run after the account was funded, the compile resumed from disk at no cost and only the three generation slots re-ran. That is the resume path working under a real fault, not a simulated one.
Estimate vs actual.
| Run | Estimated | Actual | Gap |
|---|---|---|---|
r-20260818-42e1 (3 × flux2-dev) | $0.1260 | $0.1115 | −12% |
The −12% hides a much larger structural error in both directions. The estimator charges OVERHEAD_PER_IMAGE = $0.03 per image. The real compile is one call per run, and it cost $0.0755 — 2.5× the per-image allowance on a 3-image run, and it would have been the same $0.0755 on a 300-image run.
Compile cost scales with context size, not with image count. docade's visual mission is ~8KB and goes into every compile. On a 300-image batch the estimator would over-charge overhead by ~$9; on a 1-image probe it under-charges by ~2.5×.
Recommendation: split the overhead term into a per-run compile cost (estimated from context size) and a per-image QA cost. Filed as ATL-14.
2. Model behaviour, per register
| Register | Model | What it does well | Where it fails | Verdict |
|---|---|---|---|---|
| Collectible object (plush) | flux2-dev, no reference | Silhouette legibility at 80px; consistent form across seeds; real weight and contact; fast (2–3s) and cheap ($0.012) | Renders photoreal product photography, not flat-fill illustration. Baked painterly lighting. No transparency, and no reference input either. | Not usable for this register. Not a cost problem — a steerability problem. |
| Collectible object (plush) | recraft-v4-svg | Principle 3 by construction — 26 paths, 7 flat solid fills, zero gradients. True alpha, no background to remove. A recolor is one attribute edit | Sticker-flat. No weight, no volume, reads as a logo — the mission's other named non-goal | Right answer for the wrong register. Keep it for glyph and avatar |
| Collectible object (plush) | gemini-flash-image + 1 reference plate | Flat separable fills and volume. Sits, has occlusion, reads cleanly at 80px. $0.067, ~8s | First plate carried near-black fills; output inherited a keyline and a rim light despite the prompt forbidding both by name | The register is reachable, and the fault was the plate |
| Collectible object (plush) | gemini-flash-image + retinted plate | Keyline gone, rim gone, flat fills held, weight held, reads at 80px. Principles 1–4 all pass | Still opaque (ATL-16). Small eye speculars, which are chrome | Pass. ATL-20 |
| A whole themed SET of six | gemini-flash-image + one plate | Six different creatures that read as one artist: consistent simplification, consistent light direction, consistent shading discipline. All six nameable at 80px | Untested for rarity legibility (principle 5) | Pass. ATL-19 — one plate holds a set |
| Small functional glyph | — | not yet probed | ||
| Scene | — | not yet probed | ||
| Character (Lyle) | — | not yet probed | ||
| Type | — | not yet probed | ||
| Thing on a saturated ground (prize) | — | not yet probed |
The finding that matters: register is not steerable by adjective on flux2-dev. The compiled prompt did everything right — it demanded "flat separable fills", put shading in a "separate simple pass", said "clean vector-like flat-fill illustration", said "no rim light", and named the exact hex fills. The model produced soft-lit studio plush photography anyway, three times out of three.
This is docade's own visual mission §6 predicting its own failure a day in advance: "Not painterly. A beautifully rendered plush with baked lighting violates principle 3 and quietly destroys the growth model. This is the most likely wrong answer, because it is what image models do best by default."
It was right, and prompt engineering did not beat it.
3. What the prompt actually needs
| Wanted | Naive prompt gave | What closed the gap |
|---|---|---|
| Flat fills, shading as a separate pass (principle 3) | Baked painterly shading, fabric nap, physically-plausible soft light | A reference plate. Explicit instruction failed on flux2-dev; one flat vector plate handed to gemini-flash-image moved the register in a single pass. See §4.6. |
| Volume and weight with flat fills | The vector plate gave fills without weight; the photoreal pass gave weight without fills | Naming both halves in the brief — what to KEEP from the plate (fill structure, shading discipline) and what to OVERRIDE (give it volume, it is not a sticker) |
| Reads at 80px (principle 1) | Passes — one big dome, eight thick arms, no thin features | The prompt naming the 80px judgement and forbidding thin/antler-like features. Cheap and it worked. |
| Weight, sits on something (principle 4) | Passes — arms flatten under load, broad underside occlusion | Describing how weight shows (arms flattening, occlusion) rather than asking for "weight". |
| Palette discipline | Passes — used only #4ac8ff, #f5f3ff, #262c4a | Giving the compiler the committed palette and forbidding invention. No drift observed. |
| Transparent cutout | Opaque PNG on a dark ground | Impossible on this model. See §4.1. |
4. Process findings
These are worth more than the images.
4.1 The router carried a capability that does not exist
flux2-dev was flagged alpha: true. No FLUX.2 endpoint has a transparency parameter. Verified 2026-08-18 against each endpoint's own OpenAPI schema — fal-ai/flux-2, -pro, -flex and -max/edit all accept prompt, image_size, seed, guidance, steps, output_format and safety flags, and nothing resembling a background or alpha option. Neither gpt-image-2 endpoint on fal exposes one either.
This is the same bug class as the phantom slugs that ATL-1 caught, one level up: not a model id nobody verified, but a capability nobody verified. It was worse than the slugs, because a phantom slug 404s immediately and loudly, while a phantom capability silently produces 18 unusable plush assets that look fine until someone puts one on a different ground.
Fixed: every raster model is now alpha: false; only true-vector models claim transparency. npm run check asserts that invariant, and asserts that a transparent raster request fails loudly rather than resolving.
The adapter was also sending two parameters that do not exist — transparent_background and negative_prompt — which fal silently dropped. The negative one is the more insidious: a compiler that believes negatives are being honoured writes a weaker positive prompt.
4.2 art doctor reports auth, not spendability
fal's health probe returned green while the account balance was $0.00 and every generation call 403'd with User is locked. Reason: Exhausted balance. A key can be perfectly valid on an account that cannot spend.
https://rest.alpha.fal.ai/billing/user_balance returns the balance as a bare number for a read-only GET. Doctor should show it. Filed as ATL-15.
4.3 The estimator's cost model is shaped wrong
See §1. Overhead is charged per image; compile is per run and scales with context. Filed as ATL-14.
4.4 Two subjects with the same name
Compiled prompts were originally matched back to subjects by name. docade's piece 2 is octopus | coral colorway and octopus | teal colorway — the exact run that settles Q44 — and name-matching collapsed both to one prompt. Caught before spending, by writing the check first. Now matched by index, with a guard that refuses a batch whose distinct subjects compiled to identical prompts.
4.6 Reference conditioning works where prompting did not — ATL-17
The probe, run 2026-08-18 for $0.42 all in.
Prompting flux2-dev to produce flat separable fills failed three times out of three. So the probe asked a different question: give the model an example instead of an adjective.
recraft-v4-svgproduced a flat vector plate — 26 paths, 7 solid fills, zero gradients, true alpha. Principle 3 satisfied by construction. It also fails principle 4 completely: it is a sticker.- That plate, rasterised to 512px, was handed to
gemini-flash-imageas a reference alongside a prompt that named explicitly what to keep (fill structure, shading as its own darker-value pass, no gradients) and what to override (give it the volume and weight of a stuffed object). - Two candidates came back with flat separable fills and volume — the space between the two failure modes the mission predicted.
At 80px on #262c4a, the conditioned pair is the only one of the three that reads cleanly and has weight. Comparison sheet: runs/r-20260818-4648/proofs/atl-17-comparison.png.
This is the difference that matters: flux2-dev did not respond to explicit instruction at all, so there was nothing to iterate on. The conditioned output's remaining faults — a dark keyline, specular highlights, a cyan rim on candidate b — are each already named as non-goals in the visual mission, which means they are addressable in the next prompt or as a lesson. A model that ignores the brief is a dead end; a model that obeys it imperfectly is a starting point.
A capability gap this exposed. No FLUX.2 text-to-image endpoint accepts reference images — only fal-ai/flux-2-max/edit has image_urls. The router claimed refs: 10 for dev, pro and flex. Same audit also found the base fal-ai/flux-2 has no loras field: FLUX.2 LoRA inference is a separate endpoint, fal-ai/flux-2/lora, now its own model entry. And the trainer's parameters were the FLUX.1 trainer's — training_style defaults to "subject", so a plush style set would have silently trained a subject LoRA. All three would have surfaced only after spending a training run.
4.9 The plate wins arguments with the prompt — ATL-20
The first conditioned run produced a dark keyline and a rim light. The obvious reading was that the brief had failed to forbid them. It had not. The compiled prompt said, verbatim:
No black outline or keyline anywhere; separation from #262c4a comes from value and local colour alone. No rim light and no back-edge glow of any kind. No gradients, no airbrush, no soft falloff, no painterly brushwork.
The model produced a keyline and a rim anyway — because the plate had them. recraft-v4-svg had used two near-black fills (rgb(27,1,64), rgb(15,4,55)) to do contour and deep-shadow duty, and the model elaborated those into lighting.
The durable rule: when a reference plate and the prompt disagree, the plate wins. Negations lose to images especially badly. You cannot talk a model out of something its reference is showing it — you have to fix the reference.
So the plate was retinted (scripts/tint-plate.mjs): its fills were remapped onto a six-step ramp of a single hue from docade's committed palette, where the darkest step is a dark violet rather than near-black. Re-run on the corrected plate: keyline gone, rim gone, flat fills held, weight held, reads at 80px.
4.10 Q44 is answered, and the answer is yes — a recolor is free
The retint is a deterministic offline transform of a flat vector plate. It runs in milliseconds, calls no model, and costs $0.00. Two colorways of the same sculpt, produced from one source: runs/r-20260818-0266/proofs/atl-20-tint.png.
The shading survives the recolor because the shading is a darker value of the same hue rather than baked light — which is precisely what principle 3 asks for, now demonstrated rather than asserted. This is docade's Q44, and with it the economics of ~150 new images a year: a colorway is an attribute edit.
Note what it depends on: flat separable fills in a vector source. It would not work on any of the photoreal candidates, and that is the whole argument for principle 3 in one sentence.
4.11 One plate holds a whole set — ATL-19
The six real deepsea keys from docade's manifest — octopus, jellyfish, crab, turtle, otter, whale — conditioned on the same violet plate, one candidate each, $0.62 for the set.
The manifest's hard constraint on this category is "Must read at 80x80px — Critter Match uses these as card faces", and they are seen together on a grid, so the question was never whether any one is good. It was whether all six look like the same artist made them.
They do. Consistent level of simplification, consistent light direction, consistent shading discipline, consistent palette, no keylines or rims anywhere. All six are nameable at 80px, including the jellyfish — the mission's flagged hardest case for value-only separation, which works because its bell is simply a lighter value.
Proof: runs/r-20260818-33b3/proofs/at-80-on-262c4a.png.
What this settles for ATL-3. A reference plate was supposed to teach one look while only a LoRA could teach the look across many subjects. On this evidence the plate does both. The LoRA is now an optimisation rather than a prerequisite.
What it does not settle: one set, one seed each, one prompt. Rarity legibility (principle 5) is untested, and so is whether three different themed sets hold together with each other.
4.5 What the canon got right
The visual mission's §6 non-goals predicted the exact failure mode, in advance, in writing. Principles stated as falsifiable tests ("downscale to 80px, if you can't name the animal it fails") made the verdict a five-second observation rather than an argument. Keep writing them this way.
5. Corpus write-backs
| Finding | Written to | Done |
|---|---|---|
| No transparency parameter on any FLUX.2 endpoint; no negative prompt | docs/corpus/models/flux2-dev.md + pro/flex/max-edit | ✅ |
| Same, gpt-image-2 pair and ideogram-v4 | docs/corpus/models/gpt-image-2*.md, ideogram-v4.md | ✅ |
background_color ≠ alpha; do not infer from the SVG sibling | docs/corpus/models/recraft-v4.md | ✅ |
| flux2-dev renders photoreal for "plush" and does not respond to flat-fill instruction | docs/corpus/models/flux2-dev.md | ✅ |
| First real verdict history entry for any model | docs/corpus/models/flux2-dev.md | ✅ |
6. Spend
| Calibration ceiling | $5.00 per run · $75.00 per month (raised to $200 on 2026-08-18) |
Plush baseline (r-20260818-42e1) | $0.1115 |
Vector plate (r-20260818-0266) | $0.1572 |
Reference-conditioned probe (r-20260818-4648) | $0.2517 |
Corrected-plate re-run, ATL-20 (r-20260818-f2bd) | $0.2273 |
Six-subject set, ATL-19 (r-20260818-33b3) | $0.6204 |
| Plate retint to two colorways | $0.0000 — offline transform |
| Diagnostic endpoint probes, outside the pipeline | $0.2220 — see §4.7 |
| Month to date | $1.59 of $75.00 |
4.7 Diagnostic probes bypass the ledger
Confirming that a fal 403 was transient meant POSTing to five endpoints with curl. Those calls generated real images and cost $0.222, and none of it went through art, so none of it was ledgered until it was entered by hand.
The ledger is authoritative for spend (00-CHARTER.md), so an unrecorded call is a correctness bug, not an untidiness. Two candidate fixes: an art probe <endpoint> command that submits a minimal job and ledgers it, or simply never POSTing to a generating endpoint outside the CLI. The schema probes — which are free GETs against the OpenAPI document — are fine and should stay the default diagnostic.
4.8 The estimate gap is widening in the predicted direction
| Run | Estimated | Actual | Gap | Compile |
|---|---|---|---|---|
r-20260818-42e1 | $0.1260 | $0.1115 | −12% | $0.0755 |
r-20260818-0266 | $0.1100 | $0.1572 | +43% | $0.0772 |
r-20260818-4648 | $0.1940 | $0.2517 | +30% | $0.1177 |
Three runs, and the compile cost rose 56% as the context grew — the reference section and a larger corpus entry went in. Generation cost was predicted almost exactly every time. The error is entirely in the overhead term, exactly as ATL-14 says: it is modelled per-image when it is per-run and scales with context.
7. Ready to forge?
- [x] One brief ran end-to-end and landed a candidate with full provenance
- [ ] We can say what each routed model does for this client's registers — 1 of 6, but that one now has five model rows, a working method, and a set that holds
- [x] The estimate is trustworthy, or we know by how much it is not — we know exactly how, and why
- [x] Nothing here was promoted into the library by accident — the library is still empty by construction
Not ready to forge docade-plush@1. Forging a style sheet now would encode a register the routed model cannot hit, and freeze it.
Open questions carried forward:
How do we get flat separable fills at all?Answered 2026-08-18: reference conditioning, on a model that accepts references.gemini-flash-image- one flat vector plate. Next iteration must kill the keyline, the speculars and the rim — all three are named non-goals, so they are prompt work, not model work.
Does the plush lane still need a LoRA?Not to reach the register, and not to hold a set. ATL-19 showed one plate carrying six creatures coherently. WootBuild's recommendation on ATL-3 is do not train yet — revisit if a set fails to hold, or if three themed sets drift against each other. Note the cost argument is not the deciding one: at docade's 150 images/year, gemini at $0.067 is ~$10/yr against ~$2/yr for a trainedflux2-lora. Consistency is the only thing worth $5 and a training run here, and references are currently supplying it.- The reusable recipe, and the thing worth keeping: vector plate → retint → condition. $0.08 once for the plate, $0.00 per colorway, $0.067 per conditioned render. It is cheap, it is inspectable, and every step is a file.
- Still open on the plush register: rarity legibility (principle 5, untested), whether three themed sets hold against each other, and the cutout (ATL-16) — every candidate so far is opaque.
- Transparency. Advisory 03. Blocks 6 of 10 lanes as specced, and blocks the 18 derived silhouettes entirely.
- Their Q44 (is a recolor cheap) is still unanswered. Piece 2 needs two colorways of the same sculpt, which needs either a seed-locked recolor or an edit pass — not two independent generations. Design that probe before running it; two separate octopuses would answer a question nobody asked.