WootBuild.

Studio/docade/Calibration

Stage 4.5 · calibration.md

Calibration — docade

Stage 4.5. A sandbox. Nothing here is canon: no style version, no canary, no delivery, no exemplar promotion. Cheap by construction — the point is information per dollar, not quality.

Opened 2026-08-18. Status: open. Sections 2 and 3 are one register deep out of five; the exit condition in §7 is not met yet.


1. Pipeline shakedown

Does one brief survive the whole path? Yes — first attempt, on the day the spine was built.

StepWorks?Notes
Brief → compile✅One Claude call for the whole batch. Read the visual mission with no style sheet present and derived direction from it correctly — named the committed palette hexes, targeted principles 1–4 by name, and worked out on its own that fal-ai/flux-2 takes no negative prompt, so it emitted an empty negative and phrased every constraint positively.
Route (art route --explain)✅Correctly eliminated gemini-flash-lite-image on the alpha constraint. It then routed to a model whose alpha flag was wrong — see §4.1. The mechanism worked; the data it ran on was false.
Generate✅3/3 at 2–3s each. Concurrency, per-candidate failure isolation and resume all exercised for real.
Candidate + provenance sidecar✅Full sidecar per candidate: prompt, seed, prompt hash, spec, model, cost, elapsed, QA result, lessons applied. art reindex rebuilds every DB row from these.
Ledger entry + estimate accuracy✅Every image ledgered as it landed, before anything downstream could fail. See the estimate gap below.
Review surface / placement harness✅Built this session (art proof). Renders candidates at 80px on #262c4a, which is the only ground on which principles 1 and 2 mean anything.

Failure isolation was tested involuntarily and passed. The first attempt hit HTTP 403 — User is locked. Reason: Exhausted balance on all three candidates. Each failure was recorded to its own sidecar, the compile was checkpointed, the run was marked failed, and the ledger recorded the $0.0755 of compile that had genuinely been spent. On re-run after the account was funded, the compile resumed from disk at no cost and only the three generation slots re-ran. That is the resume path working under a real fault, not a simulated one.

Estimate vs actual.

RunEstimatedActualGap
r-20260818-42e1 (3 × flux2-dev)$0.1260$0.1115−12%

The −12% hides a much larger structural error in both directions. The estimator charges OVERHEAD_PER_IMAGE = $0.03 per image. The real compile is one call per run, and it cost $0.0755 — 2.5× the per-image allowance on a 3-image run, and it would have been the same $0.0755 on a 300-image run.

Compile cost scales with context size, not with image count. docade's visual mission is ~8KB and goes into every compile. On a 300-image batch the estimator would over-charge overhead by ~$9; on a 1-image probe it under-charges by ~2.5×.

Recommendation: split the overhead term into a per-run compile cost (estimated from context size) and a per-image QA cost. Filed as ATL-14.

2. Model behaviour, per register

RegisterModelWhat it does wellWhere it failsVerdict
Collectible object (plush)flux2-dev, no referenceSilhouette legibility at 80px; consistent form across seeds; real weight and contact; fast (2–3s) and cheap ($0.012)Renders photoreal product photography, not flat-fill illustration. Baked painterly lighting. No transparency, and no reference input either.Not usable for this register. Not a cost problem — a steerability problem.
Collectible object (plush)recraft-v4-svgPrinciple 3 by construction — 26 paths, 7 flat solid fills, zero gradients. True alpha, no background to remove. A recolor is one attribute editSticker-flat. No weight, no volume, reads as a logo — the mission's other named non-goalRight answer for the wrong register. Keep it for glyph and avatar
Collectible object (plush)gemini-flash-image + 1 reference plateFlat separable fills and volume. Sits, has occlusion, reads cleanly at 80px. $0.067, ~8sFirst plate carried near-black fills; output inherited a keyline and a rim light despite the prompt forbidding both by nameThe register is reachable, and the fault was the plate
Collectible object (plush)gemini-flash-image + retinted plateKeyline gone, rim gone, flat fills held, weight held, reads at 80px. Principles 1–4 all passStill opaque (ATL-16). Small eye speculars, which are chromePass. ATL-20
A whole themed SET of sixgemini-flash-image + one plateSix different creatures that read as one artist: consistent simplification, consistent light direction, consistent shading discipline. All six nameable at 80pxUntested for rarity legibility (principle 5)Pass. ATL-19 — one plate holds a set
Small functional glyph—not yet probed
Scene—not yet probed
Character (Lyle)—not yet probed
Type—not yet probed
Thing on a saturated ground (prize)—not yet probed

The finding that matters: register is not steerable by adjective on flux2-dev. The compiled prompt did everything right — it demanded "flat separable fills", put shading in a "separate simple pass", said "clean vector-like flat-fill illustration", said "no rim light", and named the exact hex fills. The model produced soft-lit studio plush photography anyway, three times out of three.

This is docade's own visual mission §6 predicting its own failure a day in advance: "Not painterly. A beautifully rendered plush with baked lighting violates principle 3 and quietly destroys the growth model. This is the most likely wrong answer, because it is what image models do best by default."

It was right, and prompt engineering did not beat it.

3. What the prompt actually needs

WantedNaive prompt gaveWhat closed the gap
Flat fills, shading as a separate pass (principle 3)Baked painterly shading, fabric nap, physically-plausible soft lightA reference plate. Explicit instruction failed on flux2-dev; one flat vector plate handed to gemini-flash-image moved the register in a single pass. See §4.6.
Volume and weight with flat fillsThe vector plate gave fills without weight; the photoreal pass gave weight without fillsNaming both halves in the brief — what to KEEP from the plate (fill structure, shading discipline) and what to OVERRIDE (give it volume, it is not a sticker)
Reads at 80px (principle 1)Passes — one big dome, eight thick arms, no thin featuresThe prompt naming the 80px judgement and forbidding thin/antler-like features. Cheap and it worked.
Weight, sits on something (principle 4)Passes — arms flatten under load, broad underside occlusionDescribing how weight shows (arms flattening, occlusion) rather than asking for "weight".
Palette disciplinePasses — used only #4ac8ff, #f5f3ff, #262c4aGiving the compiler the committed palette and forbidding invention. No drift observed.
Transparent cutoutOpaque PNG on a dark groundImpossible on this model. See §4.1.

4. Process findings

These are worth more than the images.

4.1 The router carried a capability that does not exist

flux2-dev was flagged alpha: true. No FLUX.2 endpoint has a transparency parameter. Verified 2026-08-18 against each endpoint's own OpenAPI schema — fal-ai/flux-2, -pro, -flex and -max/edit all accept prompt, image_size, seed, guidance, steps, output_format and safety flags, and nothing resembling a background or alpha option. Neither gpt-image-2 endpoint on fal exposes one either.

This is the same bug class as the phantom slugs that ATL-1 caught, one level up: not a model id nobody verified, but a capability nobody verified. It was worse than the slugs, because a phantom slug 404s immediately and loudly, while a phantom capability silently produces 18 unusable plush assets that look fine until someone puts one on a different ground.

Fixed: every raster model is now alpha: false; only true-vector models claim transparency. npm run check asserts that invariant, and asserts that a transparent raster request fails loudly rather than resolving.

The adapter was also sending two parameters that do not exist — transparent_background and negative_prompt — which fal silently dropped. The negative one is the more insidious: a compiler that believes negatives are being honoured writes a weaker positive prompt.

4.2 art doctor reports auth, not spendability

fal's health probe returned green while the account balance was $0.00 and every generation call 403'd with User is locked. Reason: Exhausted balance. A key can be perfectly valid on an account that cannot spend.

https://rest.alpha.fal.ai/billing/user_balance returns the balance as a bare number for a read-only GET. Doctor should show it. Filed as ATL-15.

4.3 The estimator's cost model is shaped wrong

See §1. Overhead is charged per image; compile is per run and scales with context. Filed as ATL-14.

4.4 Two subjects with the same name

Compiled prompts were originally matched back to subjects by name. docade's piece 2 is octopus | coral colorway and octopus | teal colorway — the exact run that settles Q44 — and name-matching collapsed both to one prompt. Caught before spending, by writing the check first. Now matched by index, with a guard that refuses a batch whose distinct subjects compiled to identical prompts.

4.6 Reference conditioning works where prompting did not — ATL-17

The probe, run 2026-08-18 for $0.42 all in.

Prompting flux2-dev to produce flat separable fills failed three times out of three. So the probe asked a different question: give the model an example instead of an adjective.

  1. recraft-v4-svg produced a flat vector plate — 26 paths, 7 solid fills, zero gradients, true alpha. Principle 3 satisfied by construction. It also fails principle 4 completely: it is a sticker.
  2. That plate, rasterised to 512px, was handed to gemini-flash-image as a reference alongside a prompt that named explicitly what to keep (fill structure, shading as its own darker-value pass, no gradients) and what to override (give it the volume and weight of a stuffed object).
  3. Two candidates came back with flat separable fills and volume — the space between the two failure modes the mission predicted.

At 80px on #262c4a, the conditioned pair is the only one of the three that reads cleanly and has weight. Comparison sheet: runs/r-20260818-4648/proofs/atl-17-comparison.png.

This is the difference that matters: flux2-dev did not respond to explicit instruction at all, so there was nothing to iterate on. The conditioned output's remaining faults — a dark keyline, specular highlights, a cyan rim on candidate b — are each already named as non-goals in the visual mission, which means they are addressable in the next prompt or as a lesson. A model that ignores the brief is a dead end; a model that obeys it imperfectly is a starting point.

A capability gap this exposed. No FLUX.2 text-to-image endpoint accepts reference images — only fal-ai/flux-2-max/edit has image_urls. The router claimed refs: 10 for dev, pro and flex. Same audit also found the base fal-ai/flux-2 has no loras field: FLUX.2 LoRA inference is a separate endpoint, fal-ai/flux-2/lora, now its own model entry. And the trainer's parameters were the FLUX.1 trainer's — training_style defaults to "subject", so a plush style set would have silently trained a subject LoRA. All three would have surfaced only after spending a training run.

4.9 The plate wins arguments with the prompt — ATL-20

The first conditioned run produced a dark keyline and a rim light. The obvious reading was that the brief had failed to forbid them. It had not. The compiled prompt said, verbatim:

No black outline or keyline anywhere; separation from #262c4a comes from value and local colour alone. No rim light and no back-edge glow of any kind. No gradients, no airbrush, no soft falloff, no painterly brushwork.

The model produced a keyline and a rim anyway — because the plate had them. recraft-v4-svg had used two near-black fills (rgb(27,1,64), rgb(15,4,55)) to do contour and deep-shadow duty, and the model elaborated those into lighting.

The durable rule: when a reference plate and the prompt disagree, the plate wins. Negations lose to images especially badly. You cannot talk a model out of something its reference is showing it — you have to fix the reference.

So the plate was retinted (scripts/tint-plate.mjs): its fills were remapped onto a six-step ramp of a single hue from docade's committed palette, where the darkest step is a dark violet rather than near-black. Re-run on the corrected plate: keyline gone, rim gone, flat fills held, weight held, reads at 80px.

4.10 Q44 is answered, and the answer is yes — a recolor is free

The retint is a deterministic offline transform of a flat vector plate. It runs in milliseconds, calls no model, and costs $0.00. Two colorways of the same sculpt, produced from one source: runs/r-20260818-0266/proofs/atl-20-tint.png.

The shading survives the recolor because the shading is a darker value of the same hue rather than baked light — which is precisely what principle 3 asks for, now demonstrated rather than asserted. This is docade's Q44, and with it the economics of ~150 new images a year: a colorway is an attribute edit.

Note what it depends on: flat separable fills in a vector source. It would not work on any of the photoreal candidates, and that is the whole argument for principle 3 in one sentence.

4.11 One plate holds a whole set — ATL-19

The six real deepsea keys from docade's manifest — octopus, jellyfish, crab, turtle, otter, whale — conditioned on the same violet plate, one candidate each, $0.62 for the set.

The manifest's hard constraint on this category is "Must read at 80x80px — Critter Match uses these as card faces", and they are seen together on a grid, so the question was never whether any one is good. It was whether all six look like the same artist made them.

They do. Consistent level of simplification, consistent light direction, consistent shading discipline, consistent palette, no keylines or rims anywhere. All six are nameable at 80px, including the jellyfish — the mission's flagged hardest case for value-only separation, which works because its bell is simply a lighter value.

Proof: runs/r-20260818-33b3/proofs/at-80-on-262c4a.png.

What this settles for ATL-3. A reference plate was supposed to teach one look while only a LoRA could teach the look across many subjects. On this evidence the plate does both. The LoRA is now an optimisation rather than a prerequisite.

What it does not settle: one set, one seed each, one prompt. Rarity legibility (principle 5) is untested, and so is whether three different themed sets hold together with each other.

4.5 What the canon got right

The visual mission's §6 non-goals predicted the exact failure mode, in advance, in writing. Principles stated as falsifiable tests ("downscale to 80px, if you can't name the animal it fails") made the verdict a five-second observation rather than an argument. Keep writing them this way.

5. Corpus write-backs

FindingWritten toDone
No transparency parameter on any FLUX.2 endpoint; no negative promptdocs/corpus/models/flux2-dev.md + pro/flex/max-edit✅
Same, gpt-image-2 pair and ideogram-v4docs/corpus/models/gpt-image-2*.md, ideogram-v4.md✅
background_color ≠ alpha; do not infer from the SVG siblingdocs/corpus/models/recraft-v4.md✅
flux2-dev renders photoreal for "plush" and does not respond to flat-fill instructiondocs/corpus/models/flux2-dev.md✅
First real verdict history entry for any modeldocs/corpus/models/flux2-dev.md✅

6. Spend

Calibration ceiling$5.00 per run · $75.00 per month (raised to $200 on 2026-08-18)
Plush baseline (r-20260818-42e1)$0.1115
Vector plate (r-20260818-0266)$0.1572
Reference-conditioned probe (r-20260818-4648)$0.2517
Corrected-plate re-run, ATL-20 (r-20260818-f2bd)$0.2273
Six-subject set, ATL-19 (r-20260818-33b3)$0.6204
Plate retint to two colorways$0.0000 — offline transform
Diagnostic endpoint probes, outside the pipeline$0.2220 — see §4.7
Month to date$1.59 of $75.00

4.7 Diagnostic probes bypass the ledger

Confirming that a fal 403 was transient meant POSTing to five endpoints with curl. Those calls generated real images and cost $0.222, and none of it went through art, so none of it was ledgered until it was entered by hand.

The ledger is authoritative for spend (00-CHARTER.md), so an unrecorded call is a correctness bug, not an untidiness. Two candidate fixes: an art probe <endpoint> command that submits a minimal job and ledgers it, or simply never POSTing to a generating endpoint outside the CLI. The schema probes — which are free GETs against the OpenAPI document — are fine and should stay the default diagnostic.

4.8 The estimate gap is widening in the predicted direction

RunEstimatedActualGapCompile
r-20260818-42e1$0.1260$0.1115−12%$0.0755
r-20260818-0266$0.1100$0.1572+43%$0.0772
r-20260818-4648$0.1940$0.2517+30%$0.1177

Three runs, and the compile cost rose 56% as the context grew — the reference section and a larger corpus entry went in. Generation cost was predicted almost exactly every time. The error is entirely in the overhead term, exactly as ATL-14 says: it is modelled per-image when it is per-run and scales with context.

7. Ready to forge?

Not ready to forge docade-plush@1. Forging a style sheet now would encode a register the routed model cannot hit, and freeze it.

Open questions carried forward: