Promptography™Scoring method · publishedStart your shoot

The Quality Gate, audited

The judge has a name.

“Independent AI judge” is not a claim you should have to take on faith. This page names the scorer models, publishes the threshold, and walks through one real pilot order — scores, rejections, and the frame a human vetoed even though its score passed.

The scorers

Two named judges, one human veto.

Gemini VLM criticEvery delivered frame is scored by Google’s gemini-3.1-pro-preview vision model against your own uploaded references. It returns an identity score from 0 to 100 plus a defect list (warped anatomy, invented text, plastic skin, wrong gear). The verdict rule is mechanical: ship only if the score is at least 85% and no single defect disqualifies the frame.
ArcFace identity scoringIn our pilot pipeline every candidate is additionally scored with ArcFace face-embedding similarity (InsightFace buffalo_l): the generated face is compared against each reference photo as a cosine similarity from 0 to 1, with pass bars calibrated per subject from their own photos — not one global number.
Human review, finalScores miss things. In the pilot, frames with passing scores were still vetoed by eye — added muscle mass, a strained smile, a harsh shadow. Human review outranks the score, always in the reject direction: a human can kill a frame, never rescue a below-threshold one.

The threshold

85%, same at every price.

A frame that scores below 85% is regenerated at our cost (up to two extra attempts per slot). If the best attempt still cannot clear the gate, we deliver the strongest take, record it honestly as best-effort in the database, and the 7-day, 100% refund stands. If the scorer itself is unavailable we deliver rather than hold your paid order hostage, and the frame is recorded as unscored. The threshold is identical for a $9.99 First Look and the largest Portfolio — likeness is never a paywall.

Reading the numbers

Three scales. They never compare.

Quality Gate score · 0–100 · bar: 85The shipping gate. The Gemini judge grades every frame against your references before delivery; below 85 it is regenerated at our cost, and the best take ships with its score on record. Wherever this site says “85% Quality Gate,” it means this scale.
ArcFace identity match · similarity · calibrated barThe pilot audit record — 91.4% ArcFace on the barrel frame, mean 82.7% across the 7 delivered frames. It measures how closely the generated face matches your own reference photos, with the pass bar calibrated per person from how much their own photos agree with each other — Dimi’s STRONG bar is 54.8%. His 82.7% pilot mean clears that bar by 28 points; it is not a Quality Gate score, and comparing it to the 85 threshold is comparing two different rulers. The rest of the pilot’s published set was retired in a founder audit — the eye outranks the score, even above the bar.
Dual-model identity score · normalized · bar: 0.85The chip on published founder frames, on Danielle’s client shoot, and on the pilot barrel — the newer gate, two models instead of one. ArcFace and AdaFace each compare the frame with the subject’s own real reference photos, and the chip publishes the LOWER of the two reads, scaled so 0 is the average score of a different person and 1.00 is the typical agreement between two real photos of the same person. The pass bar is 0.85; frames where one reference dominates must additionally pass on the second reference alone, and a raw similarity of .90 or higher is treated as a reference-carryover alarm, not a pass. Every published chip is floored to a re-score of the exact shipped file, so no chip overstates the file it sits on. Not an ArcFace percentage, not a Quality Gate score — a third ruler.

Worked example · real pilot order

Dimi’s red-carpet frame, scored in the open.

From the pilot shoot of Dimi S. (Amsterdam, 2026-08-30, published with his consent). Pilot scores below are ArcFace cosine similarity against his own reference photos — his calibrated STRONG bar is 54.8%, derived from how similar his reference photos are to each other. The pilot delivered this brief at 94.5%. The red-carpet frame was later retired with most of the pilot set in a founder audit; the numbers below stand as the published record of how every frame is scored.

Red-carpet brief: three base candidates generated, best-of-3 selection, then crop.
CandidateArcFace score (max vs. references)Outcome
Base 194.8%Picked — delivered at 94.5% after the 4:5 crop · verdict STRONG
Base 292.7%Rejected — four foreground faces instead of three
Base 393.5%Rejected — four foreground faces instead of three

The honest counter-example from the same order: on the LinkedIn brief, base 1 scored higher (75.6%) than the frame we delivered (74.2%, 75.8% after crop) — and was still rejected by human review for a harsh shadow and a strained smile. The score never outranks the eye. Across the whole order: 21 candidates generated, 7 frames delivered, all 7 STRONG, mean ArcFace identity match 82.7%.

The pilot’s one remaining frame on the landing page — the barrel — carries its dual-model chip (identity 1.57 · PASSED); its audited ArcFace record from the pilot delivery is 91.4%. The rubric itself is on the Quality Gate page, and failed founder frames are in the reject gallery. The bar is enforced, not decorated: one audit round pulled three pilot frames from the landing page — a sub-bar rescore (89.4% against the shoot’s 90% rule) and two failed 3× zoom audits despite passing chips — and a follow-up founder audit retired the rest of the set outright. Every retirement is disclosed, never painted over. A published bar you don’t enforce is marketing, not a bar — and the score never outranks the eye.

Calibration case · September 2026

Four passing scores. Four rejections — by the founder.

The clearest calibration lesson in this method came from our own founders. Four passport-style frames shipped with audited dual-model chips of 1.10, 0.96, 1.20 and 1.26 — every one a numeric pass against the 0.85 bar, floored re-scores of the exact published files. The founder rejected all four the day they went live — his own two as the subject himself: the 0.96 frame, his ruling went, “does not look like me at all” — the last few percent is exactly what matters. The frames came off every selling surface the same day and are published, scores and verdicts included, in the reject gallery. The standing lesson is now part of the method: a facial-recognition pass is necessary, never sufficient — perceived likeness is judged by a human, the subject holds the final veto, and no founder frame publishes without the subject’s explicit approval. Your shoot inherits the same two gates: the numeric bar first, and you — backed by the refund — last.

Scored before you see it.

Same judges, same 85% threshold, at every tier. If you still disagree with the gate, the full refund stands — even after download.Start your shoot