Models · System One
Judging Four Variants, and the Gap a Decision Model Fills

Generating options is the easy half. Choosing between them is the half that eats an afternoon, and it is the half most creative tooling leaves to you.
So we ran the review pass properly: four square promo posts for a fictional harbour coffee bar, four written criteria, and a model asked to score every post on every criterion. The result was useful, and it also showed exactly the weakness that a typed decision model like Jev is designed to remove. Ciyo cannot run Jev — it produces no image and no video and sits in neither of our model registries — so this article shows what works today and marks the seam where something else would go.
Four variants, one offer
The brief was a weekend drink promotion for Pier Six Coffee, an invented harbour-side bar: the Salted Maple Cold Brew at 4.50, Saturday and Sunday only, aimed at people walking the harbour. We ran it in Ciyo on September 21, 2026. One post came back from the original brief; three more were generated as deliberate alternatives — a close-up on the glass, a cup held at the harbour rail, and a flat-lay on the counter.
Four candidates is the realistic number. It is enough that you cannot hold them all in your head, and few enough that a person will actually look at each one.

Write the criteria before you look
The discipline that makes a review pass worth running is writing down what you are judging before you see the candidates. Otherwise you rationalise the one you already liked.
Four criteria, each a question with an answer: does the drink catch the eye first; is the price readable at thumbnail size; is the reserved headline space actually clear; does it read as a weekend harbour scene. Every one of those is checkable by someone who has never seen the brand.
Read the four posts on the canvas. Score each one from 1 to 5 on four separate criteria, asking each one on its own: (1) is the drink the first thing the eye lands on, (2) is the 4.50 price readable at thumbnail size, (3) is the top headline space actually clear, (4) does it read as a weekend harbour scene. Give me a table with one row per post and one column per criterion, then name the one you would post and the single reason why.What came back
It worked. The model read all four at thumbnail scale, produced a four-by-four table with a short justification in every cell, and named a winner: the cup held at the harbour rail, “because it most directly mirrors the audience's weekend harbour walk.” That is a defensible answer, arrived at against criteria we wrote down first.
It also caught the one real failure. The counter flat-lay scored 1 on “reads as a weekend harbour scene”, with the note that there are no visible harbour cues and the weekend is conveyed only by the copy. That is exactly the kind of thing a person misses at four in the afternoon on the fourth variant.

The number that gives the weakness away
Sixteen cells. Fifteen of them came back as a 4 or a 5. The single discriminating score in the whole table was that 1.
So the table is a good failure detector and a poor ranker. It will reliably tell you which variant is broken. It will not reliably tell you which of the three surviving variants is best, because on a one-to-five scale it does not use the middle.
This is the well-known behaviour of a prose model asked to score: a written scale gets compressed towards the top, and the reasons in each cell are more informative than the digits next to them. It is not a flaw in this model or this prompt. It is what happens when a judgement has to be expressed as a sentence containing a number.
| Question | Answered well | Why |
|---|---|---|
| Is any variant broken? | Yes | The flat-lay scored 1 on harbour scene, with a specific reason |
| Which of the good ones is best? | No | Fifteen of sixteen cells were a 4 or a 5; the scale did not separate them |
| Would it score the same tomorrow? | Unknown | We ran it once; repeatability is exactly what a prose scorer is weakest at |
| Should a human still look? | Yes | Nothing here decides a brand's tone, and the winner was chosen on one reason |
Where a System One model would sit
This is the seam. A typed decision model does not write a sentence containing a score; it returns a value of a type you declared, with probabilities and a confidence number, and it returns the same value when asked again.
For this exact task that would mean four Score questions with an explicit rubric — not “1 to 5” but four named levels with written definitions — evaluated against each post, plus a Choice for the overall call with an explicit `needs_a_human` option. The combining rule would live in your own code, where you could weight readability above atmosphere for a paid placement and the other way round for an organic post.
What you would gain is not better taste. It is a score that means the same thing on every run, a confidence value to decide when a person should look, and a rule you can read and change without rewriting a prompt.
What you would lose is the reasons. The cell that told us the flat-lay has no harbour cues is prose, and a decision model does not produce prose. In practice the interesting design is both: a typed score for the ranking and the routing, and a language model for the sentence that explains it.
What Jev cannot do here, plainly
It cannot look at the posts. TypeSafe's own documentation states that state must be a string, a JSON object or an array of text values, and that images, audio and video are not supported yet. A visual review pass is outside what the model accepts today.
It cannot make the variants either. It generates no image and no video, which is why it appears in neither of Ciyo's model registries and why Ciyo cannot offer it.
So the honest version of this article's idea is a text-shaped one: a decision model can judge the brief, the caption, the alt text, the campaign metadata and the structured description of a render — the parts of creative work that are already words. The picture still needs eyes, and today those are yours or a vision model's.
Run the review pass anyway. Writing the criteria down before you look is most of the value, and it works with whatever model you have.
Review-pass questions
Can Jev score my ad variants?
Not the pictures. Its documented state is text only — a string, a JSON object or an array of text values — so images are outside what it accepts today. It could score the brief, the caption or a structured description.
Why did almost every score come back a 4 or a 5?
Because a prose model expressing a judgement as a written number compresses the scale towards the top. The written reasons in each cell carried more information than the digits.
Is the review pass worth running anyway?
Yes. It caught a real failure the same afternoon it was made, and writing the criteria down before looking is most of the benefit regardless of which model scores them.
Does Ciyo run Jev?
No. It produces no image and no video and is in neither of Ciyo's model registries. The scoring in this article was done by GPT-6 Astra, which Ciyo does run.
Run your own review pass
Generate four variants, write four criteria before you look at them, and make the agent score every variant on every one.