GPT-6 image generation · Concept review
GPT-6 Image Generation: Choose Between Two Concepts Using Your Brief

Two generated concepts can both look good and still fail the job. One may bury the product among props; the other may leave no room for the headline. Choosing by taste alone makes the decision hard to explain and easy to reverse.
In Ciyo, GPT Image 2.5 Flare renders the pictures and GPT-6 Astra can review them. OpenAI lists text as Astra's direct output, so Astra does the planning and reviewing while an image model creates the pixels. This walkthrough compares two real concepts for a fictional Fold notebook launch post, generated and reviewed on September 18, 2026. It shows a design decision, not an ad-performance test.
Write criteria you can see in the image
Before generating, turn the brief into four checks that a reviewer can answer by looking. For the Fold launch post we used: the ivory notebook is the clear hero; the mood is calm for desk workers; the upper third is plain for a headline; there are no extra branded items.
Avoid criteria such as "will convert" or "feels premium". They cannot be checked from a single image, and a model's opinion about them is not evidence.
Generate two genuinely different concepts
Use the Image Generator with GPT Image 2.5 Flare at 1:1. Change the idea, not just the colour: Concept A placed the closed notebook on a workday desk; Concept B used an overhead flat lay of the open notebook. Two near-identical images give the review nothing to decide.
Keep the fixed parts of both prompts identical, including the headline space and the no-text rule, so the comparison is fair.
Strong options make the choice harder, not easier. On September 15, 2026, Larus Canus (@MrLarus) shared four poster directions made with Images 2.5, each pairing a large colour field with a miniature scene: a spring landscape, a pottery kiln, plant printing and indigo dyeing. When every direction looks good, a fixed set of criteria turns "which one do I like" into "which one meets the brief".
Concept B for a square Instagram launch post for Fold, a fictional stationery shop. Overhead flat lay of the ivory linen notebook lying open to blank cream pages, with one pencil across the spread, on a pale grey paper background, soft even light, suggesting a fresh start. Keep the upper third plain for a headline added later. No text, no logos, no people.Post in Chinese showing four Images 2.5 poster concepts; the prompts and designs are the poster's. Our scoring workflow is an original adaptation.
Ask Astra to score, not to praise
Select both images on the canvas and ask GPT-6 Astra for a pass or fail on each criterion, one visible problem per concept and one recommendation. Tell it not to generate or edit anything. A fixed format makes the answer easy to compare with your own view.
In our run the agent panel showed an Inspect images step before the answer, so the scores referred to the actual pictures on the canvas.

Compare the verdict with the pictures
Astra failed Concept A on two criteria: the laptop, tea and props compete with the notebook, and foliage and the laptop edge intrude on the upper third. Concept B passed all four. Its named problem was that the open spread hides most of the ivory cover.
Look at the images yourself before you accept a verdict. Here the problems were visible: the dark laptop edge does pull the eye, and the flat lay does show pages rather than the cover. When the review names something you cannot see, ask where it is before acting on it.
| Criterion | Concept A: desk scene | Concept B: flat lay |
|---|---|---|
| Notebook is the clear hero | Fail | Pass |
| Calm desk-worker mood | Pass | Pass |
| Plain upper third | Fail | Pass |
| No extra branded items | Pass | Pass |

Carry the named weakness into the next round
A winning concept is rarely finished. Concept B's weakness is specific: the cover material is hard to see. That becomes the one instruction for the next version, for example a partly closed notebook that shows both pages and cover. Keep every passing criterion fixed while you fix it.
Record why the other concept lost. If the owner later asks for a lifestyle scene, the note explains which props caused the failure.
Keep the decision separate from performance
This review decides which concept follows the brief. It says nothing about which post will get more clicks. If you want performance evidence, publish variants to a real audience and measure the result; do not treat a model's preference as a test.
Use the same four criteria for every round. A stable checklist makes rounds comparable and stops the brief from drifting toward whatever the latest image happens to show.
Concept review questions
Does GPT-6 Astra create the images?
In this workflow GPT Image 2.5 Flare generated the images. Astra planned and reviewed them; OpenAI lists text as Astra's direct output modality.
Can I trust Astra's pass or fail scores?
Treat them as a second reviewer. Check each score against the image yourself, especially when the stakes are high.
Does the winning concept perform better in ads?
Unknown. The review measures fit with the brief. Only a live test with real viewers measures performance.
Decide with the brief in view
Generate two concepts in Ciyo and ask GPT-6 Astra to score them against your criteria.