Ciyo · GPT-6

GPT-6 Astra Benchmarks Explained for Designers

A brass ruler, graduated amber glass weights, and a row of ceramic tiles of rising height sit on a warm ivory surface like a bar chart.
Editorial illustration created for this guide.

OpenAI’s GPT-6 Astra launch post publishes several dozen benchmark scores across computer use, professional work, coding, science, cybersecurity, alignment, and long context. Most were produced by OpenAI. A few come from an independent index. Almost none measure what a designer cares about most: whether the final picture is good.

This article reads that table as a creative team would, using the launch post and the Artificial Analysis methodology as checked on September 7, 2026. It explains what each relevant evaluation asks a model to do, which rows favour Astra and which favour Claude, and how to use the numbers without mistaking them for a verdict on your work.

How to read the table before reading any number

The launch post states that evaluation scores are the maximum at any effort. A 72.6% may therefore come from the most expensive setting, not the one you would use for a daily brief. Footnotes matter too: the Claude BenchCAD scores reflect three modifications described in Anthropic’s system card, and two Claude rows come from Mythos, the version of Fable with fewer safeguards.

Several rows are labelled internal, including Design Tasks, Data Science Tasks, and Database Migration Tasks. Their contents are not public, so they cannot be reproduced. Treat them as OpenAI’s description of its own priorities rather than as independent evidence. Everything below is vendor-reported unless it says otherwise.

Computer use: operating the software you already own

OSWorld 2.0 asks an agent to complete tasks inside real desktop software. OpenAI reports 72.6% for Astra, 65.7% for GPT-5.6 Sol, and 70.2% for Claude Opus 5, and adds that in latency simulations Astra finished in roughly 40 minutes per task versus roughly 75 minutes for Sol. ScreenSpot-Pro, which tests locating interface elements in professional applications without tools, is reported at 92.7% for Astra and 76.9% for Sol.

Agents’ Last Exam covers professional tasks in real software, including media production, and is reported at 59.3% for Astra, 55.5% for Opus 5, and 53.6% for Sol. For a designer these rows describe the ability to drive an editor, a browser, or a 3D tool through its interface. They do not describe the taste of what gets made there.

Professional rows with a creative angle

BenchCAD tests whether a model can reconstruct 3D objects from multi-view renders by writing CAD code. Astra is reported at 95.9% with tools, against 83.3% for Sol and 84.3% for Fable 5.1 under the footnoted conditions. OpenScore String Quartets measures reading music notation from images; Astra scores 0.84 on the 1 minus OMR-NED metric versus 0.19 for Sol, a large jump in structured visual reading.

The internal Design Tasks row shows 50.0% for Astra, 47.4% for Sol, and 35.8% for Claude Fable 5, with no Fable 5.1 figure. AutomationBench, which OpenAI groups under professional work, shows 41.4% for Astra and 31.4% for Fable 5.1. These are the closest published rows to design work, and the design row is also the one you cannot inspect.

Launch-post rows a designer can use, vendor-reported, September 7, 2026
EvaluationGPT-6 AstraGPT-5.6 SolBest Claude shown
OSWorld 2.072.6%65.7%70.2% (Opus 5)
ScreenSpot-Pro, no tools92.7%76.9%87.3% (Fable 5, Mythos)
Agents’ Last Exam59.3%53.6%55.5% (Opus 5)
BenchCAD95.9%83.3%84.3% (Fable 5.1, modified)
Internal Design Tasks50.0%47.4%35.8% (Fable 5)
MRCR v2, 512K to 1M96.3%73.8%Not shown
Paired bars show four vendor-reported rows for Astra and GPT-5.6 Sol: OSWorld 2.0 at 72.6 versus 65.7, internal Design Tasks at 50.0 versus 47.4, BenchCAD at 95.9 versus 83.3, and MRCR 512K to 1M at 96.3 versus 73.8.
Bars are drawn to the reported values. Every figure is vendor-reported at the maximum score across effort levels.

The independent index tells a different story

The Artificial Analysis Intelligence Index v4.1.1 appears in OpenAI’s own table: 61.2 for Astra, 60.9 for Sol, 65.7 for Claude Fable 5.1, 63.1 for Claude Opus 5, and 58.7 for Gemini 3.8 Flash. Artificial Analysis’s September 3 article describes Astra as matching Sol on intelligence while using about 10% fewer output tokens, at a price 2.5 times higher.

The index is a weighted average of ten evaluations across agents, coding, scientific reasoning, and general categories, with agents and general work weighted at 30% each. Its current version lists AA-Briefcase, GDPval-AA v2, Terminal-Bench, SciCode, Humanity’s Last Exam, and others. None is a visual or brand evaluation, so a lead on this index is a lead on reasoning breadth, not on design.

Long context: reading the whole brand book

OpenAI MRCR v2 asks a model to find eight needles in a long document. Astra is reported at 100.0% for 256K to 512K tokens and 96.3% for 512K to 1M, against 91.5% and 73.8% for Sol. For a creative team the practical reading is that a large reference pack, such as a brand book plus past campaigns, is more likely to be used accurately.

Retrieval quality is only half the story. Above 272,000 input tokens an Astra request is billed at the long-context band, which doubles the input rate. A high MRCR score makes a one-million-token request feasible; the pricing page decides whether it is sensible for a routine job.

Alignment rows that matter when you delegate

OpenAI reports an internal computer-use safety benchmark where lower is better: 2.4% for Astra, 22.0% for Sol, and 9.5% for Fable 5.1. It also reports an internal hallucination benchmark at 4.2% for Astra versus 12.2% for Sol, and says Astra is three times less likely than Sol to misrepresent its own capabilities.

If you plan to let a model operate design software or a browser on your behalf, these rows describe the risk of unintended actions. They are internal and vendor-reported, so they justify a careful trial with approvals switched on, not unattended use on a client account.

What the benchmarks do not cover

There is no published image-generation score, because Astra does not produce images directly; the renderer is GPT Image 2 and is evaluated separately. There is no typography, legibility, colour, or brand-consistency evaluation. There is no measure of how much editing a human needs after the model’s first draft.

Cost claims in the post are estimates under OpenAI’s conditions, such as a Terminal-Bench Science result at approximately 31% lower estimated API cost than Fable 5.1. Token-efficiency quotes, including a customer report of up to 20% fewer tokens on creative workflows, describe other people’s pipelines. Your own logs are the only cost benchmark that applies to your product.

Turn the table into a test plan

Use the rows to decide what to try, not what to buy. Computer-use results suggest testing a repetitive editing task in your actual software. BenchCAD and the Blender demos suggest testing a scripted 3D scene. MRCR suggests testing whether a full brand book improves compliance with copy rules. Each test should have a pass condition written before the run.

Run the same tasks on the model you use today. Keep the effort level, references, and renderer fixed, and record time, cost, and corrections. When the internal design row says 50.0%, the useful response is to measure your own approval rate on ten real briefs, which will be a number you can defend.

A short reading list for the numbers

Read the launch post for the full table and the footnotes about effort, modifications, and Mythos. Read the Artificial Analysis methodology to see which evaluations make up the index and how they are weighted. Read the model page for the specifications behind the long-context rows. Then read your own results, dated, beside them.

For image work the decision stays practical. Astra can plan and direct; GPT Image 2 renders; a reviewer approves. In Ciyo you can test the last two steps directly, with the price of each generation shown before you run it, and compare outputs from any planner you choose.

Continue reading

Questions about Astra’s benchmarks

Is GPT-6 Astra the best model on every benchmark?

No. In OpenAI’s own table Claude Fable 5.1 leads on Humanity’s Last Exam with tools and on the Artificial Analysis Intelligence Index, and Claude Opus 5 leads on the Coding Agent Index. Astra leads on the computer-use, CAD, science, and long-context rows shown.

Do any of these benchmarks measure image quality?

No. Astra returns text and calls an image tool for pictures. The renderer, GPT Image 2, is a separate model, and no published row scores typography, layout, or brand consistency.

What does “maximum at any effort” mean for my costs?

Scores may come from the highest reasoning setting, where hidden reasoning tokens are billed as output. Your daily setting may score lower and cost less, so test at the effort you will actually use.

Measure the step the benchmarks skip

Generate the same brief with GPT Image 2 in Ciyo, keep the outputs you would publish, and record your own approval rate beside the vendor’s table.