INDEPENDENT RESEARCH / GPT-5.6 / AUGUST 16, 2026

What does more reasoning buy on a real web-design task?

I gave GPT-5.6 Luna, Terra and Sol the same sparse website-hero brief at every reasoning-effort level. One brief. One 1440×900 evaluation viewport, not shown to the models. No examples or access to my site.

This extends the familiar SVG drawing grid into an applied output: a self-contained HTML/CSS/JavaScript hero that has to combine product argument, information architecture, typography and interaction. The test covers the initial rendered viewport, not deployment readiness. Simon Willison: original 18-cell GPT-5.6 pelican matrix

3 models × 6 effort settings. Exactly one provider-completed page per planned cell; all 18 are shown. No completed output was edited, replaced or excluded.

In these 18 single samples, effort changed time, token use and cost. The matrix shows visual variation, not a measured quality ranking.

tested pages
18
model tiers
3
effort levels
6
estimated total cost
$4.74

All 18 first viewports

Rows hold reasoning effort constant; columns hold the model tier constant. Select any preview to open its original 1440×900 render.

EFFORTLunaTerraSol

NONE

Website hero generated by gpt-5.6-luna at none reasoning effortOpen 1440×900 ↗
Lunanone

$0.0065 · 1m 40.9s · 5,381 output · 0 reasoning

loaded without overflow

Website hero generated by gpt-5.6-terra at none reasoning effortOpen 1440×900 ↗
Terranone

$0.0779 · 1m 59.6s · 6,476 output · 0 reasoning

loaded without overflow

Website hero generated by gpt-5.6-sol at none reasoning effortOpen 1440×900 ↗
Solnone

$0.2274 · 2m 19.6s · 7,566 output · 0 reasoning

loaded without overflow

LOW

Website hero generated by gpt-5.6-luna at low reasoning effortOpen 1440×900 ↗
Lunalow

$0.0051 · 1m 17.4s · 4,194 output · 47 reasoning

loaded without overflow

Website hero generated by gpt-5.6-terra at low reasoning effortOpen 1440×900 ↗
Terralow

$0.0671 · 1m 42.5s · 5,574 output · 67 reasoning

loaded without overflow

Website hero generated by gpt-5.6-sol at low reasoning effortOpen 1440×900 ↗
Sollow

$0.2683 · 2m 43.3s · 8,929 output · 66 reasoning

loaded without overflow

MEDIUM

Website hero generated by gpt-5.6-luna at medium reasoning effortOpen 1440×900 ↗
Lunamedium

$0.0098 · 2m 28.3s · 8,178 output · 164 reasoning

loaded without overflow

Website hero generated by gpt-5.6-terra at medium reasoning effortOpen 1440×900 ↗
Terramedium

$0.0968 · 2m 27.3s · 8,049 output · 86 reasoning

horizontal overflow: 190 px

Website hero generated by gpt-5.6-sol at medium reasoning effortOpen 1440×900 ↗
Solmedium

$0.3680 · 3m 42.2s · 12,253 output · 250 reasoning

loaded without overflow

HIGH

Website hero generated by gpt-5.6-luna at high reasoning effortOpen 1440×900 ↗
Lunahigh

$0.0193 · 4m 52.6s · 16,101 output · 748 reasoning

loaded without overflow

Website hero generated by gpt-5.6-terra at high reasoning effortOpen 1440×900 ↗
Terrahigh

$0.1579 · 3m 58.4s · 13,141 output · 250 reasoning

loaded without overflow

Website hero generated by gpt-5.6-sol at high reasoning effortOpen 1440×900 ↗
Solhigh

$0.4208 · 4m 20.7s · 14,012 output · 280 reasoning

loaded without overflow

XHIGH

Website hero generated by gpt-5.6-luna at xhigh reasoning effortOpen 1440×900 ↗
Lunaxhigh

$0.0197 · 4m 57.9s · 16,435 output · 1,552 reasoning

loaded without overflow

Website hero generated by gpt-5.6-terra at xhigh reasoning effortOpen 1440×900 ↗
Terraxhigh

$0.1683 · 4m 14.0s · 14,008 output · 483 reasoning

loaded without overflow

Website hero generated by gpt-5.6-sol at xhigh reasoning effortOpen 1440×900 ↗
Solxhigh

$0.7644 · 7m 42.2s · 25,464 output · 5,832 reasoning

horizontal overflow: 180 px

MAX

Website hero generated by gpt-5.6-luna at max reasoning effortOpen 1440×900 ↗
Lunamax

$0.0381 · 9m 33.3s · 31,727 output · 11,419 reasoning

loaded without overflow

Website hero generated by gpt-5.6-terra at max reasoning effortOpen 1440×900 ↗
Terramax

$0.5928 · 14m 57.3s · 49,388 output · 26,542 reasoning

loaded without overflow

Website hero generated by gpt-5.6-sol at max reasoning effortOpen 1440×900 ↗
Solmax

$1.4341 · 14m 28.8s · 47,787 output · 21,401 reasoning

loaded without overflow

18 GPT-5.6 website heroes arranged by Luna, Terra, Sol and six reasoning-effort levelsOpen the complete 4574×5830 comparison sheet

What this run actually shows

01

The displayed samples did not converge on one style.

Across these single samples, each effort setting produced a different visual thesis, hierarchy and amount of detail. This describes the displayed matrix; it is not evidence that higher effort cannot improve visual quality on average.

02

Max was the largest single effort tier.

The three max cells cost $2.06—43.5% of the estimated total—and averaged 13 minutes per request. The three none cells cost $0.31 and averaged 2 minutes: roughly 6.6× the cost and 6.5× the elapsed time. The two time-separated batches limit cross-model latency interpretation.

03

Luna was cheapest in this price snapshot.

All six Luna pages loaded in the defined initial-view check for a combined estimated cost of $0.099. Sol's six pages cost $3.48—about 35× more. This compares list-price token cost, not output quality, interaction quality or deployment readiness.

04

Observed horizontal overflow was non-monotonic.

In the defined initial-load check, all 18 pages loaded with zero runtime exceptions, console errors or external requests. Overflow above 1 px appeared only in Terra medium (190 px) and Sol xhigh (180 px). This does not characterize other implementation defects or interactions.

Method: deliberately little art direction

The creative brief names only the product goal, the audience effect and a light theme. A separate developer instruction constrains the response envelope, not the design.

Exact user prompt

Create a complete, self-contained website hero for a product that helps organizations safely operate autonomous AI agents in production.

The goal of the site is to make technical decision-makers interested in the product.

Use a light theme.

Exact output envelope

Return only one complete self-contained HTML document beginning with <!DOCTYPE html>. Put all CSS and JavaScript inline. Do not use external assets, Markdown fences, or explanatory text.

Execution and validation

  1. Denominator: 18 planned cells, 18 model requests, 18 completed HTML responses and 18 displayed pages. Model-request retries: 0. Completed outputs edited, repaired, replaced or excluded: 0.
  2. Requests ran in two concurrent batches, not one contemporaneous batch. Luna started at 2026-08-15 23:04:50 UTC, with six starts inside 73 ms. Terra/Sol started at 2026-08-16 08:49:15 UTC, with twelve starts inside 96 ms—9 h 44 m 24.643 s later. Cross-model latency and behavior may therefore reflect different service conditions or aliases.
  3. The models were not told a viewport size. After generation, each exact HTML response was served locally and evaluated at 1440×900 in Chromium 150.0.7871.186. Captures show only the initial viewport; no mobile result was generated or inserted.
  4. Validation observed navigation through Page.loadEventFired, document.fonts.ready and an additional 500 ms. It recorded Runtime exceptions, console.error calls and non-loopback/data/about:blank network requests; the load-event timeout was 30 seconds.
  5. Horizontal overflow equals max(documentElement.scrollWidth, body.scrollWidth) minus window.innerWidth. No clicks, form submissions, full-page behavior or deployment behavior were tested.
  6. Every raw response started with one <!DOCTYPE html>, contained exactly one HTML document, had no Markdown wrapper or external stylesheet, script or asset reference, and matched the rendered HTML byte for byte. Per-cell SHA-256 hashes and byte counts are in the JSON dataset.
  7. The models received no screenshot, CSS, brand asset, stranmor.com content, private prompt or skill context. They did receive the output envelope above, including the external-asset prohibition; that constraint limits the design space even though it does not prescribe an aesthetic.
  8. Token accounting: 1,584 input + 294,663 output = 296,247 total. The 69,187 reasoning tokens are reported inside output_tokens_details and are not added again.

Why run this instead of another SVG test?

Visual reasoning-effort comparisons already exist. Simon Willison published an 18-cell pelican-on-a-bicycle SVG grid for this exact 3×6 GPT-5.6 matrix; PlayCode later used a MacBook SVG to make geometry easier to grade. Those tests isolate drawing. This one asks how the same effort settings change a self-contained website hero under a sparse brief. Simon Willison: original 18-cell GPT-5.6 pelican matrix

Search note, August 16, 2026: I found general GPT-5.6 benchmark tables and SVG grids, but no public 18-cell comparison in which Luna, Terra and Sol build the same self-contained website hero across all six effort settings. This is a bounded search observation, not a claim that no such page can exist.

Limits of the conclusion

  • This is one sample per model-effort cell, so it cannot estimate average visual quality or statistical variance.
  • Luna and Terra/Sol ran in two batches separated by 9 h 44 m. Cross-model latency and behavioral differences are not isolated from changing service conditions or model aliases.
  • The task tests one narrow product category: a light-theme B2B hero for safe production AI agents. Results may differ for other products, languages, or visual genres.
  • Elapsed time includes the API route and network path; it is not a pure measure of model compute.
  • Cost is estimated from token usage and the OpenAI API list prices observed on August 16, 2026. Prices can change.
  • Mechanical validation covers one initial desktop load and can observe loading, overflow, requests, landmarks, focus CSS and reduced-motion CSS. It does not test interactions, the full page, deployment readiness or whether a design is persuasive or beautiful.

Sources and data

Download the sanitized benchmark data (JSON)The public dataset includes prompt, two-batch timing, token accounting, pricing, browser protocol, per-cell HTML hashes, conformance checks and validation. Private gateway identifiers and response IDs are intentionally omitted.