INDEPENDENT RESEARCH / GPT-5.6 / AUGUST 16, 2026
What does more reasoning buy on a real web-design task?
I gave GPT-5.6 Luna, Terra and Sol the same sparse website-hero brief at every reasoning-effort level. One brief. One 1440×900 evaluation viewport, not shown to the models. No examples or access to my site.
This extends the familiar SVG drawing grid into an applied output: a self-contained HTML/CSS/JavaScript hero that has to combine product argument, information architecture, typography and interaction. The test covers the initial rendered viewport, not deployment readiness. Simon Willison: original 18-cell GPT-5.6 pelican matrix ↗
3 models × 6 effort settings. Exactly one provider-completed page per planned cell; all 18 are shown. No completed output was edited, replaced or excluded.
In these 18 single samples, effort changed time, token use and cost. The matrix shows visual variation, not a measured quality ranking.
- tested pages
- 18
- model tiers
- 3
- effort levels
- 6
- estimated total cost
- $4.74
All 18 first viewports
Rows hold reasoning effort constant; columns hold the model tier constant. Select any preview to open its original 1440×900 render.
NONE
Open 1440×900 ↗$0.0065 · 1m 40.9s · 5,381 output · 0 reasoning
loaded without overflow
Open 1440×900 ↗$0.0779 · 1m 59.6s · 6,476 output · 0 reasoning
loaded without overflow
Open 1440×900 ↗$0.2274 · 2m 19.6s · 7,566 output · 0 reasoning
loaded without overflow
LOW
Open 1440×900 ↗$0.0051 · 1m 17.4s · 4,194 output · 47 reasoning
loaded without overflow
Open 1440×900 ↗$0.0671 · 1m 42.5s · 5,574 output · 67 reasoning
loaded without overflow
Open 1440×900 ↗$0.2683 · 2m 43.3s · 8,929 output · 66 reasoning
loaded without overflow
MEDIUM
Open 1440×900 ↗$0.0098 · 2m 28.3s · 8,178 output · 164 reasoning
loaded without overflow
Open 1440×900 ↗$0.0968 · 2m 27.3s · 8,049 output · 86 reasoning
horizontal overflow: 190 px
Open 1440×900 ↗$0.3680 · 3m 42.2s · 12,253 output · 250 reasoning
loaded without overflow
HIGH
Open 1440×900 ↗$0.0193 · 4m 52.6s · 16,101 output · 748 reasoning
loaded without overflow
Open 1440×900 ↗$0.1579 · 3m 58.4s · 13,141 output · 250 reasoning
loaded without overflow
Open 1440×900 ↗$0.4208 · 4m 20.7s · 14,012 output · 280 reasoning
loaded without overflow
XHIGH
Open 1440×900 ↗$0.0197 · 4m 57.9s · 16,435 output · 1,552 reasoning
loaded without overflow
Open 1440×900 ↗$0.1683 · 4m 14.0s · 14,008 output · 483 reasoning
loaded without overflow
Open 1440×900 ↗$0.7644 · 7m 42.2s · 25,464 output · 5,832 reasoning
horizontal overflow: 180 px
MAX
Open 1440×900 ↗$0.0381 · 9m 33.3s · 31,727 output · 11,419 reasoning
loaded without overflow
Open 1440×900 ↗$0.5928 · 14m 57.3s · 49,388 output · 26,542 reasoning
loaded without overflow
Open 1440×900 ↗$1.4341 · 14m 28.8s · 47,787 output · 21,401 reasoning
loaded without overflow
Open the complete 4574×5830 comparison sheet ↗What this run actually shows
The displayed samples did not converge on one style.
Across these single samples, each effort setting produced a different visual thesis, hierarchy and amount of detail. This describes the displayed matrix; it is not evidence that higher effort cannot improve visual quality on average.
Max was the largest single effort tier.
The three max cells cost $2.06—43.5% of the estimated total—and averaged 13 minutes per request. The three none cells cost $0.31 and averaged 2 minutes: roughly 6.6× the cost and 6.5× the elapsed time. The two time-separated batches limit cross-model latency interpretation.
Luna was cheapest in this price snapshot.
All six Luna pages loaded in the defined initial-view check for a combined estimated cost of $0.099. Sol's six pages cost $3.48—about 35× more. This compares list-price token cost, not output quality, interaction quality or deployment readiness.
Observed horizontal overflow was non-monotonic.
In the defined initial-load check, all 18 pages loaded with zero runtime exceptions, console errors or external requests. Overflow above 1 px appeared only in Terra medium (190 px) and Sol xhigh (180 px). This does not characterize other implementation defects or interactions.
Method: deliberately little art direction
The creative brief names only the product goal, the audience effect and a light theme. A separate developer instruction constrains the response envelope, not the design.
Exact user prompt
Create a complete, self-contained website hero for a product that helps organizations safely operate autonomous AI agents in production.
The goal of the site is to make technical decision-makers interested in the product.
Use a light theme.Exact output envelope
Return only one complete self-contained HTML document beginning with <!DOCTYPE html>. Put all CSS and JavaScript inline. Do not use external assets, Markdown fences, or explanatory text.Execution and validation
- Denominator: 18 planned cells, 18 model requests, 18 completed HTML responses and 18 displayed pages. Model-request retries: 0. Completed outputs edited, repaired, replaced or excluded: 0.
- Requests ran in two concurrent batches, not one contemporaneous batch. Luna started at 2026-08-15 23:04:50 UTC, with six starts inside 73 ms. Terra/Sol started at 2026-08-16 08:49:15 UTC, with twelve starts inside 96 ms—9 h 44 m 24.643 s later. Cross-model latency and behavior may therefore reflect different service conditions or aliases.
- The models were not told a viewport size. After generation, each exact HTML response was served locally and evaluated at 1440×900 in Chromium 150.0.7871.186. Captures show only the initial viewport; no mobile result was generated or inserted.
- Validation observed navigation through Page.loadEventFired, document.fonts.ready and an additional 500 ms. It recorded Runtime exceptions, console.error calls and non-loopback/data/about:blank network requests; the load-event timeout was 30 seconds.
- Horizontal overflow equals max(documentElement.scrollWidth, body.scrollWidth) minus window.innerWidth. No clicks, form submissions, full-page behavior or deployment behavior were tested.
- Every raw response started with one <!DOCTYPE html>, contained exactly one HTML document, had no Markdown wrapper or external stylesheet, script or asset reference, and matched the rendered HTML byte for byte. Per-cell SHA-256 hashes and byte counts are in the JSON dataset.
- The models received no screenshot, CSS, brand asset, stranmor.com content, private prompt or skill context. They did receive the output envelope above, including the external-asset prohibition; that constraint limits the design space even though it does not prescribe an aesthetic.
- Token accounting: 1,584 input + 294,663 output = 296,247 total. The 69,187 reasoning tokens are reported inside output_tokens_details and are not added again.
Why run this instead of another SVG test?
Visual reasoning-effort comparisons already exist. Simon Willison published an 18-cell pelican-on-a-bicycle SVG grid for this exact 3×6 GPT-5.6 matrix; PlayCode later used a MacBook SVG to make geometry easier to grade. Those tests isolate drawing. This one asks how the same effort settings change a self-contained website hero under a sparse brief. Simon Willison: original 18-cell GPT-5.6 pelican matrix ↗
Search note, August 16, 2026: I found general GPT-5.6 benchmark tables and SVG grids, but no public 18-cell comparison in which Luna, Terra and Sol build the same self-contained website hero across all six effort settings. This is a bounded search observation, not a claim that no such page can exist.
Limits of the conclusion
- This is one sample per model-effort cell, so it cannot estimate average visual quality or statistical variance.
- Luna and Terra/Sol ran in two batches separated by 9 h 44 m. Cross-model latency and behavioral differences are not isolated from changing service conditions or model aliases.
- The task tests one narrow product category: a light-theme B2B hero for safe production AI agents. Results may differ for other products, languages, or visual genres.
- Elapsed time includes the API route and network path; it is not a pure measure of model compute.
- Cost is estimated from token usage and the OpenAI API list prices observed on August 16, 2026. Prices can change.
- Mechanical validation covers one initial desktop load and can observe loading, overflow, requests, landmarks, focus CSS and reduced-motion CSS. It does not test interactions, the full page, deployment readiness or whether a design is persuasive or beautiful.