GPT-6.1 Sol
builds a little world.
Light reasoning / Codex / Codex consumer-subscription run
Pilot result — comparative rating pending.Original generation telemetry is estimated or agent-reported. Import and inspection are separate.
01 / Standard evaluation viewImage tools
Export a presentation crop. Original evaluation evidence stays unchanged.
02 / Standard evaluation viewImage tools
Export a presentation crop. Original evaluation evidence stays unchanged.
Inspect the actual creation.
The published images show the supplied output. Download the original HTML source to explore it locally in an isolated browser. It relies on external libraries.
Screenshot preview · original artifact preserved unchangedLook closely. Judge carefully.
A detailed, cohesive voxel garden with a strong pagoda silhouette and working desktop controls. The opposite view holds together, while narrow-screen framing crops the garden. Pilot result — comparative rating pending.
Prompt adherence
Observed- Both standard daylight views show a five-tier pagoda with stepped roofs, upturned corners, pillars, railings, window detail and a roof ornament.
- The garden includes pink blossom and green trees, bamboo, flowers, stone lanterns, paths, a torii, fencing, a bridge and a pond. The island has visible soil sides and small surface elevation steps.
- Mouse orbit, zoom and keyboard vertical movement worked during actual browser checks.
Strong delivery of the central scene and unmistakable voxel identity. Terrain variation is relatively shallow; moss is not distinguishable as a separate feature in the standard evidence.
Visual craftsmanship
Observed- Crisp block geometry remains visible in architecture, foliage and terrain. Roof trims, railings and window divisions give the pagoda substantial small-scale detail.
- Green roofs, warm timber, pale pink blossoms and turquoise water form a consistent palette. Shadows separate overlapping roof and garden forms.
Careful architectural layering and a coherent palette make the scene feel finished. Tree crowns repeat similar clustered shapes, and the fine water detail is modest at the default viewing distance.
Composition and spatial coherence
Observed- The pagoda remains the tallest, central landmark in both front and opposite views. The front places the gate, bridge and pond around a visible route toward the building.
- The opposite view reveals bamboo and perimeter fencing; denser foreground foliage hides parts of the lower pagoda.
- The standard desktop frame slightly cuts off the island base. At a 390×844 viewport, substantial garden content falls outside the frame.
The front has a clear focal hierarchy and inviting garden depth. The rear is still coherent but more occluded; responsive camera framing is the clearest weakness.
Interaction
Observed- Orbit, wheel zoom, ArrowUp/ArrowDown movement, right-drag panning and Reset view visibly responded.
- Sunlit, Blue hour and Auto orbit toggles worked; the controls remained visible after a narrow viewport resize.
The desktop interaction set is complete and discoverable. Narrow-screen composition needs attention in a future generation, and touch usability remains untested.
Technical stability
Observed- The unchanged artifact rendered in a fresh Chrome context with reviewed CDN dependencies and no observed console errors, uncaught runtime errors or failed requests.
- Paused-camera frame pairs showed moving petals and koi. Canvas dimensions followed the tested viewport resize.
Stable in this bounded desktop session. This does not establish sustained performance, offline operation, or broad browser/device compatibility.
This assessment was not blinded: the judge had access to the model identity. AI judging involves subjective judgement and may carry bias.
- This is a subjective AI assessment of one supplied artifact, not a comparative score or evidence of general model superiority.
- The assessment context included the generating model, harness, receipt and source filename; it was not blinded.
- The exact judge model was not exposed in available session metadata.
- The frozen prompt, both standardized daylight images, source inspection and actual browser checks were used. Optional dusk imagery was not used to improve the visual rubric judgement.
- No paid judging API was used. A second independent judge, swapped-order pairwise review and comparison result are not yet available.
- Two static viewpoints and a short interaction session cannot certify every scene detail or every possible camera position.
In the browser.
macOS; fresh unauthenticated headless Chrome 154.0.8037.93; Playwright 1.63.0; 1600×1000 at device scale 1, plus 390×844 resize check; reviewed dependency allowlist.
Pre-execution source reviewpassed
Read the complete preserved HTML and exact prompt before execution. External requests are limited to reviewed Three.js 0.169.0, OrbitControls and Google Fonts; no authored navigation, form, storage or API requests were found.
Initial WebGL renderingpassed
The loading overlay disappeared and the complete pagoda/garden rendered in the fresh browser context. Both standard screenshots show the actual scene.
External-library loadingpassed
5 reviewed external resources loaded successfully, including both JavaScript modules and both font files. No failed or blocked requests occurred.
Console and runtime stabilitypassed
No console errors or uncaught page errors were observed during the 49.45-second final capture/check session.
Manual orbitpassed
A left-drag from (500,460) to (1000,460) produced the opposite view. 52.147% of tested central image pixels changed; both orientations were visually inspected.
Mouse-wheel zoompassed
A -360 wheel delta enlarged the rendered scene; 62.264% of central pixels changed. Reset restored the default framing.
Vertical camera movementpassed
Six ArrowUp key presses moved the view vertically, changing 53.15% of central pixels. Six ArrowDown presses returned near the original framing (0.483% residual, with live particles/fish).
Right-drag movementpassed
Right-drag from (650,500) to (750,540) visibly translated the garden in the frame (53.804% central pixel change).
Reset viewpassed
Reset returned from the rear camera angle to the front composition (0.455% residual difference from the earlier front capture, including live animation).
Auto orbit togglepassed
Auto orbit was initially active, was paused before evidence capture, and resumed visibly when enabled (43.612% central pixel change over 2.4 seconds).
Sunlit and Blue hour controlspassed
Blue hour changed the background, illumination and lanterns, with the selected button active. Returning to Sunlit restored daylight and its active button state.
Visible animationpassed
With auto orbit paused, frames 1.6 seconds apart showed falling petals and koi in different positions. Central changed-pixel fractions were 0.223% in daylight and 0.253% at dusk; animation continued during both modes.
Responsive canvas and controlspassed
Resizing from 1600×1000 to 390×844 produced a matching 390×844 canvas. All four buttons remained within the viewport, with no horizontal document overflow; returning to desktop restored the canvas size.
Narrow-screen scene framingfailed
At 390×844 the garden extends beyond both horizontal edges. The canvas responds, but the camera does not fit the entire diorama to the narrower viewport. The original output is preserved.
Individual firefly motionuntested
Dusk frame differences show continuing animation, but individual firefly trajectories were not separately resolved from petals and fish.
Long-session performance and extreme camera limitsuntested
No sustained frame-rate measurement, prolonged session, or exhaustive camera-limit sweep was performed.
Touch and other browser enginesuntested
Checks used desktop mouse/keyboard in Chrome. Mobile viewport resizing is not a real touch-device or Safari/Firefox test.
Offline and direct-file executionuntested
The unmodified HTML was served from a temporary local origin with internet access. Opening a downloaded file and loss-of-network behaviour were not tested.
Inspection: 2026-10-02 · Protocol: pagoda-v1-desktop-1600x1000
Passed records describe observed behaviour in one bounded browser session, not a guarantee across browsers or devices.
Pixel-change measurements support visible interactions; animation means reset frames are not pixel-identical.
Mobile framing clips the garden. Desktop default framing also slightly clips the island base.
The numbers, with their limits.
| Metric | Value | Provenance | Context |
|---|---|---|---|
| Elapsed generation time | 380seconds | estimated | Original generation duration, estimated by the generating agent. Import and capture time are separate. |
| Total input tokens | 160,000tokens | estimated | Estimated total input, including the separately reported cached input. |
| Cached input, included above | 132,000tokens | estimated | Included in the 160000 input tokens; do not add it again. |
| Output tokens | 7,800tokens | estimated | Estimated output; original receipt does not establish reasoning-token semantics. |
| Total tokens | 167,800tokens | calculated | Input plus output, calculated from estimates. |
| Tool calls | 14calls | agent-reported | Reported as measured by the generating agent; raw generation logs were not supplied and the count is not independently verified. |
| Separate cache writes | Unavailable | unavailable | No separate cache-write quantity supplied. |
| Separate reasoning tokens | Unavailable | unavailable | No separate reasoning-token quantity or inclusion semantics supplied. |
functions.exec 5exec_command 5cua_repl.js 4Illustrative API-equivalent task cost
USD · Standard pricing assumptionsA price illustration, not an invoice.
This was a consumer-subscription run. The estimate applies a named pricing snapshot to the supplied token estimates. It does not represent an actual charge.
- API-equivalent estimate; this consumer-subscription run has no per-run API invoice.
- Token counts are estimates reported in the original receipt.
- Cached input is included within total input and is charged once.
- No separately accounted tool charges; a Codex tool call does not imply an API tool fee.
- Standard-context rates assumed; no request-level lengths establish whether the 272,000-token threshold was crossed.
- Cache-write quantity is unknown. This illustration excludes separately accounted cache writes, rather than treating them as measured zero.
- standard API service tier; no regional processing multiplier assumed.
Rates per million tokens. Snapshot: 2026-10-02 · gpt-6.1-sol · Official source.
Exact model pricing sourcePricing conditions and historical receipt
- Snapshot independently checked against the exact official model pricing page on 2 October 2026.
- Standard API rates per million tokens, not a consumer subscription invoice.
- Input requests above 272000 tokens are charged at 2× input and cached-input rates and 1.5× output rates; this is per request, not total generation tokens.
- Fast service tier is 2×; Batch and Flex are 0.5×; regional processing is 1.1× where applicable.
- Cache writes are separately priced; the current receipt does not provide their quantity.
The original receipt reported $0.4400 using assumed comparable-model rates. It remains preserved as historical source evidence and is separate from this recalculation.
Historical receipt based on assumed comparable-model pricing. Retained as source evidence; these rates are not current verified pricing.
Read the original receiptExplore the data.
Each diamond is one imported run. Estimated cost and time describe this artifact, not general model performance. No frontier claim is made from one result.
Accessible data table · 1 run
| Configuration | Cost (USD) | Time (s) | Tokens | Tool calls | Rating | Provenance |
|---|---|---|---|---|---|---|
| GPT-6.1 Sol · Light / Codex | $0.1472 | 380 | 167,800 | 14 | Pending | Cost: estimate; time: estimated; tokens: calculated (see dependencies); tool calls: agent-reported; rating: pending |
Preserved, not polished.
The submitted artifact and frozen prompt are preserved byte for byte. These hashes identify the exact files behind this experiment.
- Run identity
sol-light-pagoda-001- Artifact SHA-256
a031fccbbe78ad6a079229e5049abbf00f630960464345dd2d76a703559e9bd9- Prompt SHA-256
63c96ce64ebecfc88bf59ba13e1be93154bcd56c15f7db8864c9642753f5ec4b- Generation date
- Unavailable
- Import date
- 2026-10-02
This pilot records one artifact from one configuration. It does not establish a general model ranking. Voxel Index currently refers only to performance on voxel-pagoda-v1.