Introducing the LuminaBench pilot: one preserved Voxel Pagoda creation, its evidence and the limits of what we can infer.
By LuminaBench
The unchanged Voxel Pagoda output from GPT-6.1 Sol · Light · Codex.
LuminaBench begins with one creative-coding experiment: the Voxel Pagoda prompt. The submitted result is preserved unchanged, alongside the exact version of the prompt that produced it.
The artifact comes first.
The supplied configuration is GPT-6.1 Sol with the Light setting in Codex. Our pages distinguish the finished creation, independent browser inspection, subjective AI assessment and original generation telemetry.
The two standard views and the browser-check record give the assessment something concrete to work from. A good-looking still and a functioning interactive scene are related questions, but each needs its own evidence.
Numbers that keep their context.
The original receipt required numerical estimates. Its token counts and elapsed time therefore remain labelled as estimates, and its tool count remains agent-reported. API-equivalent cost is an illustration at a dated price snapshot, not an invoice for the subscription run.
API-equivalent cost$0.1472Estimate · see assumptions
Generation time6m 20sestimated
Total tokens167,800Calculated from estimates
Tool calls14Agent-reported
The recalculated estimate is $0.1472 under the published standard-pricing assumptions. Unknown cache-write quantities and absent provider usage logs remain limitations; the arithmetic does not make the underlying estimates measured facts.
A pilot, with an open question.
One creation cannot establish a comparative ranking. Voxel Index currently describes this Pagoda v1 experiment; comparative ratings will require eligible comparisons using the same prompt and capture protocol.
Pilot result — comparative rating pending.
The useful question today is specific: what does this configuration create from this prompt, and how well can we document it? The experiment record is our first answer.