Run one prompt across several models, efforts or token budgets at once and compare the sculptures side by side.

Benchmark

Run one prompt across several models, efforts or token budgets at once and compare the sculptures side by side.


Overview

Generative 3D art answers can it build this? This page answers the question you have immediately afterwards: what do I give up by turning the knob down?

It sends one prompt to several settings at once and puts the sculptures next to each other. Part counts, triangle counts and token totals rank the runs; looking at them tells you whether the extra tokens actually bought anything — which is not the same question, and is why every cell renders a real <model-viewer> rather than a row in a table.


Live demo


What the four axes mean

Model is the biggest lever, and not in the direction you would guess. See the measured sweep below.

Effort is thinking depth. It is only sent to models that accept output_config.effort; varying it across a model that rejects the parameter would run N identical calls and present them as a comparison, so the page refuses that combination rather than producing a confident-looking result.

Max tokens bounds thinking and response text together on Opus 5, so it is not a safety net — it is a quality dial with a cliff. Set it too low and the JSON is truncated mid-object and the variant fails outright. When that happens the panel says stop=max_tokens, because "the model did not return usable JSON" is a misleading way to describe running out of budget.

Prompt version is the newest axis and the only one that varies this site's own instructions rather than the model's settings. v1 is what every other page uses. v2 teaches the model version 2 of the scene manifest — defining a shape once and placing it several times — and adds one further instruction: that parts may and should interpenetrate where they join, because a lamp sits into the top of its tower rather than balancing on it.

v2 is not the default anywhere a sculpture is generated, and that is deliberate. Its entire claim is "it produces better sculptures", and only model runs can show that. Making it the default on an argument would be exactly the mistake this page exists to prevent, so v2 runs here and nowhere else until a sweep says it earns the promotion.

Every result panel carries the number that claim turns into: of the part pairs close enough to read as joined, how many actually blend — interpenetrate — rather than merely touching. It is computed from the manifest by geometry alone, with no rendering and no second call, so it costs nothing and cannot flatter the run. Pick the prompt axis, keep everything else fixed, and the difference in that percentage is attributable to the instructions and nothing else.

Only one axis varies per run. Two moving variables make a comparison unreadable, and a full grid is a combinatorial bill.


A measured sweep, and what it suggests

One prompt — "a desert observatory, sandstone and brass, dish pointed at the sky" — at effort=low, 4,000 tokens, all three models, run concurrently:

ModelTimePartsColoursTrianglesOut tokensCost
Haiku 4.515.7 s1984,1462,450$0.014
Sonnet 522.6 s19136,8422,447$0.042
Opus 530.3 s22125,8502,837$0.079

32 s wall clock against 69 s if run one at a time. $0.135 total.

Read that carefully, because it is not the shape people expect:

half the latency.

four colours reused deliberately*. Haiku used 8; Sonnet used 13; Opus used 12. None of them obeyed, but the cheapest model came closest — and palette discipline is the single thing that most decides whether the output reads as a sculpture rather than a test scene.

That does not mean "always use Haiku". It means the default on /generative-3d — Opus 5 at medium — is a choice worth re-examining against your own prompts, which is exactly what this page is for.

Everything above is a single run per cell. Model output varies between runs, and a three-colour difference is well inside that variance. Repeat a comparison before you act on it. The table is here to show what the page produces, not to settle the question.


Why colour count is the quality metric

Part count is the obvious number and it is nearly useless: a busier sculpture is not a better one, and 25 parts of noise scores higher than 12 parts of composition.

Palette size is better for this task specifically, because the prompt makes an explicit, checkable demand:

Pick a deliberate palette of three or four colours and reuse them. A different

colour per part looks like a test scene, not a sculpture.

So distinct(baseColorFactor) measures instruction adherence on the one instruction that most affects whether the result looks composed. It is computed from the built .glb rather than the model's own claims, so it cannot be gamed by a manifest that says one thing and builds another.

It is still a proxy. The viewers are there to be looked at.


Concurrency is the whole reason this is usable

A sculpt takes about 30 seconds. Four run serially is two minutes, and nobody runs a second sweep after waiting two minutes.

with concurrent.futures.ThreadPoolExecutor(max_workers=len(variants)) as pool:
    for index, result in pool.map(one, list(enumerate(variants))):
        results.append((index, result))

Bounded by the variant count, which MAX_VARIANTS already caps at 4. The status line reports both numbers — wall time and the serial equivalent — so the saving is visible rather than asserted.

Four, not six. The excalidraw benchmark this borrows from allows six. Here each cell carries its entire .glb inline as a data URL, a few hundred KB apiece, so the callback response grows with the matrix in a way a list of drawing elements does not.


Spending

This page can spend four times per click, on a public host, with no sign-in. tier: auth does not cover that: with no Clerk keys configured — the default, and what the free deployment runs — every tier except hidden degrades to public. That degradation is correct for reading documentation and useless as a spend gate, because it fails open.

So the ceiling lives in lib/spend.py and does not depend on identifying anyone:

LimitDefaultEnv var
Calls per rolling window40 / hourMODEL_MAX_CALLS_PER_WINDOW, MODEL_WINDOW_SECONDS
Cumulative estimated spend$5.00MODEL_MAX_SPEND_USD

Both are checked before the matrix runs, priced pessimistically — every variant costed as if it used its whole output budget. The remaining allowance is shown next to the button, so a refusal is never a surprise.

Two honest limitations:

are local estimates from a static price table, and ignore cache-write premiums and any discount. The real bill comes from Anthropic.

render.yaml runs on the free plan that is the whole deployment; add workers and the effective ceiling multiplies by the worker count. A shared limit needs Redis or a database, which is a real dependency for a docs site to carry.


Source

# File: docs/benchmark/benchmark.py

"""Run one prompt across several settings at once and compare the sculptures.

/generative-3d answers "can it build this?". This page answers the question you
have immediately afterwards: **what do I give up by turning the knob down?**
— which needs the same prompt built several ways, side by side, with the cost
of each.

FOUR THINGS THAT SHAPE THE DESIGN

1. Variants run CONCURRENTLY. They are network-bound, so a four-cell matrix
   costs roughly one sculpt's wall time instead of four. A sculpt is ~30s;
   run serially, a sweep takes two minutes and nobody runs a second one.

2. The cost is shown BEFORE you press the button, not after. An estimate that
   arrives with the results is an apology rather than a decision.

3. Every cell renders its own `<model-viewer>`. The numbers rank the runs;
   only looking at them tells you whether the extra tokens bought anything.

4. MAX_VARIANTS is 4, not 6 as on the excalidraw page this borrows from. Each
   cell carries its entire `.glb` inline as a data URL — a few hundred KB —
   so the callback response grows with the matrix in a way a list of drawing
   elements does not.
"""

from __future__ import annotations

import concurrent.futures
import time
import uuid

import dash_mantine_components as dmc
from dash import Input, Output, State, callback, dcc, html, no_update

import dash_model_viewer as dmv
from lib import model_picker, overlap, spend
from lib.sculptor import (
    DEFAULT_PROMPT_VERSION, MAX_TOKENS, PROMPT_VERSION_LABELS,
    PROMPT_VERSIONS, sculpt,
)

MAX_VARIANTS = 4

EFFORT_CHOICES = spend.EFFORTS
BUDGET_CHOICES = ["2000", "4000", "8000", "16000"]

VIEWER_ATTRS = {
    "environment-image": "neutral",
    "exposure": "1.1",
    "shadow-softness": "0.7",
}


def _panel(result, viewer_id: str):
    """One cell: its settings, its numbers, and its sculpture."""
    if not result.ok:
        return dmc.GridCol(
            dmc.Paper(
                withBorder=True, p="sm",
                children=dmc.Stack(gap=4, children=[
                    dmc.Badge(result_label(result), color="red", variant="light"),
                    dmc.Text(result.reason, size="xs", c="red"),
                ]),
            ),
            span={"base": 12, "md": 6},
        )

    return dmc.GridCol(
        dmc.Paper(
            withBorder=True, p="sm",
            children=dmc.Stack(gap="xs", children=[
                dmc.Group(gap=6, children=[
                    dmc.Badge(result_label(result), variant="light"),
                    dmc.Badge(f"{result.part_count} parts", color="violet", variant="light"),
                    # The palette count is the interesting one — see the .md.
                    dmc.Badge(f"{result.palette} colours", color="grape", variant="light"),
                    dmc.Badge(f"{result.seconds:.0f}s", color="gray", variant="light"),
                    dmc.Badge(f"~${result.usd:.3f}", color="teal", variant="light"),
                ]),
                # G7a's column. Placed with the badges rather than buried in
                # the small print because it is the number the prompt axis
                # exists to move.
                dmc.Text(blend(result), size="xs", c="indigo"),
                dmc.Text(
                    f"{result.output_tokens:,} out / {result.input_tokens:,} in"
                    f"  ·  {result.triangles:,} triangles"
                    f"  ·  {len(result.glb) / 1024:.0f} KB"
                    + (f"  ·  stop={result.stop_reason}"
                       if result.stop_reason not in ("end_turn", "") else "")
                    + (f"  ·  {len(result.notes)} clamped" if result.notes else ""),
                    size="xs", c="dimmed",
                ),
                dmv.ModelViewer(
                    id=viewer_id,
                    src=result.data_url,
                    alt=f"Generated sculpture: {result.manifest.get('name', 'untitled')}",
                    camera_controls=True,
                    shadow_intensity=1,
                    attributes=VIEWER_ATTRS,
                    style={"width": "100%", "height": "300px"},
                ),
            ]),
        ),
        span={"base": 12, "md": 6},
    )


def result_label(result) -> str:
    short = result.model.replace("claude-", "")
    return (f"{short} · {result.effort} · {result.max_tokens:,} tok · "
            f"prompt {result.prompt_version}")


def blend(result) -> str:
    """G7a's interpenetration rate for one cell, from geometry alone.

    THE NUMBER THE PROMPT AXIS EXISTS TO MOVE. The owner's words were that the
    output is "a lot of rigid shapes ... able to blend better together", and
    v2's joint guidance is the attempt at it. Of the part pairs close enough to
    read as joined, what fraction actually interpenetrate rather than merely
    touching? Computed from the manifest the model returned — no rendering, no
    second call — so it costs nothing to show and cannot flatter the run.
    """
    if not result.ok or not result.manifest:
        return ""
    try:
        report = overlap.report(result.manifest)
    except Exception:                                     # noqa: BLE001
        return ""
    if report["rate"] is None:
        return f"{report['parts']} parts, no joints"
    return (f"{report['overlapping']}/{report['joints']} joints blend "
            f"({report['rate']:.0%})")


component = dmc.Stack(
    gap="md",
    children=[
        dmc.Alert(
            title="This page spends real API credits",
            color="yellow",
            children=(
                "Each variant is a separate paid model call. The estimate below "
                "updates as you change the matrix — check it before running. "
                "This host also enforces a rate limit and a spend ceiling "
                "(lib/spend.py)."
            ),
        ),
        dmc.Paper(withBorder=True, p="md", children=dmc.Stack(gap="sm", children=[
            dmc.SegmentedControl(
                id="bm-axis",
                data=[
                    {"value": "effort", "label": "Vary effort"},
                    {"value": "budget", "label": "Vary max tokens"},
                    {"value": "model", "label": "Vary model"},
                    {"value": "prompt", "label": "Vary prompt version"},
                ],
                value="model",
                fullWidth=True,
            ),
            # Only one axis varies at a time. Two moving variables make a
            # comparison unreadable, and a full grid is a combinatorial bill.
            dmc.CheckboxGroup(
                id="bm-efforts",
                label="Effort levels to compare",
                value=["low", "high"],
                children=dmc.Group([dmc.Checkbox(label=e, value=e)
                                    for e in EFFORT_CHOICES], gap="md"),
            ),
            dmc.CheckboxGroup(
                id="bm-budgets",
                label="Max-token budgets to compare",
                value=["2000", "8000"],
                children=dmc.Group([dmc.Checkbox(label=f"{int(b):,}", value=b)
                                    for b in BUDGET_CHOICES], gap="md"),
            ),
            dmc.CheckboxGroup(
                id="bm-models",
                label="Models to compare",
                value=["claude-haiku-4-5", "claude-opus-5"],
                # Children filled on first render — see lib/model_picker.py.
                # Built at import, this froze the Anthropic-only list because
                # pages are imported before run.py warms OpenAI discovery.
                children=dmc.Group(id="bm-model-boxes", gap="md"),
            ),
            dmc.CheckboxGroup(
                id="bm-prompts",
                label="Prompt versions to compare",
                description=(
                    "v1 is what every other page uses. v2 teaches defs/ref/group "
                    "and tells the model parts may interpenetrate where they "
                    "join — it is measured here before it is believed anywhere."
                ),
                value=list(PROMPT_VERSIONS),
                children=dmc.Group([
                    dmc.Checkbox(label=PROMPT_VERSION_LABELS[v], value=v)
                    for v in PROMPT_VERSIONS
                ], gap="md"),
            ),
            dcc.Interval(id="bm-model-init", interval=150, max_intervals=1),
            dmc.Text(id="bm-model-status", size="xs", c="dimmed"),
            dmc.Grid(gutter="md", children=[
                dmc.GridCol(dmc.Select(
                    id="bm-fixed-model", label="Fixed model",
                    data=[], value="claude-opus-5",
                ), span={"base": 12, "sm": 4}),
                dmc.GridCol(dmc.Select(
                    id="bm-fixed-effort", label="Fixed effort",
                    data=[{"value": e, "label": e} for e in EFFORT_CHOICES],
                    value="low",
                ), span={"base": 12, "sm": 4}),
                dmc.GridCol(dmc.NumberInput(
                    id="bm-fixed-budget", label="Fixed max tokens",
                    value=MAX_TOKENS, min=1000, max=32000, step=1000,
                ), span={"base": 12, "sm": 4}),
                dmc.GridCol(dmc.Select(
                    id="bm-fixed-prompt", label="Fixed prompt version",
                    data=[{"value": v, "label": PROMPT_VERSION_LABELS[v]}
                          for v in PROMPT_VERSIONS],
                    value=DEFAULT_PROMPT_VERSION,
                ), span={"base": 12, "sm": 4}),
            ]),
            dmc.Textarea(
                id="bm-prompt",
                label="Prompt",
                description="The same prompt goes to every variant — that is what makes them comparable",
                value="a desert observatory, sandstone and brass, dish pointed at the sky",
                autosize=True, minRows=2,
            ),
            dmc.Group([
                dmc.Button("Run benchmark", id="bm-run"),
                dmc.Text(id="bm-estimate", size="sm", c="dimmed"),
            ], justify="space-between"),
        ])),
        dmc.Alert(id="bm-status", title="Status", color="gray", children="Ready."),
        dmc.Text(
            "Running every variant concurrently — about as long as one sculpt.",
            id="bm-working", size="sm", c="dimmed", display="none",
        ),
        dcc.Loading(html.Div(id="bm-results"), type="dot"),
    ],
)


def _variants(axis, efforts, budgets, models, f_model, f_effort, f_budget,
              prompts=None, f_prompt=DEFAULT_PROMPT_VERSION):
    """(model, effort, max_tokens, prompt_version) per cell. One axis moves."""
    f_budget = int(f_budget or MAX_TOKENS)
    f_prompt = f_prompt or DEFAULT_PROMPT_VERSION
    if axis == "effort":
        return [(f_model, e, f_budget, f_prompt)
                for e in (efforts or [])[:MAX_VARIANTS]]
    if axis == "budget":
        return [(f_model, f_effort, int(b), f_prompt)
                for b in (budgets or [])[:MAX_VARIANTS]]
    if axis == "prompt":
        # The reason this axis exists: the same model, the same budget, the
        # same words — only the system prompt differs, so a difference in the
        # result is attributable to the prompt and nothing else.
        return [(f_model, f_effort, f_budget, v)
                for v in (prompts or [])[:MAX_VARIANTS]]
    return [(m, f_effort, f_budget, f_prompt)
            for m in (models or [])[:MAX_VARIANTS]]


@callback(
    Output("bm-estimate", "children"),
    Input("bm-axis", "value"),
    Input("bm-efforts", "value"),
    Input("bm-budgets", "value"),
    Input("bm-models", "value"),
    Input("bm-fixed-model", "value"),
    Input("bm-fixed-effort", "value"),
    Input("bm-fixed-budget", "value"),
    Input("bm-prompts", "value"),
    Input("bm-fixed-prompt", "value"),
)
def _estimate(axis, efforts, budgets, models, f_model, f_effort, f_budget,
              prompts, f_prompt):
    """Price the matrix BEFORE it runs. See the module docstring."""
    variants = _variants(axis, efforts, budgets, models, f_model, f_effort,
                         f_budget, prompts, f_prompt)
    if not variants:
        return "Select at least one variant."
    total = sum(spend.estimate_usd(m, t) for m, _e, t, _v in variants)
    left = spend.remaining()
    return (f"{len(variants)} variant{'s' if len(variants) != 1 else ''} · "
            f"up to ~${total:.2f} if every one uses its full budget · "
            f"{left.calls_left} calls / ${left.usd_left:.2f} left on this host")


@callback(
    Output("bm-results", "children"),
    Output("bm-status", "children"),
    Output("bm-status", "color"),
    Input("bm-run", "n_clicks"),
    State("bm-axis", "value"),
    State("bm-efforts", "value"),
    State("bm-budgets", "value"),
    State("bm-models", "value"),
    State("bm-fixed-model", "value"),
    State("bm-fixed-effort", "value"),
    State("bm-fixed-budget", "value"),
    State("bm-prompt", "value"),
    State("bm-prompts", "value"),
    State("bm-fixed-prompt", "value"),
    running=[
        (Output("bm-run", "loading"), True, False),
        (Output("bm-run", "disabled"), True, False),
        (Output("bm-working", "display"), "block", "none"),
    ],
    prevent_initial_call=True,
)
def _run(_clicks, axis, efforts, budgets, models, f_model, f_effort, f_budget,
         prompt, prompts, f_prompt):
    if not (prompt or "").strip():
        return no_update, "Write a prompt first.", "yellow"

    variants = _variants(axis, efforts, budgets, models, f_model, f_effort,
                         f_budget, prompts, f_prompt)
    if not variants:
        return no_update, "Select at least one variant to compare.", "yellow"

    # Varying effort across a model that ignores the parameter would run N
    # identical calls and present them as a comparison — worse than an error,
    # because the output looks like a result.
    if axis == "effort" and f_model not in spend.EFFORT_CAPABLE:
        return (no_update,
                f"{f_model} does not accept the effort parameter, so every variant "
                "would be identical. Pick another model, or vary a different axis.",
                "red")
    # One gate for the whole matrix, priced pessimistically.
    estimate = sum(spend.estimate_usd(m, t) for m, _e, t, _v in variants)
    verdict = spend.check(len(variants), estimate)
    if not verdict.allowed:
        return no_update, verdict.reason, "red"

    run_id = uuid.uuid4().hex[:8]
    started = time.monotonic()

    def one(item):
        index, (model, effort, budget, version) = item
        began = time.monotonic()
        try:
            # The per-call budget gate is already satisfied for the matrix as a
            # whole; re-checking per variant would fail the tail of a run that
            # was approved, which is worse than approving it once.
            result = sculpt(prompt.strip(), model=model, effort=effort,
                            max_tokens=budget, enforce_budget=False,
                            prompt_version=version)
        except Exception as exc:  # one bad variant must not lose the others
            from lib.sculptor import SculptResult
            result = SculptResult(ok=False, reason=f"{type(exc).__name__}: {exc}",
                                  model=model, effort=effort, max_tokens=budget,
                                  prompt_version=version)
        if not result.seconds:
            result.seconds = time.monotonic() - began
        return index, result

    results = []
    with concurrent.futures.ThreadPoolExecutor(max_workers=len(variants)) as pool:
        for index, result in pool.map(one, list(enumerate(variants))):
            results.append((index, result))
    results = [r for _, r in sorted(results, key=lambda x: x[0])]

    elapsed = time.monotonic() - started
    ok = [r for r in results if r.ok]
    total_cost = sum(r.usd for r in results)
    serial = sum(r.seconds for r in results)

    status = (f"{len(ok)}/{len(results)} variants in {elapsed:.0f}s "
              f"(≈{serial:.0f}s if run one at a time) · ~${total_cost:.3f} total")
    if axis == "model" and any(m not in spend.EFFORT_CAPABLE for m, _e, _t in variants):
        status += "  ·  effort not sent to models that reject it"

    grid = dmc.Grid(gutter="md", children=[
        _panel(r, f"bm-view-{run_id}-{i}") for i, r in enumerate(results)
    ])
    return grid, status, "green" if len(ok) == len(results) else "yellow"


@callback(
    Output("bm-model-boxes", "children"),
    Output("bm-fixed-model", "data"),
    Output("bm-model-status", "children"),
    Output("bm-run", "disabled"),
    Input("bm-model-init", "n_intervals"),
)
def _fill_models(_n):
    """Fill both model controls when the page is VIEWED.

    Page modules are imported while Dash registers pages, which is BEFORE
    run.py calls openai_client.warm(). Reading the list at import froze the
    Anthropic-only set into this page and no OpenAI key could ever change it.
    """
    options = spend.model_options()
    boxes = [dmc.Checkbox(label=m["label"], value=m["value"]) for m in options]
    # Disabled with no provider key. This page spends up to four times per
    # click, so an enabled button on a keyless host is the worst of the three.
    return boxes, options, model_picker.status_line(), not spend.any_provider_available()

:defaultExpanded: false :withExpandedButton: true


This site runs without provider keys

The AI demos on this page are off on modelviewer.2plot.dev, deliberately. This is a documentation site; it carries no provider keys and does no production spend (owner's decision, 2026-09-12).

So on the public site you will see the model picker empty, the generate control disabled, and a line saying so. That is the expected state, not a fault — please do not file it.

To try it, run the site locally with your own keys:

git clone https://github.com/pip-install-python/dash-model-viewer
cd dash-model-viewer
printf 'ANTHROPIC_API_KEY=sk-ant-...\nCHATGPT_API_KEY=sk-...\n' > .env
pip install -r requirements.txt && python run.py

Either key alone is enough — the picker offers whichever provider it finds. lib/spend.py's ceiling then applies locally: a rolling call limit and a cumulative dollar estimate, so an accident costs a few cents rather than a weekend.

Everything on this page that does not need a key still works and is worth reading for it: the JSON schema, the clamps, the sRGB-to-linear conversion and the dependency-free glTF writer in lib/glb.py are all plain Python.


Source: /benchmark

Note for AI agents: This is the static, prerendered view of an interactive Dash application served because we detected a non-JS user agent. Full prose docs: