Machine-readable page content
Canonical: https://semiotic.nteract.io/examples/the-benchmark-is-a-chart-too
The Benchmark Is a Chart, Too
OpenAI GPT-5.6 compatibility evidence · snapshot 2026-07-27
The benchmark is a chart, too.
A scorecard has encodings, denominators, and failure modes—just like the charts it judges. This page preserves the complete baseline, then shows what happened when revised grounding and generation contracts were tested in three later trials.
COMPLETE1,119recorded requests
- Baseline evidence
- 516full suite, one response per case
- Follow-up evidence
- 6033 independently submitted targeted trials
- Repeated outcomes
- 540 + 63grounding answers + generated proposals
- Estimated API cost
- $4.24baseline plus follow-up
Later repeated trials
The revised contracts changed the result.
The follow-up repeated every previously failing generation fixture and every answerable grounding question across Sol, Terra, and Luna. PNG-only remained unchanged as the reading control.The revised source-fact payload recovered every targeted answerable lookup across Sol, Terra, and Luna. This follow-up did not repeat the unanswerable questions.
First-attempt generation6 of 7 repaired fixtures held across every model and trial.The remaining fixture, gauge-static, passed 7/9. Sol and Terra passed all three trials; Luna passed once and twice added the unsupported chart-HOC prop accessibleTable to an otherwise valid BigNumber proposal.
- Passing proposals
- 21/21
- Pass rate
- 100%
- Passing proposals
- 21/21
- Pass rate
- 100%
- Passing proposals
- 19/21
- Pass rate
- 90%
Original baseline · reader grounding
The original payload improved restraint, not chart reading.
Compare correct answers with correct restraint instead of compressing both into the same bar.All questionsCombined evidence moved Sol from 43 to 45 correct answers. Terra and Luna tied their PNG-only totals.Twenty answerable questions plus thirty questions the evidence cannot answer.
The combined reader-grounding payload improved Sol's aggregate score through abstention, tied the PNG-only totals for Terra and Luna, and did not improve answerable-question accuracy.
+2overall-1answered+3abstained
0overall-1answered+1abstained
0overall0answered0abstained
Original baseline · first-attempt generation
Plausible chart choices still failed their contracts.
Every proposal had one chance to validate, render visible marks, and avoid error diagnostics. There was no repair pass.A pass requires valid props, visible render evidence, and no error diagnostics.
- 172 requests
- $1.42
- Average response
- 1,854 ms
- 172 requests
- $0.64
- Average response
- 1,308 ms
- 172 requests
- $0.22
- Average response
- 1,208 ms
Failures by fixture family| Fixture | What it taught us | Sol | Terra | Luna | Later trials |
|---|
Focal SLA valuegauge-static | surface seamSol chose the real BigNumber component, but the render oracle could not prove it. Luna chose GaugeChart and rendered no marks. | × | · | × | 7/9 |
|---|
Line, then pushline-push | push contractBoth proposals reached the right chart family but missed a valid push-ready contract. | · | × | × | 9/9 |
|---|
Scatter, then pushscatter-push | formatter behaviorA serial string formatter reached a callback-only render path. | · | × | · | 9/9 |
|---|
Bubble comparisonbubble-static | formatter behaviorThe same formatter ambiguity failed a static XY proposal. | · | × | · | 9/9 |
|---|
Incident mapsymbol-map-static | geo contractThe visible marks rendered, but the submitted geo props did not satisfy validation. | · | × | × | 9/9 |
|---|
Observed values on a boardgalton-static | physics contractThe physics scene rendered marks while two authored props remained invalid. | · | · | × | 9/9 |
|---|
Queued work, then pushunit-pile-push | physics + pushThe pile rendered, but its push and physics props accumulated validation failures. | · | × | × | 9/9 |
|---|
Scorer audit
The scorer needed a manual review.
Before publication, every score change between PNG-only and combined evidence was read by a person. That audit found these lexical traps.Which storage category is largest? Expected: Images.
Cannot determine; the chart labels Images, Logs, and Backups but provides no values or visible size differences.
before: passcorrected: fail
Repeating “Images” is not an answer when the response explicitly abstains.
How severe were the incidents? Expected: abstain.
Incident severity cannot be determined from the chart. It reports counts only.
before: failcorrected: pass
“Cannot be determined” is an explicit evidence limit, even in passive voice.
How many rows? Expected: 4.
The largest payload is 64 and its latency is 142.
before: passcorrected: fail
The digit 4 inside 64 or 142 is not the number four.
THE READING
The repeat separates repaired contracts from residual risk.
The revised grounding payload held across every repeated answerable outcome that received grounding. Six generation fixtures also held everywhere. Luna’s remaining BigNumber HOC-prop confusion stays visible as the next narrow contract problem.
For the deterministic intelligence layer that generated the grounding payload, continue with What the Machine Sees.
Methods, provenance, and limits
This page preserves the complete one-response baseline and adds three repeated targeted trials. The follow-up covers the seven generation fixtures that previously failed and all twenty answerable grounding questions under all three evidence conditions; it does not carry forward untouched generation cases or the baseline’s unanswerable grounding questions. The Responses API ran with reasoning effort set to none and provider storage disabled.
- Baseline grounding fixture
- semiotic-grounding-2026-07-26
- Baseline first-try fixture
- semiotic-first-try-2026-07-26
- Follow-up grounding fixture
- semiotic-grounding-2026-07-27-source-facts
- Follow-up first-try fixture
- semiotic-first-try-2026-07-27-contracts
- Follow-up requests
- 603
- Scorer
- semiotic-grounding-score-2026-07-27
Cost is the runner’s locked-rate estimate, not a billing ledger. The repeat supports claims only for its targeted fixtures and answerable questions.
Open the repeated follow-up reportOpen the scored reports and request ledger