Quality probes
On tasks with objectively checkable outcomes, does re-released Fable 5 (2026-07-03+) outperform Opus 4.8 enough to justify its 2x API rate once Fable leaves the Max subscription?
Verdictnull-result
A native dashboard over the curated public benchmark snapshot: cost trends, per-model activity and the experiment record — real data rendered as themeable, accessible charts, not an embedded frame.
Data as of 2026-07-20.
24
Benchmark cycles
9
Models tracked
9
Frozen tasks
832
Total sessions
Benchmark-suite cost (USD) per model as the CLI advances. Each line is one model; gaps are versions where that model was not run.
| CLI version | claude-haiku-4-5 | claude-opus-4-7 | claude-opus-4-8 | claude-sonnet-4-6 | claude-fable-5 | claude-sonnet-5 |
|---|---|---|---|---|---|---|
| 2.1.156 | $0.68 | $2.91 | $1.20 | $1.30 | — | — |
| 2.1.170 | — | — | — | — | $2.60 | — |
| 2.1.196 | — | — | — | $1.38 | $7.93 | $1.48 |
| 2.1.201 | — | — | $2.91 | — | $8.40 | — |
| 2.1.202 | $0.14 | — | $0.79 | — | $1.59 | $0.46 |
| 2.1.204 | $0.14 | — | $0.75 | — | $1.59 | $0.64 |
| 2.1.205 | $0.14 | — | $0.69 | — | $1.41 | $0.66 |
| 2.1.206 | $0.13 | — | $0.71 | — | $1.45 | $0.62 |
| 2.1.207 | $0.16 | — | $1.08 | — | $1.17 | $0.47 |
| 2.1.208 | $0.14 | — | $0.45 | — | $1.05 | $0.43 |
| 2.1.210 | $0.14 | — | $0.50 | — | $1.05 | $0.39 |
| 2.1.211 | $0.14 | — | $0.44 | — | $1.04 | $0.39 |
| 2.1.212 | $0.13 | — | $0.45 | — | $1.04 | $0.40 |
| 2.1.214 | $0.14 | — | $0.47 | — | $1.00 | $0.41 |
| 2.1.215 | $0.13 | — | $0.40 | — | $0.99 | $0.40 |
Total benchmark output tokens attributed to each model in the snapshot.
| Model | Family | Sessions | Output tokens |
|---|---|---|---|
| claude-fable-5 | fable | 157 | 341,301 |
| claude-opus-4-8 | opus | 180 | 164,977 |
| claude-sonnet-4-6 | sonnet | 90 | 81,464 |
| claude-haiku-4-5 | haiku | 90 | 58,767 |
| claude-sonnet-5 | sonnet | 90 | 53,471 |
| claude-opus-4-7 | opus | 90 | 52,384 |
| global.anthropic.claude-haiku-4-5-20251001-v1:0 | haiku | 45 | 28,669 |
| global.anthropic.claude-opus-4-7 | opus | 45 | 25,699 |
| global.anthropic.claude-sonnet-4-6 | sonnet | 45 | 23,653 |
On tasks with objectively checkable outcomes, does re-released Fable 5 (2026-07-03+) outperform Opus 4.8 enough to justify its 2x API rate once Fable leaves the Max subscription?
Verdictnull-result
At what task difficulty (if any) does Fable 5 begin to outperform Opus 4.8 on objectively checkable outcomes, and does the premium scale with difficulty?
Verdictno-breakpoint-found
How does a /btw message appear in the JSONL transcript and block-counts DB?
Verdictconfirmed(high confidence)
Do sequential Edit tool calls on files with repetitive structure produce anchor-drift failures, and at what minimum N does the first failure occur?
Verdictconfirmed(high confidence)
When Edit calls succeed without failure, which approach consumes fewer total blocks -- N sequential Edits or Write + Bash (Python rewrite)?
Verdictconfirmed(high confidence)
At what N does a Python rewrite (Write + Bash) become more token-efficient than N sequential Edit calls, and how does required context depth K affect that crossover?
Verdictconfirmed(high confidence)
On tasks with objectively checkable outcomes, does re-released Fable 5 (2026-07-03+) outperform Opus 4.8 enough to justify its 2x API rate once Fable leaves the Max subscription?
Verdictnull-result(medium confidence)
At what task difficulty (if any) does Fable 5 begin to outperform Opus 4.8 on objectively checkable outcomes, and does the premium scale with difficulty?
Verdictrefuted(high confidence)