google/brotli

Brotli compression format

14,673 c
441
Generated Behavioral Tests
95.5%
Best Score
GPT 5.5 (xhigh)

Hover for details · Click a model to open it · Drag to zoom the x-axis · Shift-drag to pan · The line marks the Pareto frontier (best score per cost)

25 runs
# Model Score Cost Calls
1 GPT 5.5 (xhigh) 95.5% $14.46 116 trace →
2 Muse Spark 1.3 (max) 95.2% $4.45 258 trace →
3 Muse Spark 1.3 (xhigh) 93.7% $1.16 210 trace →
4 Claude Sonnet 4.6 90.7% $19.93 394 trace →
5 GLM-5.2 85.9% $9.68 135 trace →
6 Claude Opus 4.7 84.6% $5.68 110 trace →
7 GPT 5.5 (high) 82.1% $4.86 59 trace →
8 Muse Spark 1.1 (xhigh) 82.1% $3.19 247 trace →
9 Claude Opus 4.6 80.3% $10.59 262 trace →
10 Gemini 3.7 Flash 79.6% $1.69 121 trace →
11 Muse Spark 1.2 (xhigh) 78.9% $0.40 264 trace →
12 Gemini 3.5 Flash 69.6% $6.04 199 trace →
13 GPT-5.6 Sol 68.7% $0.52 12 trace →
14 Gemini 3.6 Flash 61.7% $3.46 133 trace →
15 Gemini 3.1 Pro 52.8% $1.03 69 trace →
16 Gemini 3 Flash 46.7% $0.34 100 trace →
17 GPT 5 mini 1.6% $0.02 14 trace →
18 Claude Opus 5 (xhigh) 1.1% $49.72 221 trace →
19 GPT 5.5 1.1% $1.13 22 trace →
20 GPT 5.4 0.9% $0.52 18 trace →
21 Claude Haiku 4.5 0.9% $0.63 122 trace →
22 GPT-5.6 Sol (xhigh) 0.7% $7.31 43 trace →
23 Claude Opus 4.8 (xhigh) 0.7% $28.02 173 trace →
24 GPT 5.4 mini 0.7% $0.05 11 trace →
25 Claude Opus 4.7 (xhigh) 0.5% $6.32 94 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run