segmentio/chamber

CLI for managing secrets

2,588 go medium
1,748
Generated Behavioral Tests
93.1%
Best Score
Claude Opus 5 (xhigh)

Hover for details · Click a model to open it · Drag to zoom the x-axis · Shift-drag to pan · The line marks the Pareto frontier (best score per cost)

25 runs
# Model Score Cost Calls
1 Claude Opus 5 (xhigh) 93.1% $35.48 216 trace →
2 Muse Spark 1.3 (max) 89.1% $5.69 324 trace →
3 Gemini 3.7 Flash 88.2% $1.62 126 trace →
4 GPT 5.5 (xhigh) 88.0% $15.51 99 trace →
5 Claude Opus 4.8 (xhigh) 86.4% $12.86 128 trace →
6 GPT-5.6 Sol (xhigh) 85.5% $6.80 54 trace →
7 Muse Spark 1.3 (xhigh) 82.8% $1.35 227 trace →
8 Claude Opus 4.6 82.0% $11.99 232 trace →
9 GLM-5.2 80.0% $15.24 155 trace →
10 Muse Spark 1.1 (xhigh) 79.8% $0.38 319 trace →
11 Claude Opus 4.7 (xhigh) 79.3% $7.80 140 trace →
12 Muse Spark 1.2 (xhigh) 79.3% $0.27 182 trace →
13 GPT 5.5 (high) 79.0% $3.86 46 trace →
14 Gemini 3.6 Flash 78.5% $3.09 101 trace →
15 GPT-5.6 Sol 72.4% $0.95 13 trace →
16 GPT 5.5 71.3% $1.20 18 trace →
17 Gemini 3.5 Flash 65.8% $5.22 169 trace →
18 Claude Sonnet 4.6 62.1% $14.17 359 trace →
19 GPT 5.4 56.2% $0.24 9 trace →
20 Claude Haiku 4.5 43.0% $0.78 117 trace →
21 Gemini 3 Flash 40.3% $0.21 77 trace →
22 GPT 5 mini 20.9% $0.04 24 trace →
23 Gemini 3.1 Pro 19.9% $1.34 66 trace →
24 GPT 5.4 mini 14.5% $0.03 9 trace →
25 Claude Opus 4.7 9.4% $3.95 92 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run