313 English and Chinese questions for evaluating multi-entity web search, with reference answers dated 2026-09-18.
- Pooled reference answers · Strict reference answers
- Dataset description · Statistics
- Per-question results · Run records
Every configuration uses the same questions, answering model (gpt-5-mini, low reasoning effort), shared evidence-grounding rules and Entity-F1 scorer. Entity identities and aliases are shared across configurations. Results include both reference sets and all six pairwise comparisons, with Holm correction; tables are sorted by descending F1 with equal visual emphasis.
Octen uses broad_search followed by a reader; Exa, Parallel and Tavily use search-agent loops. The configured limit is eight subqueries/search actions with five retained results each. Backends, excerpts and answering workflows differ. See configuration and measurement details.
Pooled references:
| Configuration | F1 | Precision | Recall | API calls | Searches | Tokens | E2E (s) | Source domains |
|---|---|---|---|---|---|---|---|---|
| Octen broad_search | 0.5428 | 0.5709 | 0.5849 | 1.00 | 7.98 | 20,911 | 10.20 | 22.73 |
| Tavily-ultrafast | 0.4920 | 0.5039 | 0.5316 | 6.72 | 6.72 | 83,552 | 35.63 | 17.26 |
| Exa-instant | 0.4835 | 0.4726 | 0.5484 | 6.05 | 6.05 | 88,031 | 32.47 | 15.34 |
| Parallel-turbo | 0.4632 | 0.4620 | 0.5204 | 6.46 | 6.46 | 70,511 | 34.39 | 13.21 |
Strict reference scores
| Configuration | F1 | Precision | Recall |
|---|---|---|---|
| Octen broad_search | 0.5342 | 0.5352 | 0.6124 |
| Tavily-ultrafast | 0.4655 | 0.4537 | 0.5336 |
| Exa-instant | 0.4542 | 0.4218 | 0.5507 |
| Parallel-turbo | 0.4361 | 0.4130 | 0.5250 |
On these 313 questions, Octen broad_search ranks first in pooled F1, precision and recall, with higher F1 than all 3 agent configurations (Holm-adjusted p < 0.01); compared with Octen, the agent configurations use 6.1–6.7× as many logical API calls and 3.4–4.2× as many recorded downstream tokens, and take 3.2–3.5× as long end to end.
Quality metrics average all 313 tasks; other columns average completed runs. API calls count logical retrieval invocations; searches count recorded subqueries/search actions. Tokens are recorded downstream LLM usage; source domains count distinct retrieved domains per run.
All pairwise comparisons · Machine-readable summary
Python 3.10 or newer. Install dependencies and fill in the keys in .env:
pip install -e '.[analysis]'
cp .env.example .env
set -a; source .env; set +a
widesearch run data/tasks.jsonl --out my-run \
--arms exa-instant-agent,octen-broad-search,parallel-turbo-agent,tavily-ultrafast-agent \
--repeats 1 --concurrency 5Re-score the published answers without API calls:
python tools/regrade.py --grades results/grades.jsonl \
--gold data/tasks.jsonl --out /tmp/pooled.jsonl
python tools/regrade.py --grades results/grades.jsonl \
--gold data/tasks_strict.jsonl --out /tmp/strict.jsonl
python tools/release_report.py --checkMIT. Cite WideSearch-Bench 2026.09 and the commit used.
