These are my two attempts using an image build from the main branch as commit 0883bd1.
docker run --rm -it --entrypoint=aiperf --env HF_TOKEN=$HF_TOKEN -v ~/Downloads/aiperf/app:/app --user=$(id -u) \
OMIT/osc/aiperf:main-0883bd1a-osc-r0 profile \
--model qwen35-122b \
--tokenizer Qwen/Qwen3.5-122B-A10B \
--tokenizer-trust-remote-code \
--url https://OMIT/api \
--endpoint-type chat --streaming \
--api-key OMIT \
--no-server-metrics \
--accuracy-benchmark mmlu_pro \
--num-requests 30 \
--concurrency 10 \
--accuracy-verbose \
--use-server-token-count \
--extra-inputs '{"temperature": 0}'
docker run --rm -it --entrypoint=aiperf --env HF_TOKEN=$HF_TOKEN -v ~/Downloads/aiperf/app:/app --user=$(id -u) \
OMIT/osc/aiperf:main-0883bd1a-osc-r0 profile \
--model qwen35-122b \
--tokenizer Qwen/Qwen3.5-122B-A10B \
--tokenizer-trust-remote-code \
--url https://OMIT/api \
--endpoint-type chat --streaming \
--api-key OMIT \
--no-server-metrics \
--accuracy-benchmark mmlu_pro \
--accuracy-tasks "computer science",physics,math \
--num-requests 30 \
--concurrency 10 \
--accuracy-verbose \
--use-server-token-count \
--extra-inputs '{"temperature": 0}'
When I request 3 different tasks (second command above), only one is reported and it's always math which is only 1 of the requested categories:
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━┓
┃ Task ┃ Correct ┃ Total ┃ Unparsed ┃ Accuracy ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━┩
│ math │ 28 │ 30 │ 1 │ 93.33% │
├────────────────────────────────┼─────────┼───────┼──────────┼──────────┤
│ OVERALL │ 28 │ 30 │ 1 │ 93.33% │
└────────────────────────────────┴─────────┴───────┴──────────┴──────────┘
CLI Command: aiperf profile --model 'qwen35-122b' --tokenizer 'Qwen/Qwen3.5-122B-A10B' --tokenizer-trust-remote-code --url 'https://ai-alpha.osc.edu/api'
--endpoint-type 'chat' --streaming --api-key '<redacted>' --no-server-metrics --accuracy-benchmark 'mmlu_pro' --accuracy-tasks 'computer
science,physics,math' --num-requests 30 --concurrency 10 --accuracy-verbose --use-server-token-count --extra-inputs '{"temperature": 0}'
Benchmark Duration: 229.82 sec
The inputs show the other categories but other files do not besides showing the requested value:
$ grep -HnR -c "computer science" ~/Downloads/aiperf/app/artifacts/
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_results.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.json:2
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/logs/aiperf.log:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/inputs.json:410
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_console.txt:0
$ grep -HnR -c "physics" ~/Downloads/aiperf/app/artifacts/
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_results.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.json:2
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/logs/aiperf.log:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/inputs.json:1299
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_console.txt:0
$ grep -HnR -c "math" ~/Downloads/aiperf/app/artifacts/
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_export.jsonl:30
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_results.csv:1
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.json:4
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/logs/aiperf.log:3
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/inputs.json:1503
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.csv:2
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_console.txt:1
When I omit --accuracy-tasks (first command above) I still only get one reported category, and it's always business:
Accuracy Benchmark Results
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━┓
┃ Task ┃ Correct ┃ Total ┃ Unparsed ┃ Accuracy ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━┩
│ business │ 23 │ 30 │ 3 │ 76.67% │
├────────────────────────────────┼─────────┼───────┼──────────┼──────────┤
│ OVERALL │ 23 │ 30 │ 3 │ 76.67% │
└────────────────────────────────┴─────────┴───────┴──────────┴──────────┘
CLI Command: aiperf profile --model 'qwen35-122b' --tokenizer 'Qwen/Qwen3.5-122B-A10B' --tokenizer-trust-remote-code --url 'https://ai-alpha.osc.edu/api'
--endpoint-type 'chat' --streaming --api-key '<redacted>' --no-server-metrics --accuracy-benchmark 'mmlu_pro' --num-requests 30 --concurrency 10
--accuracy-verbose --use-server-token-count --extra-inputs '{"temperature": 0}'
Benchmark Duration: 320.84 sec
There are other categories in the inputs but no where else, so only business is being reported:
$ grep -HnR -c "physics" ~/Downloads/aiperf/app/artifacts/
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_results.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.json:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/logs/aiperf.log:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/inputs.json:1798
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_console.txt:0
$ grep -HnR -c "computer science" ~/Downloads/aiperf/app/artifacts/
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_results.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.json:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/logs/aiperf.log:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/inputs.json:410
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.csv:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_console.txt:0
$ grep -HnR -c "business" ~/Downloads/aiperf/app/artifacts/
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_export.jsonl:30
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/accuracy_results.csv:1
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.json:2
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/logs/aiperf.log:6
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export.jsonl:0
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/inputs.json:975
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_aiperf.csv:2
/Users/tdockendorf/Downloads/aiperf/app/artifacts//qwen35-122b-openai-chat-concurrency10/profile_export_console.txt:1
These are my two attempts using an image build from the main branch as commit 0883bd1.
When I request 3 different tasks (second command above), only one is reported and it's always
mathwhich is only 1 of the requested categories:The inputs show the other categories but other files do not besides showing the requested value:
When I omit
--accuracy-tasks(first command above) I still only get one reported category, and it's alwaysbusiness:There are other categories in the inputs but no where else, so only
businessis being reported: