I've been reading backend/pyspur/evals/ — evaluator.py and common.py (the multiple-choice/math-equivalence grading helpers, normalize_response/normalize_extracted_answer/extract_answer_with_regex) plus evals/tasks/ — which is effectively PySpur's own eval-harness layer sitting on top of the visual workflow graph.
I maintain EvalPort, a JSON-Schema-based interchange spec (TestCase/Suite/Grader/Result/ResultSet) for portable LLM eval data, built so eval data isn't locked into one tool's format. It's early-stage (~35 shipped adapters, no notable star count — being upfront about that).
What I don't see in pyspur/evals is any way to load a task set that wasn't authored specifically for PySpur, or to export a completed eval run's graded results in a form another framework's tooling could consume/compare against. An EvalPort Suite (portable TestCases with reference_outputs, comparable to what evaluator.py's tasks already carry) could feed evals/tasks, and the graded output (per-task score/verdict from evaluator.py) could export as an EvalPort ResultSet.
I'd propose a standalone, optional adapter for exactly that: import an EvalPort Suite as a pyspur.evals task set, and export evaluator.py's run output as an EvalPort ResultSet. No required dependency, no change to the workflow graph or existing eval task format.
Happy to build this as a PR into pyspur (e.g. backend/pyspur/evals/evalport.py), or as a standalone package in EvalPort's own adapters/ directory with zero footprint on this repo — whichever you'd prefer. Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
— Sahi, independent contributor (not affiliated with PySpur)
I've been reading
backend/pyspur/evals/—evaluator.pyandcommon.py(the multiple-choice/math-equivalence grading helpers,normalize_response/normalize_extracted_answer/extract_answer_with_regex) plusevals/tasks/— which is effectively PySpur's own eval-harness layer sitting on top of the visual workflow graph.I maintain EvalPort, a JSON-Schema-based interchange spec (
TestCase/Suite/Grader/Result/ResultSet) for portable LLM eval data, built so eval data isn't locked into one tool's format. It's early-stage (~35 shipped adapters, no notable star count — being upfront about that).What I don't see in
pyspur/evalsis any way to load a task set that wasn't authored specifically for PySpur, or to export a completed eval run's graded results in a form another framework's tooling could consume/compare against. An EvalPortSuite(portableTestCases withreference_outputs, comparable to whatevaluator.py's tasks already carry) could feedevals/tasks, and the graded output (per-task score/verdict fromevaluator.py) could export as an EvalPortResultSet.I'd propose a standalone, optional adapter for exactly that: import an EvalPort
Suiteas apyspur.evalstask set, and exportevaluator.py's run output as an EvalPortResultSet. No required dependency, no change to the workflow graph or existing eval task format.Happy to build this as a PR into pyspur (e.g.
backend/pyspur/evals/evalport.py), or as a standalone package in EvalPort's ownadapters/directory with zero footprint on this repo — whichever you'd prefer. Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md— Sahi, independent contributor (not affiliated with PySpur)