Agent PM: OpenAI Agents-powered product management orchestrator with automated PRDs, tickets, and comms.
-
Updated
Sep 27, 2026 - TypeScript
Agent PM: OpenAI Agents-powered product management orchestrator with automated PRDs, tickets, and comms.
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
AI Agent EvalOps dashboard for regression testing, drift detection, run comparison, and quality analysis.
EvalOps platform for continuous AI-quality evaluation
CI-ready LLM regression harness for document-grounded RAG assistants: hard checks, pairwise judging, bootstrap CIs, and artifact reports.
Evidence-aware benchmark and PR advisory prototype for AI coding-agent context safety.
RAG pipeline evaluation for .NET — LLM-as-judge scoring for Faithfulness, Answer Relevance, Context Precision, and Context Recall. The evaluation layer your .NET RAG pipeline is missing.
LLM gateway with explainable routing, streaming failover boundaries, budget reservations, and evaluation regression gates. Zero-dependency Python demo.
Hermes plugin: route LLM calls through the EvalOps gateway and report agent registration + tool spans to the EvalOps platform.
To associate your repository with the evalops topic, visit your repo's landing page and select "manage topics."