Eval framework. Define correct, test against it, get results.
-
Updated
Feb 17, 2026 - Go
Eval framework. Define correct, test against it, get results.
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
RAG context precision and context recall evaluator measuring retrieved chunk relevance and rank quality
RAG context precision and context recall evaluator measuring retrieved chunk relevance and rank quality
A web-based interactive demo for the GuessArena evaluation framework
One-stop CLI for running structured skill evals across all agent harnesses
Curated AI agent evaluation skills from Microsoft's Eval Guide — plan, generate, run, and interpret eval suites for Copilot Studio agents
4-model parallel planning workflow with eval framework — Claude, Gemini, Codex, GLM-5 · OpenClaw ecosystem
Binary safety verdicts (SAFE/HELD/LEAK/MISS/BROKE) + persona fan-out for LLM pipeline evals
Observability layer for multi-step AI pipelines — traces execution, auto-diagnoses root causes, and builds a growing eval dataset from human feedback
Self-hosted evaluation framework for tool-calling LLM agents. Define tasks in YAML with expected tool calls and scoring rubrics, run against Claude or GPT-4, get scored reports via CLI and FastAPI dashboard. Lightweight, local-first, no lock-in.
Evaluation framework for testing LLM outputs locally. Define prompt templates and custom scorers, run evals against OpenAI or Anthropic models, store results in SQLite, and browse via Rich CLI dashboard or FastAPI UI. Lightweight, self-contained, extensible Python library.
Open-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctness and safety, audits for security issues, and visualizes traces in a web dashboard.
Open-source evaluation framework for LLM agents. Run head-to-head A/B tests, score with LLM-as-judge rubrics, and visualize results in a Streamlit dashboard. Model-agnostic, self-hosted, zero external infrastructure.
Python tool for extracting validated, schema-defined JSON from documents (PDF, HTML, Markdown, plaintext) using LLMs, with source grounding and a built-in eval harness. Works with Anthropic Claude or OpenAI.
Open-source evaluation framework for AI agents. Define test suites with rubrics, run your agent, get LLM-as-judge scores against criteria, inspect full execution traces, and diff runs to catch behavioral regressions.
Define YAML rubrics, run agents through test scenarios, get LLM-judged per-criterion scores with full trajectory traces, and analyze results in an interactive web dashboard.
Lightweight Python evaluation framework for LLM prompts. Define test cases in YAML, run against OpenAI and Anthropic models, score outputs with LLM-as-judge, view results in a web dashboard. Built for regression testing, prompt comparison, and quality assurance.
Multi-model LLM evaluation using Promptfoo — benchmarks Claude, Gemini, and GPT-4.1-mini on a review classification task, combining deterministic JSON checks with LLM-as-a-judge assertions for field names, valid values, and correctness.
A FastAPI WebSocket service and CI-friendly CLI for running LangSmith evaluations on demand — bundles LLM-as-judge (Claude) and heuristic evaluators, syncs datasets idempotently, and gates deploys via threshold checks.
To associate your repository with the eval-framework topic, visit your repo's landing page and select "manage topics."