Companies
Every company in the AI evals space we've reviewed. Independent — we don't accept vendor sponsorships, and reviews are updated as products change.
Arize AI
7.2ML observability platform extended into LLMs, with the open-source Phoenix framework as a popular standalone trace viewer.
Braintrust
9.1Eval-driven dev platform combining traces, datasets, scorers, and a playground in one product.
Comet (Opik)
7.4Open-source LLM evaluation and observability from a mature MLOps team — credible Langfuse alternative.
Datadog
6.4APM giant with bolted-on LLM observability for OpenAI and Anthropic calls.
DeepEval (Confident AI)
7.6pytest-style LLM evaluation framework with synthetic dataset generation and CI/CD-native testing.
Evidently AI
7.0Open-source ML and LLM evaluation framework with strong methodology docs — building blocks, not a finished platform.
Fiddler
7.2Enterprise ML governance platform extended to LLMs and generative AI, with audit-ready traces and in-environment evaluations.
Galileo
7.5Agent reliability platform with cheap, fast evaluators that can run on every request in production.
HUD
7.8Open-source platform for building RL environments and evals for computer-use agents — used by frontier labs, ships its own benchmarks.
Label Studio
7.0Open-source data annotation platform with rubric enforcement, escalation workflows, and audit trails — extended to LLM review.
Langfuse
8.4Open-source LLM observability with evals, prompt management, and best-in-class tracing.
LangSmith
7.5Observability and evaluation built by the LangChain team — best-in-class if your stack is LangChain or LangGraph.
LiteLLM
8.0Open-source Python SDK and proxy that translates requests across 100+ LLM providers into the OpenAI format.
Maxim AI
6.8AI quality evaluation platform with prebuilt and custom scorers, designed to plug into existing observability stacks.
MLflow
6.6Open-source MLOps standard with LLM tracing, evaluation, and prompt management bolted on top.
OpenRouter
8.2Single OpenAI-compatible endpoint to 500+ models across 60+ providers, billed pay-as-you-go.
Portkey
7.8Full-stack AI gateway with the broadest model catalog, built-in guardrails, and enterprise-grade governance.
Promptfoo
7.4Open-source CLI for evaluating LLM prompts and red-teaming applications, with YAML/JSON configs that live next to your code.
PromptHub
6.8Git-style version control for prompts — branch, commit, merge, and CI-gate prompt changes.
PromptLayer
7.0Visual prompt editor and version control built for non-technical teams.
RAGAS
7.5Open-source evaluation framework purpose-built for RAG pipelines, with reference-free metrics that became the industry standard.
SuperAnnotate
6.8Annotation platform with strong tooling for measuring and resolving disagreements between human reviewers and automated scorers.
Vellum
7.0Visual workflow builder with built-in observability for low-code agent development.
Weights & Biases Weave
6.8LLM tracing, evaluation, and prompt management embedded inside the Weights & Biases ML platform.
ZenML
6.8Open-source MLOps and LLMOps framework for building reproducible, infrastructure-agnostic AI pipelines.