$ai-evals
/companies

Companies

Every company in the AI evals space we've reviewed. Independent — we don't accept vendor sponsorships, and reviews are updated as products change.

Category
Pricing
other
25 companies

Arize AI

7.2

ML observability platform extended into LLMs, with the open-source Phoenix framework as a popular standalone trace viewer.

observabilityML monitoringLLM evals·freemium

Braintrust

9.1

Eval-driven dev platform combining traces, datasets, scorers, and a playground in one product.

LLM evalsobservabilityprompt management·freemium

Comet (Opik)

7.4

Open-source LLM evaluation and observability from a mature MLOps team — credible Langfuse alternative.

observabilityLLM evalsMLOps·open-source

Datadog

6.4

APM giant with bolted-on LLM observability for OpenAI and Anthropic calls.

observabilityAPM·paid

DeepEval (Confident AI)

7.6

pytest-style LLM evaluation framework with synthetic dataset generation and CI/CD-native testing.

LLM evals·freemium

Evidently AI

7.0

Open-source ML and LLM evaluation framework with strong methodology docs — building blocks, not a finished platform.

LLM evalsML monitoring·open-source

Fiddler

7.2

Enterprise ML governance platform extended to LLMs and generative AI, with audit-ready traces and in-environment evaluations.

AI governanceagent observability·enterprise

Galileo

7.5

Agent reliability platform with cheap, fast evaluators that can run on every request in production.

agent observabilityLLM evals·freemium

HUD

7.8

Open-source platform for building RL environments and evals for computer-use agents — used by frontier labs, ships its own benchmarks.

agent observabilityRL environmentsbenchmarks·open-source

Label Studio

7.0

Open-source data annotation platform with rubric enforcement, escalation workflows, and audit trails — extended to LLM review.

annotationLLM evals·open-source

Langfuse

8.4

Open-source LLM observability with evals, prompt management, and best-in-class tracing.

observabilityLLM evalsprompt management·open-source

LangSmith

7.5

Observability and evaluation built by the LangChain team — best-in-class if your stack is LangChain or LangGraph.

observabilityLLM evalsprompt management·freemium

LiteLLM

8.0

Open-source Python SDK and proxy that translates requests across 100+ LLM providers into the OpenAI format.

LLM gatewaymulti-provider routing·open-source

Maxim AI

6.8

AI quality evaluation platform with prebuilt and custom scorers, designed to plug into existing observability stacks.

LLM evals·freemium

MLflow

6.6

Open-source MLOps standard with LLM tracing, evaluation, and prompt management bolted on top.

MLOpsobservabilityLLM evals·open-source

OpenRouter

8.2

Single OpenAI-compatible endpoint to 500+ models across 60+ providers, billed pay-as-you-go.

LLM gatewaymulti-provider routing·paid

Portkey

7.8

Full-stack AI gateway with the broadest model catalog, built-in guardrails, and enterprise-grade governance.

LLM gatewaymulti-provider routingAI governance·freemium

Promptfoo

7.4

Open-source CLI for evaluating LLM prompts and red-teaming applications, with YAML/JSON configs that live next to your code.

LLM evalsred-teaming·open-source

PromptHub

6.8

Git-style version control for prompts — branch, commit, merge, and CI-gate prompt changes.

prompt management·freemium

PromptLayer

7.0

Visual prompt editor and version control built for non-technical teams.

prompt management·freemium

RAGAS

7.5

Open-source evaluation framework purpose-built for RAG pipelines, with reference-free metrics that became the industry standard.

LLM evalsRAG evaluation·open-source

SuperAnnotate

6.8

Annotation platform with strong tooling for measuring and resolving disagreements between human reviewers and automated scorers.

annotationLLM evals·paid

Vellum

7.0

Visual workflow builder with built-in observability for low-code agent development.

prompt managementagent observability·freemium

Weights & Biases Weave

6.8

LLM tracing, evaluation, and prompt management embedded inside the Weights & Biases ML platform.

observabilityLLM evalsprompt managementMLOps·freemium

ZenML

6.8

Open-source MLOps and LLMOps framework for building reproducible, infrastructure-agnostic AI pipelines.

MLOpsLLM evals·freemium