The field guide to AI eval tools, written for engineers.
We test, compare, and rate the platforms teams use to measure LLMs and agents in production — observability, offline & online evals, prompt management, gateways, and red-teaming.
- 25
- companies reviewed
- Jul 23, 2026
- last updated
Featured companies
Braintrust
9.1Eval-driven dev platform combining traces, datasets, scorers, and a playground in one product.
Fiddler
7.2Enterprise ML governance platform extended to LLMs and generative AI, with audit-ready traces and in-environment evaluations.
Galileo
7.5Agent reliability platform with cheap, fast evaluators that can run on every request in production.
HUD
7.8Open-source platform for building RL environments and evals for computer-use agents — used by frontier labs, ships its own benchmarks.
Langfuse
8.4Open-source LLM observability with evals, prompt management, and best-in-class tracing.
LiteLLM
8.0Open-source Python SDK and proxy that translates requests across 100+ LLM providers into the OpenAI format.
Recent editorial
Arize vs LangSmith (2026)
A monitoring-first LLM platform against a LangChain-native observability tool. Two different origins — here's which fits which team.
Braintrust vs Arize (2026)
An LLM-native eval platform against a monitoring-first tool that grew into LLMs. The right pick depends on whether you want to watch models or ship LLM features.
Braintrust vs Langfuse (2026)
The best closed-source eval platform against the best open-source one. The decision comes down to one question — do you have to self-host?