Best AI LLM Evaluation Tools in 2026

In short: Weights & Biases is ranked #1 of 30 as of 8 October 2026, ahead of Evidently AI and Vellum. The best-ranked option with a free plan is Evidently AI. The lowest first paid tier on this page is Opik at $19/mo.

When you need to assess language-model behavior, evaluation tools offer different methods, model support, and safety checks to consider. Compare evaluation methods and safety evaluations, then look at model support for the systems you work with. Prompt versioning, API access, and deployment add further dimensions for comparing how a tool could fit into your workflow. Weights & Biases, Evidently AI, and Vellum are among the entries to examine. Free-plan availability and paid-from pricing help you weigh access and cost against the evaluation capabilities listed.

30 AI LLM evaluation tools ranked on what their makers publish — plans and prices, free tiers, platforms and the facts on their own pages.

30ranked
20free plans on this page
$19/molowest paid tier
8 Oct 2026last checked
Input list AI LLM Evaluation Tools 25 channels on this page · 68 of 200 spec lines stated by the makers
Ch Tool Free planPaid fromEvaluation methodsModel supportSafety evaluationsDeploymentPrompt versioningAPI access Spec sheet Score
01 Weights & Biases Free planYesPaid from$60/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 2/8spec lines stated 8.0
02 Evidently AI Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 7.8
03 Vellum Free planYesPaid from$30/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessYes 3/8spec lines stated 7.7
04 Opik Free planYesPaid from$19/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 2/8spec lines stated 7.5
05 Promptfoo Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 7.4
06 Maxim AI Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningYesAPI accessnot stated 2/8spec lines stated 7.3
07 DeepEval Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsYesDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 2/8spec lines stated 7.2
08 Giskard Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 7.1
09 LangWatch Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 7.1
10 Rhesis AI Free planYesPaid fromnot statedEvaluation methodsoffline evaluation, online trace metrics, LLM-as-a-judge, custom metrics, single-turn testing, multi-turn testing, adversarial red-teamingModel supportOpenAI, Anthropic, Google Gemini, Azure OpenAI, Mistral, Cohere, Groq, Together AI, Perplexity, Replicate, Ollama, vLLM, LiteLLM ProxySafety evaluationsYesDeploymentnot statedPrompt versioningnot statedAPI accessYes 5/8spec lines stated 7.1
11 Galileo Free planYesPaid from$100/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningYesAPI accessnot stated 3/8spec lines stated 7.0
12 Langfuse Free planYesPaid from$29/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 2/8spec lines stated 7.0
13 Braintrust Free planYesPaid from$249/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningYesAPI accessnot stated 3/8spec lines stated 6.9
14 LangSmith Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 6.9
15 NVIDIA NeMo Evaluator Free plannot statedPaid fromnot statedEvaluation methodsBuilt-in benchmarks, exact match, fuzzy match, multiple-choice regex, answer-line, numeric match, code sandbox, LLM-as-judge, JSON-schema validation, regression comparison, quality gatesModel supportOpenAI-compatible model endpoints, NVIDIA API Catalog, vLLM, NIM, local vLLM, SGLang, TensorRT-LLM, vision-language modelsSafety evaluationsYesDeploymentself-hostedPrompt versioningNoAPI accessYes 6/8spec lines stated 6.9
16 Arize Phoenix Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 6.8
17 Confident AI Free planYesPaid from$200/moEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningYesAPI accessnot stated 3/8spec lines stated 6.8
18 HoneyHive Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentnot statedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 6.8
19 OpenCompass Free plannot statedPaid fromnot statedEvaluation methodsobjective; subjective; discriminative; generative; LLM-as-a-judgeModel supportHugging Face models; API-based models; custom models; OpenAI; Anthropic; Gemini; Qwen; GLM; DeepSeekSafety evaluationsYesDeploymentself-hostedPrompt versioningnot statedAPI accessnot stated 4/8spec lines stated 6.8
20 OpenAI Evals Free plannot statedPaid fromnot statedEvaluation methodsbasic exact/match evaluations, model-graded evaluations, custom evaluation logic, academic benchmarks, meta-evaluationsModel supportOpenAI API models and custom CompletionFunction implementationsSafety evaluationsYesDeploymenthybridPrompt versioningYesAPI accessYes 6/8spec lines stated 6.7
21 Pydantic Evals Free planYesPaid fromnot statedEvaluation methodsDeterministic checks; custom evaluators; LLM judges; G-Eval; performance checks; report evaluators; span-based evaluation; agentic trajectory evaluationModel supportOpenAI; Anthropic; Gemini; xAI; Bedrock; Cerebras; Cohere; Groq; Hugging Face; Mistral; OpenRouter; and other listed Pydantic AI providersSafety evaluationsYesDeploymentself-hostedPrompt versioningYesAPI accessYes 7/8spec lines stated 6.7
22 UpTrain Free plannot statedPaid fromnot statedEvaluation methodspreconfigured checks; custom prompt evaluations; custom Python evaluations; model-graded evaluations; classification; chain-of-thought classification; regression testing; experimentsModel supportOpenAI; Azure; Claude; Mistral; Together AI; Anyscale; Ollama; Hugging Face; Replicate; custom endpointsSafety evaluationsYesDeploymenthybridPrompt versioningYesAPI accessYes 6/8spec lines stated 6.7
23 LM Evaluation Harness Free plannot statedPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentself-hostedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 6.6
24 Ragas Free planYesPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentself-hostedPrompt versioningYesAPI accessnot stated 3/8spec lines stated 6.6
25 TruLens Free plannot statedPaid fromnot statedEvaluation methodsnot statedModel supportnot statedSafety evaluationsnot statedDeploymentself-hostedPrompt versioningnot statedAPI accessnot stated 1/8spec lines stated 6.6
Compare all 25 in a table
#ToolScoreFree planFromFree planPaid fromEvaluation methodsModel support
1Weights & Biases8.0Free plan$60/moYes60 /mo——
2Evidently AI7.8Free plan$80/moYes———
3Vellum7.7Free plan$30/moYes30 /mo——
4Opik7.5Free plan$19/moYes19 /mo——
5Promptfoo7.4Free planFreeYes———
6Maxim AI7.3Free plan$29/moYes———
7DeepEval7.2Free planFreeYes———
8Giskard7.1Free planFreeYes———
9LangWatch7.1Free plan€29/moYes———
10Rhesis AI7.1Free planFreeYes—offline evaluation, online trace metrics, LLM-as-a-judge, custom metrics, single-turn testing, multi-turn testing, adversarial red-teamingOpenAI, Anthropic, Google Gemini, Azure OpenAI, Mistral, Cohere, Groq, Together AI, Perplexity, Replicate, Ollama, vLLM, LiteLLM Proxy
11Galileo7.0Free plan$100/moYes100 /mo——
12Langfuse7.0Free plan$29/moYes29 /mo——
13Braintrust6.9Free plan$249/moYes249 /mo——
14LangSmith6.9Free plan$39/moYes———
15NVIDIA NeMo Evaluator6.9Free planFree——Built-in benchmarks, exact match, fuzzy match, multiple-choice regex, answer-line, numeric match, code sandbox, LLM-as-judge, JSON-schema validation, regression comparison, quality gatesOpenAI-compatible model endpoints, NVIDIA API Catalog, vLLM, NIM, local vLLM, SGLang, TensorRT-LLM, vision-language models
16Arize Phoenix6.8Free planFreeYes———
17Confident AI6.8Free plan$200/moYes200 /mo——
18HoneyHive6.8Free planFreeYes———
19OpenCompass6.8No———objective; subjective; discriminative; generative; LLM-as-a-judgeHugging Face models; API-based models; custom models; OpenAI; Anthropic; Gemini; Qwen; GLM; DeepSeek
20OpenAI Evals6.7No———basic exact/match evaluations, model-graded evaluations, custom evaluation logic, academic benchmarks, meta-evaluationsOpenAI API models and custom CompletionFunction implementations
21Pydantic Evals6.7No—Yes—Deterministic checks; custom evaluators; LLM judges; G-Eval; performance checks; report evaluators; span-based evaluation; agentic trajectory evaluationOpenAI; Anthropic; Gemini; xAI; Bedrock; Cerebras; Cohere; Groq; Hugging Face; Mistral; OpenRouter; and other listed Pydantic AI providers
22UpTrain6.7No———preconfigured checks; custom prompt evaluations; custom Python evaluations; model-graded evaluations; classification; chain-of-thought classification; regression testing; experimentsOpenAI; Azure; Claude; Mistral; Together AI; Anyscale; Ollama; Hugging Face; Replicate; custom endpoints
23LM Evaluation Harness6.6Free planFree————
24Ragas6.6No—Yes———
25TruLens6.6Free planFree————

Is your tool on this list?

Numbered spots on this list can be sponsored, and a sponsored row is labelled as paid.

Questions about this list

Which AI LLM evaluation tool is ranked first on Specifiction?

Weights & Biases is ranked #1 of 30 with a score of 8.0. Evidently AI is second and Vellum third.

How many of these have a free plan?

20 of the 25 on this page publish a free plan on their own pricing pages.

Which is the cheapest paid option?

On this page, Opik has the lowest first paid tier we found: $19/mo.

How is this list ranked?

Ranked on what each maker publishes, the fullest spec sheet first: how deeply the product is documented, the platforms it runs on, a free tier or trial to test it, and its standing.

More in AI Tools

All AI tools lists