Tech riderRev. 4 Oct 2026
LM Evaluation Harness
- 1Runs onAPI, Linux, self-hosted
- 2CostsFree plan
- 3Deploymentself-hosted
3 lines stated Written from the maker's own pages: github.com

Overview
LM Evaluation Harness is ranked #21 of 30 in AI LLM evaluation tools on Specifiction. It runs on API, Linux, Self-hosted. There is a free plan.
LM Evaluation Harness plans and pricing
All plansLM Evaluation Harness Free MIT-licensed software · model backends and task dependencies may require separate optional installs github.com · 4 Oct 2026
Compared on AI LLM evaluation tools
- Deployment
- self-hostedgithub.com
Facts
- Purpose
- LM Evaluation Harness is a unified framework for evaluating generative language models across evaluation tasks.github.com · 4 Oct 2026
- Benchmarks
- The project lists over 60 standard academic LLM benchmarks with hundreds of subtasks and variants.github.com · 4 Oct 2026
- Model support
- Supported model options include Hugging Face Transformers, GPT-NeoX, Megatron-DeepSpeed, and vLLM.github.com · 4 Oct 2026
- API integrations
- The harness supports hosted APIs including OpenAI and Anthropic, plus local servers compatible with OpenAI APIs.github.com · 4 Oct 2026
- Custom evaluation
- Users can define custom prompts and evaluation metrics, and can configure tasks with YAML.github.com · 4 Oct 2026
- Plugins
- Model backends, filters, metrics, and aggregations can be registered from external packages through entry points or loaded from a local module.github.com · 4 Oct 2026
- Reproducibility
- The project says its use of publicly available prompts supports reproducibility and comparability between papers.github.com · 4 Oct 2026
- Installation
- The base package provides the core framework while model backends are installed separately as optional extras.github.com · 4 Oct 2026
- Supported platforms
- Documented runtime options include Windows ML and ONNX Runtime providers for CPU, CUDA, DirectML, NPU, ROCm, and other accelerators.github.com · 4 Oct 2026
- Security and license
- The repository provides the software under the MIT License, which states that it is provided “AS IS” without warranty.github.com · 4 Oct 2026
- Support
- The project directs users to open a GitHub issue or join the EleutherAI Discord for support.github.com · 4 Oct 2026
- Notable limitation
- For model APIs that do not provide logits or log probabilities, the documentation limits use to generate-until tasks.github.com · 4 Oct 2026
- Users
- The project says the harness serves as the backend for Hugging Face’s Open LLM Leaderboard and is used by research papers and organizations.github.com · 4 Oct 2026
Best LM Evaluation Harness alternatives
See all 20 All accessCh 01 Weights & Biases Free planFree trialAPI from $60/mo8.0 All accessCh 02 Evidently AI Free planFree trialAPI from $80/mo7.8 All accessCh 03 Vellum Free planAndroidBrowser from $30/mo7.7 All accessCh 04 Opik Free planFree trialAPI from $19/mo7.5 All accessCh 05 Promptfoo Free planAPILinux Free to start7.4 All accessCh 06 Maxim AI Free planFree trialAPI from $29/mo7.3
Where it ranks on Specifiction
- Best AI LLM Evaluation Tools in 2026#21 of 30
- Best LLM Evaluation Tools in 2026#11 of 29
Is LM Evaluation Harness yours?
Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.
Sources
- github.com/EleutherAI/lm-evaluation-harness· checked 4 Oct 2026
- github.com/EleutherAI/lm-evaluation-harness/blob/m· checked 4 Oct 2026




