Tech riderRev. 20 Sept 2026
DecodingTrust
- 1Runs onAPI, self-hosted
- 2CostsNot stated by the maker
- 3Deployment optionsself-hosted
- 4Safety evaluationsYes
3 lines stated Written from the maker's own pages: decodingtrust.github.io

Overview
DecodingTrust is ranked #17 of 29 in LLM evaluation tools on Specifiction. It runs on API, Self-hosted.
Compared on LLM evaluation tools
- Deployment options
- self-hosteddecodingtrust.github.io
- Safety evaluations
- Yesdecodingtrust.github.io
Facts
- Purpose
- DecodingTrust is a research project for assessing trustworthiness in GPT models and helping researchers and practitioners understand LLM capabilities, limitations, and deployment risks.decodingtrust.github.io · 4 Oct 2026
- Evaluation areas
- The benchmark covers toxicity, stereotype and bias, adversarial robustness, out-of-distribution robustness, privacy, adversarial demonstrations, machine ethics, and fairness.decodingtrust.github.io · 4 Oct 2026
- Models
- The project says its evaluations mainly focus on GPT-3.5 and GPT-4, and it also supports causal LLMs hosted on Hugging Face or locally.github.com · 4 Oct 2026
- Resources
- The project provides a dataset and evaluation scripts organized by trustworthiness area.decodingtrust.github.io · 4 Oct 2026
- Reproducibility
- The benchmark uses timestamped GPT-3.5 and GPT-4 model versions to support consistent results and reproducibility.github.com · 4 Oct 2026
- Installation
- The project recommends cloning the repository and installing it in editable mode with pip so the data, code, and configurations remain together.github.com · 4 Oct 2026
- Supported architecture
- The repository says it supports the ppc64le architecture on IBM Power-9 platforms.github.com · 4 Oct 2026
- License
- The dataset and project are distributed under the CC BY-SA 4.0 license.decodingtrust.github.io · 4 Oct 2026
- Content warning
- The project warns that its data contains model outputs that may be considered offensive.decodingtrust.github.io · 4 Oct 2026
- Model coverage limit
- The repository says its benchmark mainly focuses on GPT-3.5-turbo-0301 and GPT-4-0314 for consistent conclusions and results.github.com · 4 Oct 2026
- Support
- Questions and suggestions can be sent by GitHub issue or pull request, or by email to [email protected].github.com · 4 Oct 2026
- Intended users
- The project describes its resources as intended to help researchers and practitioners assess LLM capabilities, limitations, and risks.decodingtrust.github.io · 4 Oct 2026
Best DecodingTrust alternatives
See all 20 All accessCh 01 Promptfoo Free planAPILinux Free to start7.5 All accessCh 02 DeepEval Free planLinuxMac Free to start7.4 All accessCh 03 Maxim AI Free planFree trialAPI from $29/mo7.4 All accessCh 04 Giskard Free planAPILinux Free to start7.2 All accessCh 05 Braintrust Free planAPIself-hosted from $249/mo7.1 All accessCh 06 Galileo Free planAPIself-hosted from $100/mo7.1
Where it ranks on Specifiction
- Best LLM Evaluation Tools in 2026#17 of 29
Is DecodingTrust yours?
Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.
Sources
- decodingtrust.github.io· checked 4 Oct 2026
- github.com/AI-secure/DecodingTrust· checked 4 Oct 2026






