Vendor website benchmark · category
LLM observability and evaluation vendors, measured on their own websites
Tracing, evaluation and prompt management for language models. Website measurement only.
Listed
13
Currently measurable
13
Median website score
74
Categories are editorial and unscored; membership never moves a score. Ungraded vendors sort after measured ones and are never coerced to zero. Score version 25.
| Vendor | Website score | Reach | Read | Quote | Findings | llms.txt | Entity map | Measured |
|---|---|---|---|---|---|---|---|---|
| Arize AI arize.com | 84 C | 88 | 70 | 91 | 2 | present | absent | 2026-09-18 00:18Z |
| LiteLLM litellm.ai | 84 B | 100 | 80 | 67 | 1 | absent | absent | 2026-09-18 03:05Z |
| Braintrust braintrust.dev | 83 C | 93 | 75 | 77 | 3 | present | absent | 2026-09-18 02:39Z |
| PromptLayer promptlayer.com | 83 C | 89 | 75 | 84 | 3 | present | absent | 2026-09-19 02:17Z |
| Fiddler AI fiddler.ai | 79 B | 93 | 70 | 69 | 0 | present | absent | 2026-09-18 02:57Z |
| Vellum vellum.ai | 76 C | 92 | 65 | 66 | 3 | present | absent | 2026-09-18 03:19Z |
| Patronus AI patronus.ai | 74 C | 96 | 67 | 53 | 5 | absent | absent | 2026-09-18 22:20Z |
| Helicone helicone.ai | 72 C | 85 | 80 | 47 | 3 | present | absent | 2026-09-18 03:01Z |
| Portkey portkey.ai | 71 C | 93 | 70 | 44 | 3 | present | absent | 2026-09-19 01:17Z |
| Humanloop humanloop.com | 70 C | 95 | 75 | 33 | 1 | present | absent | 2026-09-18 03:02Z |
| OpenRouter openrouter.ai | 70 D | 80 | 45 | 80 | 2 | present | absent | 2026-09-18 19:20Z |
| Langfuse langfuse.com | 68 C | 93 | 60 | 43 | 2 | present | absent | 2026-09-18 03:04Z |
| WhyLabs whylabs.ai | 57 D | 67 | 80 | 20 | 6 | absent | absent | 2026-09-18 03:28Z |
Compare within this category
Open any two to four of these side by side: /directory/compare/arize/litellm — slugs in any order resolve to one canonical page.