Evaluation

LLM Benchmarks & Evaluation Harnesses

Model rankings shift every few weeks. Rather than embed a scoreboard that goes stale, this page links straight to the leaderboards and harnesses the industry actually uses — plus a live feed of the latest model news.

Evaluation Harnesses

Leaderboards tell you where models rank. Harnesses are the actual tooling to run those evaluations yourself, against your own prompts and data.

Already in this map's Capability Directory: Braintrust, DeepEval, Patronus AI, and Ragas are the production-grade eval platforms referenced under AI Evaluation & Quality Assurance.

Latest Model News

Live-fetched on page load from the Hugging Face Blog RSS feed — one of the highest-signal public feeds for new model releases.

Loading latest items…

Fetched client-side via a public RSS-to-JSON proxy — no server on this site, so availability isn't guaranteed. If it's down, use the direct link shown in its place.