Benchmark Radar: A Living Database for AI Benchmarks and Evaluation
This repository profile is provided by osrepos.com, an open source repository discovery platform.

Summary
Benchmark Radar is an extensive open-source project that tracks over 20,710 AI benchmark, evaluation, dataset, and data-quality records from 37 public sources. It provides daily updates, linked evidence, and tools for researchers and developers to discover and analyze AI benchmarks. This project is essential for anyone needing to stay current with AI evaluation trends and model performance.
Repository Information
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Introduction
Benchmark Radar, developed by ktiwu01, is a comprehensive open-source initiative designed to centralize and track AI benchmarks and evaluation data. It continuously collects over 20,710 records related to AI benchmarks, evaluations, datasets, and data quality from 37 diverse public sources. With daily updates and linked evidence, it serves as a living database for the rapidly evolving field of AI evaluation, including LLM evaluation, agentic benchmarks, coding, reasoning, and safety assessments.
The project offers a web dashboard, a downloadable dataset, and a command-line interface (CLI), making it a versatile resource for researchers and evaluation engineers. It aims to help users quickly find relevant evaluations, understand reported scores, and track how model performance changes over time.
Why Use Benchmark Radar and Its Benefits
Benchmark Radar offers significant advantages for anyone involved in AI research and development:
- Comprehensive Coverage: It aggregates data from 37 public sources, providing a broad view of the AI benchmark landscape. This includes papers, repositories, datasets, and model reports.
- Daily Updates: The database is updated daily, ensuring you have access to the latest benchmark signals and evaluation results.
- Evidence-Based Tracking: Each record includes linked evidence, allowing users to inspect the original sources and understand the context of reported scores.
- Historical Performance Analysis: Users can track how model scores change over time for specific benchmarks, offering insights into progress and saturation.
- Searchable Catalog and Leaderboards: The platform provides a searchable catalog, daily findings, and a benchmark leaderboard, making it easy to discover new benchmarks and compare model performance.
- Offline Access: The complete dataset is downloadable, and a CLI tool allows for local queries, enabling offline analysis and integration into custom workflows.
- Community-Driven: It encourages contributions, allowing the community to add new benchmarks, model cards, and sources.
Installation
To query Benchmark Radar locally using its command-line interface (CLI), you can use npx:
npx skills add ktwu01/benchmark-radar
This command installs the necessary tool and downloads the data to your computer the first time you use it. For detailed setup and usage instructions, refer to the CLI setup and usage guide.
Examples
Benchmark Radar provides several ways to explore its data:
- Web Dashboard: Visit the official dashboard to see today's findings, benchmark trends, scores, and model coverage. The "Today" page offers a ranked feed of newly found benchmarks and a daily briefing with cited evidence.
- Benchmark Frontier: Explore the "Benchmark Frontier" to see which difficult benchmarks have been tested most, compare reported scores, release dates, and the number of models tested.
- Data Download: You can download the complete dataset, which includes the benchmark catalog, detail records, and daily discovery snapshots in a single ZIP file.
- RSS Feed: Stay updated with new benchmark signals daily by subscribing to the RSS feed.
Links
Here are some essential links for Benchmark Radar:
- GitHub Repository: https://github.com/ktwu01/benchmark-radar
- Official Dashboard: https://benchmark-radar.org/
- Download Complete Dataset: https://github.com/ktwu01/benchmark-radar/releases/download/cli-data/benchmark-radar-data.zip
- Hugging Face Dataset: https://huggingface.co/datasets/ktwu01/benchmark-radar
- arXiv Technical Report: https://arxiv.org/abs/2609.11115
- CLI Setup and Usage Guide: https://github.com/ktwu01/benchmark-radar/blob/main/skills/benchmark-radar/SKILL.md
- RSS Feed: https://benchmark-radar.org/feed.xml
- Contribution Guide: https://github.com/ktwu01/benchmark-radar/blob/main/CONTRIBUTING.md
Source repository
Open the original repository on GitHub.