{"name":"Benchmark Radar: A Living Database for AI Benchmarks and Evaluation","description":"Benchmark Radar is an extensive open-source project that tracks over 20,710 AI benchmark, evaluation, dataset, and data-quality records from 37 public sources. It provides daily updates, linked evidence, and tools for researchers and developers to discover and analyze AI benchmarks. This project is essential for anyone needing to stay current with AI evaluation trends and model performance.","github":"https://github.com/ktwu01/benchmark-radar","url":"https://osrepos.com/repo/ktwu01-benchmark-radar","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/ktwu01-benchmark-radar","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/ktwu01-benchmark-radar.md","json":"https://osrepos.com/repo/ktwu01-benchmark-radar.json","topics":["AI Benchmark","LLM Evaluation","Agentic Benchmarking","Datasets","Python","Evaluation Framework","Model Cards","AI Research"],"keywords":["AI Benchmark","LLM Evaluation","Agentic Benchmarking","Datasets","Python","Evaluation Framework","Model Cards","AI Research"],"stars":null,"summary":"Benchmark Radar is an extensive open-source project that tracks over 20,710 AI benchmark, evaluation, dataset, and data-quality records from 37 public sources. It provides daily updates, linked evidence, and tools for researchers and developers to discover and analyze AI benchmarks. This project is essential for anyone needing to stay current with AI evaluation trends and model performance.","content":"## Introduction\n\nBenchmark Radar, developed by ktiwu01, is a comprehensive open-source initiative designed to centralize and track AI benchmarks and evaluation data. It continuously collects over 20,710 records related to AI benchmarks, evaluations, datasets, and data quality from 37 diverse public sources. With daily updates and linked evidence, it serves as a living database for the rapidly evolving field of AI evaluation, including LLM evaluation, agentic benchmarks, coding, reasoning, and safety assessments.\n\nThe project offers a web dashboard, a downloadable dataset, and a command-line interface (CLI), making it a versatile resource for researchers and evaluation engineers. It aims to help users quickly find relevant evaluations, understand reported scores, and track how model performance changes over time.\n\n## Why Use Benchmark Radar and Its Benefits\n\nBenchmark Radar offers significant advantages for anyone involved in AI research and development:\n\n*   **Comprehensive Coverage**: It aggregates data from 37 public sources, providing a broad view of the AI benchmark landscape. This includes papers, repositories, datasets, and model reports.\n*   **Daily Updates**: The database is updated daily, ensuring you have access to the latest benchmark signals and evaluation results.\n*   **Evidence-Based Tracking**: Each record includes linked evidence, allowing users to inspect the original sources and understand the context of reported scores.\n*   **Historical Performance Analysis**: Users can track how model scores change over time for specific benchmarks, offering insights into progress and saturation.\n*   **Searchable Catalog and Leaderboards**: The platform provides a searchable catalog, daily findings, and a benchmark leaderboard, making it easy to discover new benchmarks and compare model performance.\n*   **Offline Access**: The complete dataset is downloadable, and a CLI tool allows for local queries, enabling offline analysis and integration into custom workflows.\n*   **Community-Driven**: It encourages contributions, allowing the community to add new benchmarks, model cards, and sources.\n\n## Installation\n\nTo query Benchmark Radar locally using its command-line interface (CLI), you can use `npx`:\n\nbash\nnpx skills add ktwu01/benchmark-radar\n\n\nThis command installs the necessary tool and downloads the data to your computer the first time you use it. For detailed setup and usage instructions, refer to the [CLI setup and usage guide](https://github.com/ktwu01/benchmark-radar/blob/main/skills/benchmark-radar/SKILL.md \"CLI setup and usage guide\" target=\"_blank\").\n\n## Examples\n\nBenchmark Radar provides several ways to explore its data:\n\n*   **Web Dashboard**: Visit the [official dashboard](https://benchmark-radar.org/ \"Benchmark Radar Dashboard\" target=\"_blank\") to see today's findings, benchmark trends, scores, and model coverage. The \"Today\" page offers a ranked feed of newly found benchmarks and a daily briefing with cited evidence.\n*   **Benchmark Frontier**: Explore the \"Benchmark Frontier\" to see which difficult benchmarks have been tested most, compare reported scores, release dates, and the number of models tested.\n*   **Data Download**: You can download the [complete dataset](https://github.com/ktwu01/benchmark-radar/releases/download/cli-data/benchmark-radar-data.zip \"Download complete dataset\" target=\"_blank\"), which includes the benchmark catalog, detail records, and daily discovery snapshots in a single ZIP file.\n*   **RSS Feed**: Stay updated with new benchmark signals daily by subscribing to the [RSS feed](https://benchmark-radar.org/feed.xml \"Benchmark Radar RSS Feed\" target=\"_blank\").\n\n## Links\n\nHere are some essential links for Benchmark Radar:\n\n*   **GitHub Repository**: [https://github.com/ktwu01/benchmark-radar](https://github.com/ktwu01/benchmark-radar \"GitHub Repository\" target=\"_blank\")\n*   **Official Dashboard**: [https://benchmark-radar.org/](https://benchmark-radar.org/ \"Benchmark Radar Dashboard\" target=\"_blank\")\n*   **Download Complete Dataset**: [https://github.com/ktwu01/benchmark-radar/releases/download/cli-data/benchmark-radar-data.zip](https://github.com/ktwu01/benchmark-radar/releases/download/cli-data/benchmark-radar-data.zip \"Download complete dataset\" target=\"_blank\")\n*   **Hugging Face Dataset**: [https://huggingface.co/datasets/ktwu01/benchmark-radar](https://huggingface.co/datasets/ktwu01/benchmark-radar \"Hugging Face Dataset\" target=\"_blank\")\n*   **arXiv Technical Report**: [https://arxiv.org/abs/2609.11115](https://arxiv.org/abs/2609.11115 \"arXiv Technical Report\" target=\"_blank\")\n*   **CLI Setup and Usage Guide**: [https://github.com/ktwu01/benchmark-radar/blob/main/skills/benchmark-radar/SKILL.md](https://github.com/ktwu01/benchmark-radar/blob/main/skills/benchmark-radar/SKILL.md \"CLI Setup and Usage Guide\" target=\"_blank\")\n*   **RSS Feed**: [https://benchmark-radar.org/feed.xml](https://benchmark-radar.org/feed.xml \"Benchmark Radar RSS Feed\" target=\"_blank\")\n*   **Contribution Guide**: [https://github.com/ktwu01/benchmark-radar/blob/main/CONTRIBUTING.md](https://github.com/ktwu01/benchmark-radar/blob/main/CONTRIBUTING.md \"Contribution Guide\" target=\"_blank\")","metrics":{"detailViews":1,"githubClicks":0},"dates":{"published":null,"modified":"2026-09-29T11:56:18.000Z"}}