Open Source Datasets
Datasets are organized collections of information used to analyze patterns, train and evaluate machine learning models, and support research or software development. They can include text, images, audio, tabular records, or other formats. Well-documented, reliable data helps teams reproduce results, compare methods, and build systems suited to real-world needs while reducing the effort of collecting and preparing information from scratch.
Open source dataset resources include curated collections, data-generation and annotation tools, cleaning and validation utilities, and benchmark suites. When choosing one, check its license and permitted uses, data provenance, documentation, update and maintenance activity, format, and compatibility with your workflow. Researchers, developers, analysts, educators, and organizations can use these resources to explore questions, develop models, or test systems, while accounting for privacy, representation, and data quality.
2 repositories · updated September 29, 2026

Benchmark Radar: A Living Database for AI Benchmarks and Evaluation
Benchmark Radar is an extensive open-source project that tracks over 20,710 AI benchmark, evaluation, dataset, and data-quality records from 37 public sources. It provides daily updates, linked evidence, and tools for researchers and developers to discover and analyze AI benchmarks. This project is essential for anyone needing to stay current with AI evaluation trends and model performance.

txtinstruct: Building Instruction-Tuned Models with Custom Data
txtinstruct is a Python framework designed for training instruction-tuned models. It focuses on supporting open data and models, enabling users to build their own instruction-following datasets and train models without licensing ambiguity. This project simplifies the process of creating custom instruction-tuned solutions.