Open Source Reinforcement Learning Projects
Reinforcement learning is a machine learning approach in which an agent learns to make decisions by interacting with an environment. It uses feedback, often expressed as rewards, to improve a policy over time. This approach is useful for problems involving sequences of actions and changing conditions, such as robotics, resource allocation, game playing, and adaptive control. It can help optimize decisions where outcomes depend on both current choices and their longer-term effects.
Open source tools in this area include algorithm libraries, simulation environments, training frameworks, benchmarks, and utilities for evaluation or visualization. When choosing one, consider its maintenance activity, license, documentation, hardware and software requirements, supported algorithms, and compatibility with your data and existing workflows. Researchers, students, and practitioners can use these tools to study learning methods, run experiments, and develop agents for practical tasks.
5 repositories · updated October 3, 2026

VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models
VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

Uni-Agent: A Scalable Framework for Training Long-Horizon AI Agents
Uni-Agent is a powerful Python framework designed for training long-horizon agents at scale. It allows users to integrate existing agent harnesses, unify diverse agent tasks through an extensible interface, and run thousands of sessions concurrently for efficient data collection and training.

EnvHarness: Dynamically Adapting Environments for Agent Learning
EnvHarness empowers large language models (LLMs) acting as autonomous agents to learn more effectively from interactive environments. It achieves this by wrapping static environments with plug-in components, making them dynamically controllable without altering their internal code. This innovative approach allows environments to target specific agent weaknesses and continuously teach as agents improve, leading to more effective and efficient learning outcomes.

RL4LMs: Fine-Tune Language Models with Reinforcement Learning
RL4LMs is a Python library for training language models against custom reward functions using on-policy reinforcement learning. It suits NLP researchers and developers who need configurable training components for text-generation tasks.

verifiers: Build Environments for LLM Training and Evaluation
verifiers is a Python library for building environments to train and evaluate large language models. It fits teams developing LLM reinforcement-learning workflows or reusable evaluations, especially those using Prime Intellect's training tools.