Open Source Generative AI Projects
Generative AI uses models to create new content, such as text, images, audio, video, and code, in response to prompts or other inputs. It can help with tasks that involve drafting, summarizing, search, design, and content transformation. Building these systems also raises practical challenges, including model selection, data handling, evaluation, inference costs, and integration with existing applications.
Open source tools in this area include model training and fine-tuning frameworks, inference servers, application libraries, evaluation and monitoring utilities, and creative interfaces. When choosing a tool, consider its license, maintenance activity, documentation, hardware and software requirements, supported models, and compatibility with your workflow. These tools are useful to developers, researchers, creators, and organizations that want to build, adapt, or operate generative AI systems with greater control.
67 repositories · updated October 3, 2026

VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models
VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

dify-official-plugins: Extend Dify with Models and Tools
Official plugins for Dify add models, tools, agent strategies, and HTTP webhook extensions to the platform. Use this repository when building or managing Dify applications that need these integrations.

awesome-ai-agents: A Curated List of AI Agent Resources
awesome-ai-agents is a comprehensive, curated list of resources for building and understanding AI agents. It covers frameworks, tools, platforms, research papers, and more, making it an essential guide for anyone exploring the rapidly evolving field of autonomous AI systems.

awesome-ai: A Curated List of 400+ AI APIs, Tools, and Frameworks
The awesome-ai repository by edwardtay offers a comprehensive, curated list of over 400 AI APIs, tools, frameworks, and platforms. Spanning more than 40 categories, it serves as an invaluable resource for developers and researchers navigating the vast landscape of artificial intelligence. This list helps users discover solutions for LLMs, agents, image/video generation, MLOps, and more.

NVIDIA NeMo Speech: Scalable Generative AI for Speech Models
NVIDIA NeMo Speech is a powerful, scalable generative AI framework designed for researchers and developers focused on Large Language Models, Multimodal, and Speech AI. It provides tools for Automatic Speech Recognition (ASR) and Text-to-Speech (TTS), enabling efficient creation, customization, and deployment of new AI models using existing code and pre-trained checkpoints. This framework supports a wide range of applications, from real-time streaming ASR to high-quality multilingual TTS.

rag-zero-to-hero-guide: Learn Retrieval-Augmented Generation
A learning guide to retrieval-augmented generation, from core concepts to evaluation and advanced approaches. It combines explanations, Jupyter notebook implementations, tool references, and survey papers for learners building RAG knowledge.

mergoo: Combine Fine-Tuned LLM Experts into Routed Models
Mergoo combines fine-tuned language models or LoRA adapters into routed expert models, then supports training the resulting model. It is aimed at teams that want one model to draw on specialized experts rather than use them separately.

app.yumcut.com: Generate Short Videos from Prompts
YumCut turns a prompt into a vertical short video with generated scripts, voice, visuals, and captions. It is aimed at creators and teams who want to automate short-form production and control their deployment and provider choices.

lamini: Access the Lamini API from Python
Lamini is a Python client and SDK for working with Lamini's generative AI API. It suits developers who want to connect Python applications to Lamini without building API requests from scratch.

xTuring: Fine-Tune and Run Personalized LLMs
xTuring is a Python library for preparing data, fine-tuning, evaluating, and running open-source language models locally or in a private cloud. It is aimed at developers who want model customization through a high-level API and methods such as LoRA and quantization.

torchtune: Fine-Tune and Post-Train Large Language Models
torchtune is a PyTorch library for configuring and running LLM post-training workflows, from supervised fine-tuning to preference optimization. It is aimed at developers who want editable recipes and model-specific configs, but is no longer actively maintained.

TensorRT-LLM: Optimize LLM Inference on NVIDIA GPUs
TensorRT-LLM is a Python framework and runtime for efficient LLM and visual-generation inference on NVIDIA GPUs. It suits teams deploying models on NVIDIA hardware that need optimized kernels and configurable single- or distributed-GPU execution.