Open Source LLM Projects
Discover 274 open source LLM repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. LLM projects here are most often combined with Python, AI Agents and AI. Last updated October 4, 2026.
274 repositories · updated October 4, 2026

open-r1: Reproduce DeepSeek-R1 Training and Evaluation
Open R1 is Hugging Face’s toolkit and research project for reproducing the DeepSeek-R1 pipeline with open datasets, training scripts, and evaluation workflows. It is no longer maintained; its training work has moved to TRL.

qwen-code: Use AI Agents to Work on Code
Qwen Code is an AI coding agent for working with repositories through a terminal and other interfaces. It supports multiple model providers and can also be used from editors, desktop, web, chat, scripts, or SDKs.

toon: Compact JSON Data for LLM Prompts
TOON is a lossless, human-readable encoding of JSON designed to reduce prompt token use, especially for uniform records. Its TypeScript SDK and CLI convert between JSON and TOON, with format-specific trade-offs documented in the repository.

llama-cpp-python: Run llama.cpp Models from Python
Python bindings for llama.cpp let developers run GGUF language models locally through a high-level API, low-level C bindings, or an OpenAI-compatible server. Useful when you want local inference and control over hardware backends.

FuncVul: Detect Vulnerabilities in Code Functions
FuncVul is a research model for detecting vulnerable code chunks within C/C++ and Python functions. It uses fine-tuned GraphCodeBERT and provides six labeled datasets for evaluating function-level vulnerability detection.

instructor: Extract Validated Structured Data from LLMs
Instructor turns LLM responses into validated, typed Python objects using Pydantic models. It is useful for applications that need dependable extraction across providers, with retries and streaming handled through a consistent API.

LlamaFactory: Fine-Tune Large Language and Vision Models
LlamaFactory provides CLI and web interfaces for fine-tuning a broad range of language and vision models. It supports parameter-efficient methods and preference training, with workflows for training, inference, and model export.

text-generation-inference: Serve Large Language Models
A toolkit for serving large language models through text-generation APIs. It provides GPU-oriented inference features such as continuous batching, token streaming, tensor parallelism, and quantization, but is now in maintenance mode.

weave: Trace and Evaluate Generative AI Applications
Weave helps developers trace language model inputs, outputs, and function calls, then organize and evaluate application runs. It suits teams building AI features that need debugging and repeatable comparisons across experiments and production.

deepchat: Run AI Agents in a Desktop Assistant
DeepChat is a cross-platform desktop client for chatting with cloud and local AI models and running agents. It suits people who want tools, skills, and resumable sessions in one local-first workspace.

data-prep-kit: Prepare Data for LLM Applications
Data-Prep-Kit is a toolkit for cleaning, transforming, and enriching unstructured data used in LLM training and RAG pipelines. It offers reusable transforms that run with Python or Ray, from local experiments to larger-scale processing.

rag-web-ui: Build Knowledge-Base Q&A with RAG
RAG Web UI is a self-hostable web application for creating question-answering systems over your own documents. It combines document ingestion, vector search, and configurable cloud or local language models, with a web interface and OpenAPI access.