Open Source Large Language Models
Large language models are machine learning systems trained on large amounts of text to understand and generate language. They can answer questions, summarize documents, translate, extract information, and help produce code. Open source work in this area supports adapting models to specialized tasks, running them on different hardware, and inspecting or evaluating their behavior. These capabilities can help automate language-intensive work, while requiring attention to accuracy, privacy, and potential bias.
Open source tools include model weights, training and fine-tuning frameworks, inference engines, evaluation suites, and utilities for routing, security, and deployment. When choosing one, check its license and usage terms, maintenance activity, hardware and software requirements, documentation, and compatibility with your existing systems. These tools are useful to researchers, developers, organizations, and individuals building or studying language-based applications.
8 repositories · updated October 3, 2026

AWESOME-OCR-LLM: Track OCR Research in the LLM Era
AWESOME-OCR-LLM is a curated reading list for OCR research shaped by large language and vision-language models. It helps researchers and practitioners follow work across document parsing, understanding, visual text generation, and evaluation.

colibri: Run Large MoE Models on Local Hardware
colibri is a C inference engine that streams Mixture-of-Experts model weights from disk, so large models can run without fitting entirely in RAM or VRAM. It supports CPU-only use and optional GPU backends, with speed depending on hardware and storage.

RL4LMs: Fine-Tune Language Models with Reinforcement Learning
RL4LMs is a Python library for training language models against custom reward functions using on-policy reinforcement learning. It suits NLP researchers and developers who need configurable training components for text-generation tasks.

evalplus: Rigorously Evaluate LLM-Generated Code
EvalPlus evaluates code generated by language models with expanded correctness tests for HumanEval and MBPP, plus efficiency checks through EvalPerf. It is for researchers and developers comparing models or validating generated code more rigorously.

xgrammar: Constrain Language Model Output to Structured Formats
XGrammar is a library for constrained decoding that helps language models produce outputs matching JSON, regular expressions, or context-free grammars. It suits teams integrating structured output into LLM inference and applications that need reliable machine-readable responses.

GLM-5: Power Long-Horizon Coding and Agentic Tasks
GLM-5 is a family of large language models from Z.ai for complex coding, systems engineering, and long-running agent tasks. The repository provides model downloads, serving guidance, and fine-tuning references.

litgpt: Train, Fine-Tune, and Deploy Large Language Models
LitGPT provides implementations and workflows for pretraining, fine-tuning, evaluating, and serving a range of large language models. It suits developers and researchers who want configurable training recipes and direct control over model code.

FinGPT: Adapt Language Models for Financial Tasks
FinGPT provides financial datasets, fine-tuned language models, benchmarks, and workflows for tasks such as sentiment analysis and forecasting. It suits researchers and developers adapting open models to finance, with local GPU inference or supported cloud APIs.