LLaMA-Factory: Unified Efficient Fine-Tuning for 100+ LLMs & VLMs
This repository profile is provided by osrepos.com, an open source repository discovery platform.

Summary
LLaMA-Factory is an open-source project offering a unified and efficient framework for fine-tuning over 100 large language models (LLMs) and vision-language models (VLMs). Recognized at ACL 2024, it provides a comprehensive suite of tools and algorithms for various training approaches. This repository simplifies the complex process of adapting powerful models for specific tasks with ease and scalability.
Repository Information
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Introduction
LLaMA-Factory, developed by hiyouga, is a highly popular and robust framework designed for the unified and efficient fine-tuning of a vast array of large language models (LLMs) and vision-language models (VLMs). With over 62,000 stars and 7,500 forks on GitHub, it stands out as a go-to solution for researchers and developers in the AI community. The project, written primarily in Python and licensed under Apache-2.0, was recognized at ACL 2024 for its significant contributions to the field of efficient model adaptation.
Installation
Getting started with LLaMA-Factory is straightforward. You can install it directly from the source or use a pre-built Docker image.
To install from source, clone the repository and install the necessary dependencies:
git clone --depth 1 https://github.com/hiyouga/LLaMA-Factory.git
cd LLaMA-Factory
pip install -e ".[torch,metrics]" --no-build-isolation
For users preferring Docker, a pre-built image is available, simplifying environment setup:
docker run -it --rm --gpus=all --ipc=host hiyouga/llamafactory:latest
Examples
LLaMA-Factory provides intuitive command-line interface (CLI) commands for common tasks such as fine-tuning, inference, and model merging. Here are quickstart examples for the Llama3-8B-Instruct model:
To perform LoRA fine-tuning:
llamafactory-cli train examples/train_lora/llama3_lora_sft.yaml
To run inference with the fine-tuned model:
llamafactory-cli chat examples/inference/llama3_lora_sft.yaml
To merge the LoRA adapters back into the base model:
llamafactory-cli export examples/merge_lora/llama3_lora_sft.yaml
Additionally, LLaMA-Factory offers a user-friendly Web UI for fine-tuning models in your browser:
llamafactory-cli webui
Why Use LLaMA-Factory
LLaMA-Factory is a powerful tool for anyone working with large language models, offering a wide range of features and benefits:
- Extensive Model Support: It supports over 100 models, including popular ones like LLaMA, LLaVA, Mistral, Mixtral-MoE, Qwen, DeepSeek, Yi, and Gemma, ensuring compatibility with the latest advancements.
- Diverse Training Approaches: The framework integrates various methods such as supervised fine-tuning (SFT), reward modeling, PPO, DPO, KTO, and ORPO, catering to different training paradigms.
- Scalable and Efficient Tuning: It supports 16-bit full-tuning, freeze-tuning, LoRA, and 2/3/4/5/6/8-bit QLoRA via multiple quantization techniques, allowing for efficient training even on limited hardware.
- Advanced Algorithms and Tricks: LLaMA-Factory incorporates cutting-edge algorithms like GaLore, BAdam, APOLLO, DoRA, LongLoRA, and PiSSA, alongside practical tricks such as FlashAttention-2, Unsloth, and RoPE scaling for enhanced performance.
- Comprehensive Experiment Monitoring: It integrates with popular experiment monitors like LlamaBoard, TensorBoard, Wandb, and SwanLab, providing robust tracking and visualization capabilities.
- Faster Inference: The platform offers faster inference through an OpenAI-style API, Gradio UI, and CLI, leveraging backends like vLLM and SGLang for high-throughput deployments.
Links
Explore LLaMA-Factory further through these official resources:
- GitHub Repository: hiyouga/LLaMA-Factory
- Official Documentation: LLaMA-Factory Docs
- Colab Notebook: Open in Colab
- Discord Community: Join Discord
- Twitter: Follow @llamafactory_ai
Related repositories
Similar repositories that may be relevant next.
SwarmLLM: Run Local AI Models and Team Up for Giant Distributed Inference
October 2, 2026
SwarmLLM is a free, open-source application that allows you to run AI chat models directly on your own computer. It uniquely enables multiple computers to team up over the internet, collectively running models too large for a single machine. This platform offers an OpenAI and Anthropic-compatible API, all without requiring accounts or cryptocurrency.
Memoh: An Open-Source Multi-Agent Platform with Dedicated AI Workspaces
September 26, 2026
Memoh is an innovative open-source multi-agent platform designed to provide each AI agent with its own dedicated cloud computer. This includes a filesystem, desktop, browser, network, and persistent long-term memory, ensuring agents remain online 24/7. Users can integrate their own API keys or host existing AI models, fostering a versatile and always-on environment for AI development and deployment.
Graphon: A Python Graph Execution Engine for Agentic AI Workflows
September 26, 2026
Graphon is an innovative Python-based graph execution engine designed for building agentic AI workflows. It provides a robust framework for orchestrating complex AI tasks, featuring event-driven execution, graph validation, and shared runtime state. This evolving repository already includes a functional engine, built-in nodes, and end-to-end examples for developers.

llm-d Router: Intelligent Orchestration for Multi-Phase LLM Inference
September 25, 2026
The `llm-d-router` is a sophisticated Go service designed as an intelligent entry point for large language model (LLM) inference requests. It orchestrates complex, multi-phase LLM inference pipelines across specialized worker pools, routing requests through an Inference Gateway to disaggregated vLLM workers. By exposing OpenAI-compatible APIs, it simplifies integration and leverages Kubernetes for scalable, efficient deployments.
Source repository
Open the original repository on GitHub.
30 counted GitHub visits