Qwen: Run and Fine-Tune Pretrained Language Models

Summary
Qwen is Alibaba Cloud’s Python repository for running and adapting pretrained and chat language models, with an emphasis on Chinese and English. It includes inference, quantization, fine-tuning, and deployment paths, but is no longer actively maintained; the project points users to Qwen2.
At a glance
- Language
- Python
- License
- Apache-2.0
- Stars
- 21.9k
- Forks
- 2k
- Added to OSRepos
- April 30, 2026
- Last analyzed
- October 4, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Qwen provides code and guidance for using the Qwen family of pretrained language models and chat models. It addresses the practical work around loading checkpoints, generating responses, adapting models to downstream tasks, and serving them through local interfaces or APIs.
The repository is most relevant to developers and researchers who want to work directly with these model releases and their tooling. The README says this repository is no longer actively maintained because of substantial codebase differences, and directs users to Qwen2 for the newer project.
Key Features
- Supports base and chat model checkpoints in several model sizes, with links to Hugging Face and ModelScope.
- Provides inference examples using Transformers and ModelScope, including conversational and batch inference.
- Documents Int4 and Int8 GPTQ models, as well as quantization of the attention KV cache.
- Includes fine-tuning scripts for full-parameter training, LoRA, and Q-LoRA, with DeepSpeed and FSDP support described in the README.
- Offers deployment examples using vLLM, FastChat, a web demo, a CLI demo, and an OpenAI-style API.
- Covers integrations and routes for CPU inference, multiple GPUs, Docker, and Alibaba Cloud’s DashScope API.
Use Cases
- Researchers can run the published Qwen checkpoints to explore language-model behavior across Chinese and English tasks.
- Application developers can adapt chat models with LoRA or Q-LoRA when a task calls for specialized responses and they have suitable training hardware.
- Teams evaluating local inference can compare model sizes and quantization options against their GPU memory constraints.
- Developers building a prototype can use the included web, CLI, or API deployment examples to expose a model to users or other software.
Project Facts
- Language: Python
- License: Apache-2.0
- Stars: 21.9k
- Forks: 2k
- Topics: chinese, flash-attention, large-language-models, llm, natural-language-processing, pretrained-models
- Archived: No
Getting Started
The README specifies Python 3.8+, PyTorch 1.12+, and Transformers 4.32+ as requirements. Install the repository dependencies with:
pip install -r requirements.txt
Then follow the README for checkpoint selection, inference examples, and hardware-specific setup. The model checkpoints are distributed separately through Hugging Face and ModelScope.
Alternatives
- torchtune: torchtune focuses on configurable LLM post-training recipes, while Qwen also provides model inference, quantization, and deployment workflows.
Considerations
- The repository is no longer actively maintained. For the successor project, see Qwen2.
- Running larger checkpoints and fine-tuning can require substantial GPU memory. The README includes memory figures and quantized options, but actual needs depend on model size and configuration.
- The examples rely on compatible versions of Python, PyTorch, Transformers, CUDA, and optional acceleration or training packages. Checkpoint and dependency compatibility may require care.
- Model weights are hosted separately from this code repository. Review the relevant model cards and terms before choosing a checkpoint or deploying it.
Source repository
Open the original repository on GitHub.
13 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.

agent-anvil: Test AI Agent Tool Use in CI
October 1, 2026
Agent Anvil evaluates tool-using AI agents through scenario-based runs, trace checks, and optional semantic grading. It helps teams catch unsafe or incorrect tool behavior before it reaches production.