LoRA (Low-Rank Adaptation)
LoRA, or low-rank adaptation, is a method for adapting pretrained neural networks without updating all their parameters. It keeps the original model weights frozen and trains small, additional matrices that represent changes for a particular task. This can reduce the memory, storage, and compute needed for fine-tuning, making it practical to customize large language, image, and speech models with more limited resources. The resulting adapters can often be stored and shared separately from the base model.
Open source tools in this area include libraries for training adapters, workflows for managing or merging them, and utilities for loading them during inference. When choosing a tool, check support for your model architecture and training setup, hardware and dependency requirements, license, maintenance activity, and compatibility with your deployment workflow. LoRA is useful to researchers, developers, and organizations that want to adapt pretrained models without maintaining a full separately fine-tuned copy for every task.
1 repository · updated July 6, 2026
