A parameter-efficient fine-tuning method that adds small trainable rank-decomposed matrices to a frozen pre-trained model — drastically reduces fine-tuning compute and storage vs full fine-tuning.
LoRA freezes the pre-trained model's weights and adds two small trainable matrices (A and B) to each target layer whose product approximates the weight update full fine-tuning would have applied. Typical rank is 8-64, reducing trainable parameters by 1000-10000×. After training, the LoRA weights can be merged back into the base model or kept separate (enabling multi-adapter switching). LoRA + QLoRA (quantized LoRA) make fine-tuning 70B-parameter models feasible on a single consumer GPU.
Fine-tuning Llama 3 70B on a domain-specific dataset with LoRA — uses 24GB of GPU RAM instead of the 280GB+ that full fine-tuning would require.
LoRA democratized LLM fine-tuning — any team with one good GPU can adapt large models to their domain instead of paying for full-model training.
Need help implementing this in your business?
Get Started