An umbrella term for techniques that fine-tune large pre-trained models by updating only a small subset of parameters — LoRA, adapters, prefix tuning, and prompt tuning all qualify.
PEFT methods solve the cost and storage problems of full fine-tuning. Full fine-tuning a 70B-parameter model requires hundreds of GB of GPU RAM and produces a separate 70B-parameter checkpoint per fine-tuned model. PEFT methods update <1% of parameters, fitting on a single GPU and producing tiny adapter files (~50MB) that can be loaded on top of the base model at inference time. HuggingFace's `peft` library is the standard implementation.
Maintaining 20 different fine-tuned variants of a single base model as 20 small PEFT adapters (~50MB each) instead of 20 full 70GB model checkpoints.
PEFT is the practical path to multi-tenant or multi-domain LLM deployments — train once, ship many cheap adapters tailored per use case.
Need help implementing this in your business?
Get Started