03 · ADAPTER
VerifiedLoRA and QLoRA
Adapt a frozen base with a low-rank update; measure memory fit on the target hardware.
- W frozen + Δ = W + BA
- Only A and B train
- QLoRA: base is 4-bit
- Adapters are reversible
KNOW THESE FIRST
The budget, window, and …LoRA trains only A and B adapters in W' = W + scale × BA. The default random A and zero B initialization preserves base behavior; other initialization methods need separate verification.
QLoRA stores the frozen base in 4-bit form while adapters and critical computation can use higher precision. It expands the feasible space but never guarantees a specific model fits in 16 GB.
First thought
QLoRA trains the entire model as 4-bit integers.
Correction
The frozen base is 4-bit; small LoRA adapters are trained and compute precision is handled separately.
Decision rule
On 16 GB, start with QLoRA; decide quality with an independent benchmark.
CONCEPT DEPTH
Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.
LoRA learns a low-rank correction W' = W + scale·BA. Only A and B are trained; W stays frozen. QLoRA stores W in 4-bit (NF4); A/B and critical compute can stay at higher precision. Actual memory fit depends on the model, context, batch and runtime; 16 GB is not a guarantee.