03 · ADAPTER

Verified

LoRA and QLoRA

Adapt a frozen base with a low-rank update; measure memory fit on the target hardware.

QUICK LOOK
  • W frozen + Δ = W + BA
  • Only A and B train
  • QLoRA: base is 4-bit
  • Adapters are reversible

KNOW THESE FIRST

Wfrozen+B×A×α/rLoRA: only A and B trainQLoRA: W → 4-bit (NF4)CONCEPT DIAGRAM

The budget, window, and …LoRA trains only A and B adapters in W' = W + scale × BA. The default random A and zero B initialization preserves base behavior; other initialization methods need separate verification.

QLoRA stores the frozen base in 4-bit form while adapters and critical computation can use higher precision. It expands the feasible space but never guarantees a specific model fits in 16 GB.

01

First thought

QLoRA trains the entire model as 4-bit integers.

02

Correction

The frozen base is 4-bit; small LoRA adapters are trained and compute precision is handled separately.

03

Decision rule

On 16 GB, start with QLoRA; decide quality with an independent benchmark.

CONCEPT DEPTH

Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.

UNDERGRAD

LoRA learns a low-rank correction W' = W + scale·BA. Only A and B are trained; W stays frozen. QLoRA stores W in 4-bit (NF4); A/B and critical compute can stay at higher precision. Actual memory fit depends on the model, context, batch and runtime; 16 GB is not a guarantee.