05 · TRAINING MECHANICS
VerifiedEpochs, Steps, and Gradient Accumulation
Separate micro-steps from real weight updates and calculate the correct OOM response.
- Micro-step = 1 batch pass
- Optimizer step = 1 weight update
- Eff. batch = μ × accum × GPU
- On OOM: lower μ, raise accum
KNOW THESE FIRST
The budget, window, and …A micro-step is one forward/loss/backward pass for a micro batch. One optimizer step updates weights after accumulation completes.
Effective batch = micro batch × accumulation × GPU count. For OOM, reduce micro batch or sequence length; increase accumulation if the target effective batch must be preserved.
First thought
Accumulation 8 means the weights update eight times.
Correction
Eight backward passes accumulate gradients, then one optimizer step updates weights once.
Decision rule
Find the largest safe micro batch first; record step counts as optimizer steps.
CONCEPT DEPTH
Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.
A micro-step is one forward/loss/backward pass for a micro batch. An optimizer step is one weight update after accumulation completes. Effective batch = μ × accumulation × GPU count. On OOM, the first thing to drop is micro batch or sequence length; if the effective batch must be preserved, accumulation is increased.