05 · TRAINING MECHANICS

Verified

Epochs, Steps, and Gradient Accumulation

Separate micro-steps from real weight updates and calculate the correct OOM response.

QUICK LOOK
  • Micro-step = 1 batch pass
  • Optimizer step = 1 weight update
  • Eff. batch = μ × accum × GPU
  • On OOM: lower μ, raise accum

KNOW THESE FIRST

fwdlossbwdaccum?repeatoptimizer stepupdates weights onceCONCEPT DIAGRAM

The budget, window, and …A micro-step is one forward/loss/backward pass for a micro batch. One optimizer step updates weights after accumulation completes.

Effective batch = micro batch × accumulation × GPU count. For OOM, reduce micro batch or sequence length; increase accumulation if the target effective batch must be preserved.

01

First thought

Accumulation 8 means the weights update eight times.

02

Correction

Eight backward passes accumulate gradients, then one optimizer step updates weights once.

03

Decision rule

Find the largest safe micro batch first; record step counts as optimizer steps.

CONCEPT DEPTH

Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.

UNDERGRAD

A micro-step is one forward/loss/backward pass for a micro batch. An optimizer step is one weight update after accumulation completes. Effective batch = μ × accumulation × GPU count. On OOM, the first thing to drop is micro batch or sequence length; if the effective batch must be preserved, accumulation is increased.