02 · CONTEXT
VerifiedTokens, Context, and Attention
Unify token budgets, fixed weights, and temporary attention state in one mental model.
- Token = chunk of text
- Context = total token budget
- Attention = focus selection
- KV cache = temporary memory
KNOW THESE FIRST
The budget, window, and …Context is the total token budget for system text, template, user input, RAG content, and response reserve. Token IDs and segmentation are model-specific; Turkish cost must be measured with the target tokenizer.
During inference WQ/WK/WV stay fixed. Q/K/V vectors and attention scores depend on the input; KV cache temporarily stores past K/V activations.
First thought
Doubling context always quadruples total VRAM.
Correction
Classic attention work grows roughly 4×, but fixed weights and runtime details mean total VRAM need not grow at the same rate.
Decision rule
Measure the real length distribution; establish a safe 1024–2048 baseline first.
CONCEPT DEPTH
Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.
BPE is a subword algorithm; SentencePiece is a toolkit supporting multiple algorithms. Token count depends on vocabulary, normalization and text. Avoid a fixed Turkish cost multiplier; measure paired meanings with the same tokenizer and template.