Glossary
Term glossary
Short definitions for the terms used across the Atlas. Each entry explains one concept and asserts no measurement or verified result.
30 terms
- Adapter
adapter Small plug-in module
A small trainable module that adds task behavior without modifying the base model.
- Alpha (α)
alpha Scale factor
Multiplier for the LoRA correction; α/r in standard LoRA, α/√r in rsLoRA.
- Attention
attention Selecting which token to focus on
The mechanism that assigns different weights to past tokens at each step.
- Base model
base-model Raw pretrained model
A starting model that only predicts the next token and is not aligned to follow instructions.
- Benchmark
benchmark Standard test set
An independent, fixed question/task set used to evaluate the model.
- Catastrophic forgetting
forgetting Losing prior ability
When training on new data erases previously learned abilities.
- Chat template
chat-template Messages to token sequence
Template that converts role/content records into a model's special-token format.
- Context
context Total token budget
The total set of token positions the model conditions on at one step.
- Data leakage
leakage Test data leaks into training
When test/validation samples accidentally appear in training, inflating scores.
- Effective batch
effective-batch Real batch size
μ × accumulation × GPU count. The actual batch used in a weight update.
- Epoch
epoch 1 pass over the data
One full pass of the training set through the model.
- GGUF
gguf Local inference format
Quantized model format used by llama.cpp, Ollama, and similar local runtimes.
- Gradient accumulation
accumulation Cumulative gradient
Method of combining gradients from several micro-steps into one optimizer step.
- GRPO
grpo Group Relative Policy Optimization
An RL method that updates the policy using an automatically verifiable reward signal.
- Instruct model
instruct-model Instruction-aligned model
A model aligned with RLHF or SFT to follow conversation and instructions.
- KV cache
kv-cache Cached K/V vectors
Temporary memory of past K/V vectors to avoid recomputation.
- LoRA
lora Low-Rank Adaptation
A method that adds a small scale × BA correction to a frozen W matrix to reduce trainable parameters.
- Loss
loss Error measure
A numeric measure of the difference between the model's prediction and the actual target.
- Merge
merge Bake the adapter into base
Adding the LoRA weights back into the base to obtain a single model.
- Micro batch
micro-batch Examples per pass
Number of examples processed in a single forward/loss/backward pass.
- OOM
oom Out of memory
GPU memory exhausted; typically fixed by lowering micro-batch or context.
- Optimizer step
step 1 weight update
A single update of model weights using accumulated gradients.
- Overfitting
overfitting Memorizing the training data
When the model learns the training data too well and degrades on unseen data.
- QLoRA
qlora 4-bit base + LoRA
Stores the frozen base in 4-bit (NF4) while training adapters at higher precision.
- RAG
rag Retrieval-Augmented Generation
Fetching relevant passages from an external source and adding them to the prompt.
- Rank (r)
rank Adapter capacity
Inner dimension of LoRA's low-rank matrices; sets capacity and parameter count.
- Response-only masking
masking Train only the answer
Masking the loss so only the assistant's response contributes.
- SFT
sft Supervised Fine-Tuning
Fine-tuning on labeled examples in a supervised way.
- Token
token Numeric piece of text
The smallest unit a model processes; can be a word, sub-word, or punctuation. Produced by a tokenizer.
- VRAM
vram GPU memory
On-GPU memory holding the model, activations, and KV cache.