Glossary

Term glossary

Short definitions for the terms used across the Atlas. Each entry explains one concept and asserts no measurement or verified result.

30 terms

Adapteradapter

Small plug-in module

A small trainable module that adds task behavior without modifying the base model.

Alpha (α)alpha

Scale factor

Multiplier for the LoRA correction; α/r in standard LoRA, α/√r in rsLoRA.

Attentionattention

Selecting which token to focus on

The mechanism that assigns different weights to past tokens at each step.

Base modelbase-model

Raw pretrained model

A starting model that only predicts the next token and is not aligned to follow instructions.

Benchmarkbenchmark

Standard test set

An independent, fixed question/task set used to evaluate the model.

Catastrophic forgettingforgetting

Losing prior ability

When training on new data erases previously learned abilities.

Chat templatechat-template

Messages to token sequence

Template that converts role/content records into a model's special-token format.

Contextcontext

Total token budget

The total set of token positions the model conditions on at one step.

Data leakageleakage

Test data leaks into training

When test/validation samples accidentally appear in training, inflating scores.

Effective batcheffective-batch

Real batch size

μ × accumulation × GPU count. The actual batch used in a weight update.

Epochepoch

1 pass over the data

One full pass of the training set through the model.

GGUFgguf

Local inference format

Quantized model format used by llama.cpp, Ollama, and similar local runtimes.

Gradient accumulationaccumulation

Cumulative gradient

Method of combining gradients from several micro-steps into one optimizer step.

GRPOgrpo

Group Relative Policy Optimization

An RL method that updates the policy using an automatically verifiable reward signal.

Instruct modelinstruct-model

Instruction-aligned model

A model aligned with RLHF or SFT to follow conversation and instructions.

KV cachekv-cache

Cached K/V vectors

Temporary memory of past K/V vectors to avoid recomputation.

LoRAlora

Low-Rank Adaptation

A method that adds a small scale × BA correction to a frozen W matrix to reduce trainable parameters.

Lossloss

Error measure

A numeric measure of the difference between the model's prediction and the actual target.

Mergemerge

Bake the adapter into base

Adding the LoRA weights back into the base to obtain a single model.

Micro batchmicro-batch

Examples per pass

Number of examples processed in a single forward/loss/backward pass.

OOMoom

Out of memory

GPU memory exhausted; typically fixed by lowering micro-batch or context.

Optimizer stepstep

1 weight update

A single update of model weights using accumulated gradients.

Overfittingoverfitting

Memorizing the training data

When the model learns the training data too well and degrades on unseen data.

QLoRAqlora

4-bit base + LoRA

Stores the frozen base in 4-bit (NF4) while training adapters at higher precision.

RAGrag

Retrieval-Augmented Generation

Fetching relevant passages from an external source and adding them to the prompt.

Rank (r)rank

Adapter capacity

Inner dimension of LoRA's low-rank matrices; sets capacity and parameter count.

Response-only maskingmasking

Train only the answer

Masking the loss so only the assistant's response contributes.

SFTsft

Supervised Fine-Tuning

Fine-tuning on labeled examples in a supervised way.

Tokentoken

Numeric piece of text

The smallest unit a model processes; can be a word, sub-word, or punctuation. Produced by a tokenizer.

VRAMvram

GPU memory

On-GPU memory holding the model, activations, and KV cache.