07 · PROTOCOL

Verified

Chat Template Mismatch

The same semantic dataset does not mean forcing the same rendered token sequence on every model.

QUICK LOOK
  • Store role/content independently
  • Each model has its own template
  • Inspect render manually
  • Same template at inference

KNOW THESE FIRST

role: systemrole: userrole: assistanttemplaterender<|im_start|>...<|im_end|>Mask: -100 (prompt) | real id (assistant)Correct mask: only the answer affects lossWrong mask: the answer may be excludedCONCEPT DIAGRAM

The budget, window, and …Store role/content records independently, then render with each model's own tokenizer and chat template. Preserve one validated template through training, evaluation, and inference.

If response-only masking delimiters do not match the real rendering, assistant answers can fall outside the loss or prompt tokens can be trained accidentally.

01

First thought

QLoRA repairs the wrong template while adapting.

02

Correction

QLoRA changes memory and optimization; it cannot fix incorrect control tokens or mask boundaries.

03

Decision rule

Inspect BOS/EOS and role boundaries in at least one rendered example.

CONCEPT DEPTH

Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.

UNDERGRAD

Each model has its own chat template (ChatML, Llama-3, Qwen, Phi). Store role/content records independently of the model; re-render with each model's own tokenizer and template. In response-only masking, if the assistant boundaries do not match the real render, the answer can fall outside the loss.