QWEN3 4B · FINE-TUNING

Evidence record

A historical source record dated 2026-08-08, separating pipeline success from model quality. This website does not run training or connect to a GPU.

Verified30 / 30training steps completed
Verified0.8245final training loss
Verified✓adapter save and reload
Observed9.62 / 15.99 GiBobserved snapshot; not peak

VERDICT

Pipeline passed. Quality gain was not proven.

There was no evaluation split, and the base model was stronger in all three fixed quality comparisons. This run proves a smoke-test pipeline, not a model improvement.

Turkish explanationBase more coherentFine-tuned repetition/language mixing
Conditional JSONBase more validFine-tuned incomplete/loose
Missing informationBase more cautiousFine-tuned made more assumptions

Unknowns

  • Real peak VRAM was not measured.
  • There was no independent validation/test split.
  • There was no fixed seed or repeated run.
  • Statistical confidence of the quality delta is unknown.

READING LIST

Classic Papers

Mathematical basis of each concept. Short summary; use as a starting point before reading the original paper.

Observed
2017 · Vaswani et al.

Attention Is All You Need (Transformer)

TAKEAWAY

Without RNN/CNN, only multi-head self-attention can do sequence modeling. Q/K/V linear projection, scaled dot-product attention, parallel training.

WHY IT MATTERS

The building block of modern LLMs. Mathematical basis of token/attention/context concepts.

CITATION

Vaswani, A. et al. (2017). Attention Is All You Need. NeurIPS.

2021 · Hu et al.

LoRA: Low-Rank Adaptation of Large Language Models

TAKEAWAY

Learns a low-rank weight update; the paper reports up to 10,000× fewer trainable parameters in its GPT-3 175B comparison. This is not a universal reduction factor.

WHY IT MATTERS

Original source of the LoRA concept. The r × (d_in + d_out) formula and alpha/r scale come from here.

CITATION

Hu, E.J. et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685.

2023 · Dettmers et al.

QLoRA: Efficient Finetuning of Quantized LLMs

TAKEAWAY

NF4 (4-bit normal float), double quantization, and paged optimizer states fine-tune a 65B model on a single 48 GB GPU. This demonstrates memory savings but does not guarantee that a particular model fits in 16 GB.

WHY IT MATTERS

Original source of QLoRA, NF4, double-quantization, page optimizer concepts.

CITATION

Dettmers, T. et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS.

2024 · Shao et al. (DeepSeek)

DeepSeekMath: GRPO

TAKEAWAY

Group Relative Policy Optimization: simpler than PPO, no value model, in-group advantage normalization. Became widespread in reasoning models.

WHY IT MATTERS

The method behind Week 11's GRPO experiment. Foundation for automatically verifiable rewards.

CITATION

Shao, Z. et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300.

CITATION KIT

Sentences for Thesis and Reports

Evidence-labeled draft sentences. Verify the original source and cite it directly before using them.

VerifiedLoRAusage: For the LoRA core formula.
LoRA limits trainable parameters to r × (d_in + d_out) while keeping the original W matrix frozen; this dramatically lowers VRAM and storage requirements compared to full fine-tuning.
Verifiedevaluationusage: To defend the eval split requirement.
Without an independent test set, a quality claim from fine-tuning cannot be made; train-split loss is not generalization evidence (Bishop, 2006; Goodfellow et al., 2016).
Observedtokensusage: To discuss Turkish tokenization cost.
Turkish token cost varies by tokenizer; context and inference cost must be measured with the target model's actual tokenizer.
Verifiedstepsusage: To explain training mechanics.
Effective batch size = micro batch × gradient accumulation × GPU count; on OOM, the first variable to drop is micro batch or sequence length, with accumulation raised to compensate for effective batch.
Verifiedtemplatesusage: Why chat template is critical.
The same role/content record turns into different token sequences under different models' chat templates; therefore semantic records must be stored independently of the model, and rendering must use each model's own template.