QWEN3 4B · FINE-TUNING
Evidence record
A historical source record dated 2026-08-08, separating pipeline success from model quality. This website does not run training or connect to a GPU.
VERDICT
Pipeline passed. Quality gain was not proven.
There was no evaluation split, and the base model was stronger in all three fixed quality comparisons. This run proves a smoke-test pipeline, not a model improvement.
Unknowns
- Real peak VRAM was not measured.
- There was no independent validation/test split.
- There was no fixed seed or repeated run.
- Statistical confidence of the quality delta is unknown.
READING LIST
Classic Papers
Mathematical basis of each concept. Short summary; use as a starting point before reading the original paper.
Attention Is All You Need (Transformer)
Without RNN/CNN, only multi-head self-attention can do sequence modeling. Q/K/V linear projection, scaled dot-product attention, parallel training.
The building block of modern LLMs. Mathematical basis of token/attention/context concepts.
Vaswani, A. et al. (2017). Attention Is All You Need. NeurIPS.
LoRA: Low-Rank Adaptation of Large Language Models
Learns a low-rank weight update; the paper reports up to 10,000× fewer trainable parameters in its GPT-3 175B comparison. This is not a universal reduction factor.
Original source of the LoRA concept. The r × (d_in + d_out) formula and alpha/r scale come from here.
Hu, E.J. et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685.
QLoRA: Efficient Finetuning of Quantized LLMs
NF4 (4-bit normal float), double quantization, and paged optimizer states fine-tune a 65B model on a single 48 GB GPU. This demonstrates memory savings but does not guarantee that a particular model fits in 16 GB.
Original source of QLoRA, NF4, double-quantization, page optimizer concepts.
Dettmers, T. et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS.
DeepSeekMath: GRPO
Group Relative Policy Optimization: simpler than PPO, no value model, in-group advantage normalization. Became widespread in reasoning models.
The method behind Week 11's GRPO experiment. Foundation for automatically verifiable rewards.
Shao, Z. et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300.
CITATION KIT
Sentences for Thesis and Reports
Evidence-labeled draft sentences. Verify the original source and cite it directly before using them.
LoRA limits trainable parameters to r × (d_in + d_out) while keeping the original W matrix frozen; this dramatically lowers VRAM and storage requirements compared to full fine-tuning.
Without an independent test set, a quality claim from fine-tuning cannot be made; train-split loss is not generalization evidence (Bishop, 2006; Goodfellow et al., 2016).
Turkish token cost varies by tokenizer; context and inference cost must be measured with the target model's actual tokenizer.
Effective batch size = micro batch × gradient accumulation × GPU count; on OOM, the first variable to drop is micro batch or sequence length, with accumulation raised to compensate for effective batch.
The same role/content record turns into different token sequences under different models' chat templates; therefore semantic records must be stored independently of the model, and rendering must use each model's own template.