ComparisonWith this toy splitting rule, the Turkish/English piece-count ratio is 0.77×. This ratio does not measure real tokenizer efficiency or inference cost.
V02
VRAM Budget Visualizer
Separates model, adapter, optimizer, gradient, and activation shares visually. Numbers are approximate; real measurement should use nvidia-smi.
Assumptions: 32 layers; 7 square LoRA matrices per layer; KV width is one quarter of hidden width; FP16 KV, FP32 adapters and Adam states. Actual architecture, quantization overhead, workspaces, and allocator overhead are excluded. This teaching model cannot guarantee fit.
Model weights1.86 GiB69%
Adapter0.03 GiB1%
Optimizer0.07 GiB3%
Gradients0.03 GiB1%
Activations0.69 GiB26%
Simplified estimate is below budgetKV cache (inference) estimate: 0.16 GiB. KV cache adds memory at inference time; it is not counted during training.
V03
Training Loss Simulator
Change learning rate, batch, epoch, and overfit risk; watch train/val loss curves live. Models exponential decay + overfit rise; does not replace a real run.
0.816Final train loss
1.994Final validation loss
420Best validation step
421Overfit onset
V04
Attention Heatmap
Shows how much each token in a sentence 'attends' to every other as a heatmap. Color intensity = attention weight. This is a simulation; it does not reflect a real model's attention.
This
teaching
example
shows
illustrative
attention
relationships
between
words
This
teaching
example
shows
illustrative
attention
relationships
between
words
·
10
10
5
5
5
5
4
1
12
·
10
8
6
8
6
6
4
10
10
·
10
8
10
7
7
3
8
8
10
·
10
10
10
8
4
6
6
10
10
·
10
10
12
6
6
8
10
10
10
·
10
10
11
5
6
8
10
10
10
·
14
12
4
4
5
5
5
10
10
·
17
2
2
2
2
2
2
3
3
·
Highest attention“words” → “words” (84%). The diagonal (self-attention) dominates this example; real patterns depend on the layer, head, and input.