ComparisonTurkish produces about 0.83× more tokens than its English counterpart. This directly affects the context window and inference cost.
V02
VRAM Budget Visualizer
Separates model, adapter, optimizer, gradient, and activation shares visually. Numbers are approximate; real measurement should use nvidia-smi.
Model weights1.86 GiB97%
Adapter0.00 GiB0%
Optimizer0.00 GiB0%
Gradients0.00 GiB0%
Activations0.06 GiB3%
Fits the budgetKV cache (inference) estimate: 0.00 GiB. KV cache adds memory at inference time; it is not counted during training.
V03
Training Loss Simulator
Change learning rate, batch, epoch, and overfit risk; watch train/val loss curves live. Models exponential decay + overfit rise; does not replace a real run.
0.860Final train loss
2.183Final val loss
420Best val step
421Overfit onset
V04
Attention Heatmap
Shows how much each token in a sentence 'attends' to every other as a heatmap. Color intensity = attention weight. This is a simulation; it does not reflect a real model's attention.
The
model
tokenizes
Turkish
into
more
tokens
than
English
The
model
tokenizes
Turkish
into
more
tokens
than
English
·
10
10
5
5
5
5
4
1
12
·
10
8
6
8
6
6
4
10
10
·
10
8
10
7
7
3
8
8
10
·
10
10
10
8
4
6
6
10
10
·
10
10
12
6
6
8
10
10
10
·
10
10
11
5
6
8
10
10
10
·
14
12
4
4
5
5
5
10
10
·
17
2
2
2
2
2
2
3
3
·
Highest attention“English” → “English” (84%). The diagonal (self-attention) usually dominates.