08 · EVIDENCE
PlannedDataset and Evaluation Discipline
Do not claim quality without an independent test set, stable schema, and safety examples.
- 55% standard, 15% missing
- Train ≠ test
- Multi-metric benchmark
- Reject if safety drops
KNOW THESE FIRST
The budget, window, and …This atlas's illustrative mix (not a universal recommendation): 55% standard, 10% paraphrase, 15% missing information, 10% negative, and 10% escalation examples. Freeze a test set from a different asset or source.
A 100-question benchmark weighs domain accuracy, format, safety, uncertainty, and general ability retention. Reusing training data for evaluation is not generalization evidence.
First thought
Loss on the train split is independent quality evidence.
Correction
Separate validation and test sets are required; loss alone does not represent the real objective.
Decision rule
Freeze the benchmark before training and test base/adapted models with identical generation settings.
CONCEPT DEPTH
Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.
Train/validation/test splits must be frozen from separate sources. This atlas’s illustrative data mix (not universal): 55% standard, 10% paraphrase, 15% missing info, 10% negative, 10% escalation. A 100-question benchmark scores with a weighted average of domain/format/safety/uncertainty/retention. Reusing training data for evaluation is not generalization evidence.