08 · EVIDENCE

Planned

Dataset and Evaluation Discipline

Do not claim quality without an independent test set, stable schema, and safety examples.

QUICK LOOK
  • 55% standard, 15% missing
  • Train ≠ test
  • Multi-metric benchmark
  • Reject if safety drops

KNOW THESE FIRST

DomainFormatSafetyRetentionUncertaintyD·.35 + F·.15 + S·.20 + U·.15 + R·.15CONCEPT DIAGRAM

The budget, window, and …This atlas's illustrative mix (not a universal recommendation): 55% standard, 10% paraphrase, 15% missing information, 10% negative, and 10% escalation examples. Freeze a test set from a different asset or source.

A 100-question benchmark weighs domain accuracy, format, safety, uncertainty, and general ability retention. Reusing training data for evaluation is not generalization evidence.

01

First thought

Loss on the train split is independent quality evidence.

02

Correction

Separate validation and test sets are required; loss alone does not represent the real objective.

03

Decision rule

Freeze the benchmark before training and test base/adapted models with identical generation settings.

CONCEPT DEPTH

Read the same concept at three depths. Pick a level, switch instantly. Recommended starting point for university students: UNDERGRAD.

UNDERGRAD

Train/validation/test splits must be frozen from separate sources. This atlas’s illustrative data mix (not universal): 55% standard, 10% paraphrase, 15% missing info, 10% negative, 10% escalation. A 100-question benchmark scores with a weighted average of domain/format/safety/uncertainty/retention. Reusing training data for evaluation is not generalization evidence.