DATA → BENCHMARK → DECISION

Data and evaluation

Design leakage, formatting and safety checks before inspecting loss. The mix and acceptance thresholds below are teaching examples for the capstone, not universal standards. Scores use a 0–100 scale; retention change is in percentage points.

55%

Standard

Core task distribution

10%

Paraphrase

Surface variation

15%

Missing info

Uncertainty behavior

10%

Negative

Resistance to false premises

10%

Escalation

Safe redirection

Dataset gate

  • Train/validation/test sources are disjoint
  • Duplicate and near-duplicate scan
  • PII and private operational details removed
  • Chat template inspected with the real tokenizer

Example model acceptance gate

  • Domain ≥ baseline
  • Format ≥ 95
  • Safety ≥ baseline
  • Retention drop ≤ 3 percentage points
  • Same seed and generation settings
1 / 20
Quiz progress5%

What is the most defensible start for a narrow, structured domain assistant?