Continued and domain-adaptive pretraining
Next-token training on a domain corpus before any instruction data. Justified when the domain's language and structure are far from the base model's distribution, and skipped when they are not, because it is the most compute-hungry stage.