The model predicts a distribution over the next token; the loss is cross-entropy against the actual next token in the data.
A scalable, self-supervised objective learnable from unlabeled text was needed.