Next-Token Prediction
What this page will cover
Section titled “What this page will cover”- Input vs. target: the same window shifted by one token
- Every position is a mini guessing game; the loss measures how wrong the guess was
- Gradient descent nudges billions of weights to make better guesses
- Intuition: compressing the training data forces the model to learn structure