Skip to content

Next-Token Prediction

  • Input vs. target: the same window shifted by one token
  • Every position is a mini guessing game; the loss measures how wrong the guess was
  • Gradient descent nudges billions of weights to make better guesses
  • Intuition: compressing the training data forces the model to learn structure