Skip to content

Chunking & Packing

  • Concatenate documents into one long token stream
  • Cut it into fixed-length windows (the context length)
  • Group windows into batches for the GPUs
  • Aside: this is not the same as RAG chunking (see the API section)