Back to paper
QuestionRe: Section 1 · Eq. 1

Do the power-law exponents depend on tokenizer choice?

AMamir_r· Stanford NLP· about 1 month ago

The paper fits loss as a power law in N, D, and C. But the effective dataset size in tokens depends heavily on tokenization — BPE with a 50k vocabulary will compress text differently than a 32k one. Has anyone seen the exponents shift meaningfully across tokenizer families, or are they stable enough to be considered hardware-independent laws?

1 Reply

Sign in to reply and react.
LElenafabout 1 month ago

Good question. The Chinchilla paper (Hoffmann et al. 2022) refits with similar tokenization and gets slightly different exponents, which suggests they aren't universal constants. Vocabulary size probably matters at the margin, but the qualitative shape holds.