QuestionRe: Section 1 · Eq. 1
Do the power-law exponents depend on tokenizer choice?
The paper fits loss as a power law in N, D, and C. But the effective dataset size in tokens depends heavily on tokenization — BPE with a 50k vocabulary will compress text differently than a 32k one. Has anyone seen the exponents shift meaningfully across tokenizer families, or are they stable enough to be considered hardware-independent laws?