Is chain-of-thought just a better input distribution, or something deeper?
The paper shows CoT boosts reasoning, but the mechanism is unclear. One hypothesis: CoT simply converts hard one-step problems into easier multi-step ones that happen to lie in the training distribution. Another: it induces genuine intermediate computation. Section 5 mentions emergent properties at 100B+ parameters — does this exclude the distribution-shift explanation, or are both in play?