Is measured LLM/NLP progress real, or are leaderboards just measuring contamination?
Pretraining on web-scale data means test sets routinely leak into training. When a model tops GLUE/MMLU/etc., how much is genuine capability vs the answers being in the training corpus? Is any public benchmark trustworthy anymore?