Back to paper
Critique

Do these models 'understand', or are they stochastic parrots? (the question won't die)

DOdortiz· about 1 month ago

CLIP maps web-scale correlations between text and images. One camp: sophisticated pattern-matching with no grounding ('stochastic parrots'). Other camp: understanding is prediction at scale. Is 'understanding' a meaningful, testable distinction here — or a goalpost that moves every time a benchmark falls?

5 Replies

Sign in to reply and react.
AMamir_rabout 1 month ago

'Stochastic parrot' has become a thought-terminating cliché. If a system generalizes to novel compositions it never saw, 'just correlation' explains nothing. Give me a behavioral test that separates 'real' understanding from very good prediction.

DOdortizabout 1 month ago

The test is systematic generalization: CLIP struggles with 'red cube on blue sphere' vs 'blue cube on red sphere'. Compositional binding is exactly what understanding should give you for free, and it's shaky.

LElenafabout 1 month ago

It's empirical, not philosophical. Compositionality benchmarks show partial structure — not none, not full. 'Parrot vs mind' is a false binary both sides use to avoid measuring.

WEweizhabout 1 month ago

Careful with the grounding argument: humans also learn enormously from correlation. The useful question isn't 'understanding yes/no' but 'which inductive structures emerge from prediction, and which provably don't'.

SEseoyeonabout 1 month ago

구성성(compositionality)이 핵심인 것 같아요. CLIP가 색·객체 바인딩을 자주 헷갈리는 건 부분적 구조만 학습했다는 증거죠. 완전한 이해도, 완전한 앵무새도 아닌.