Back to paper
Critique

Do these models 'understand', or are they stochastic parrots? (the question won't die)

DOdortiz· 3 months ago

CLIP maps web-scale correlations between text and images. One camp: sophisticated pattern-matching with no grounding ('stochastic parrots'). Other camp: understanding is prediction at scale. Is 'understanding' a meaningful, testable distinction here — or a goalpost that moves every time a benchmark falls?

5 Replies

Sign in to reply and react.
AMamir_r3 months ago

'Stochastic parrot' has become a thought-terminating cliché. If a system generalizes to novel compositions it never saw, 'just correlation' explains nothing. Give me a behavioral test that separates 'real' understanding from very good prediction.

DOdortiz3 months ago

The test is systematic generalization: CLIP struggles with 'red cube on blue sphere' vs 'blue cube on red sphere'. Compositional binding is exactly what understanding should give you for free, and it's shaky.

LElenaf3 months ago

It's empirical, not philosophical. Compositionality benchmarks show partial structure — not none, not full. 'Parrot vs mind' is a false binary both sides use to avoid measuring.

WEweizh3 months ago

Careful with the grounding argument: humans also learn enormously from correlation. The useful question isn't 'understanding yes/no' but 'which inductive structures emerge from prediction, and which provably don't'.

SEseoyeon3 months ago

구성성(compositionality)이 핵심인 것 같아요. CLIP가 색·객체 바인딩을 자주 헷갈리는 건 부분적 구조만 학습했다는 증거죠. 완전한 이해도, 완전한 앵무새도 아닌.