Do these models 'understand', or are they stochastic parrots? (the question won't die)
CLIP maps web-scale correlations between text and images. One camp: sophisticated pattern-matching with no grounding ('stochastic parrots'). Other camp: understanding is prediction at scale. Is 'understanding' a meaningful, testable distinction here — or a goalpost that moves every time a benchmark falls?