filippo.maretti
All posts
August 11, 202618 min

Parrots 🦜

Why the stochastic parrot metaphor doesn't hold up any more

AI

The "stochastic parrot" is a metaphor used to describe language models in a detached, materialist way. Basically, it's the thing boomers still stuck on GPT-3.5 drop in the comments of a LinkedIn post where some enthusiastic kid writes that ChatGPT just solved a maths problem for them. The phrase got famous because it is sharp and memorable, but the problem is that it isn't true any more.

This article is an attempt to put that phrase back in its place. Which means taking it seriously enough to explain why it is no longer enough.


In 2021 the parrot was right

March 2021, the FAccT conference. Emily Bender, Timnit Gebru, Angelina McMillan-Major and one Shmargaret Shmitchell publish On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜. The emoji is in the original title, I didn't add it (lol).

Inside that paper the parrot takes up a paragraph at most: a system that haphazardly stitches together bits of language it has already seen, following the probability with which those bits sit next to each other, never passing through meaning. There is no chapter devoted to it. It's an image tossed in to get the point across in one line.

And it worked far too well: it ate everything else (which, by the way, is still very interesting and very current, but we'll come back to that at the end).

There's also something almost nobody knows: the serious philosophical argument, the one that explains why a model shouldn't understand, isn't in this paper. It's in a piece of work from the year before, again by Bender, this time with Alexander Koller, and it's about hyper-intelligent octopuses. It's one of the best thought experiments of the last twenty years.

The octopus

Two castaways end up on two different islands. Luckily an old telegraph cable runs between them, and they start writing to each other: they chat about this and that, what they ate, what the weather is like. On the seabed near the cable lives a hyper-intelligent octopus. It has never seen an island, nor a human being, but it taps the cable and spends years listening in. At some point it has learned the structure of those conversations so well that it decides to cut the cable and answer itself, passing as the other castaway.

Being the Stephen Hawking of octopuses, for weeks it imitates the two castaways without breaking a sweat, because "what's the weather like?" and "I'm dead tired" are conversations that follow a sort of implicit script.

Then one day one of the castaways is attacked by a bear. They write to their friend saying they've built a catapult out of coconuts and sticks, and urgently ask how to use it to defend themselves.

The octopus panics. It hasn't the faintest idea what a bear is.

This is the heart of the whole thing, and it is a serious argument: can you learn the form of a language without having access to what that language refers to? The octopus has the syntax. It doesn't have the world. And as long as the conversation stays on the surface, the difference doesn't show.

And that's where the parrot comes from: the exact same argument, just with a more familiar animal.


The problem isn't the thesis. It's the date.

The paper is from March 2021, and it describes with precision the models that existed at that moment: BERT, GPT-2, GPT-3, Switch-C.

Now let's line the calendar back up:

Bender and colleagues didn't get the map wrong: they mapped the territory in front of them beautifully. It's just that the territory then changed, and changed a lot.

Because a parrot repeats sounds. It doesn't hold together a chain of 50 steps where each one depends on the ones before. It doesn't take a principle from physics and one from biology and use them to propose a solution that wasn't anywhere in its training.

In 2026, systems built on language models solved International Mathematical Olympiad problems at gold-medal level. Not old problems, the kind you could find in the training data: problems written that year, designed specifically not to resemble anything already seen. And you don't need to go all the way to the Olympiad: open a codebase you know inside out and ask Claude to implement something new or fix a decade-old bug, and you'll be surprised how well it does it. Then sure, underneath there is always and only a machine guessing the next word. But that this stuff comes out of that stuff, this consistently, is still hard to swallow. You can call what happens in there whatever you like. But by now, calling it a parrot makes no sense.


Where the parrot is still right

And here comes the honest part, otherwise this article turns into a brochure.

Because models do fail. And they fail in ways that seem to prove the parrot right.

There are two results that always come up, and they're the best arguments available to the people who think the opposite. They're worth a close look, because they ended up in two very different places.

Change the numbers and it all falls apart

The first is the work on grade-school maths problems: take the same problem, change the names and the numbers, and accuracy drops. Add an irrelevant but plausible sentence like "five of those apples were a bit smaller than average", and some models collapse by tens of percentage points, because that sentence looks like data and so the model tries to use it. (Note that the example is a simplification so we can follow along, the original problems were considerably more devious)

It's the most cited anti-model paper of the last two years. And in 2026 something unpleasant happened to it: someone went and redid the sums.

Re-analysing twenty models with the right statistical tools, only half show a genuinely significant drop. And a detail surfaces that the original authors had explicitly ruled out: the rewritten problems contain systematically larger numbers than the ones they started from. Translation: a slice of the collapse isn't "you changed the names and it got lost", it's "you gave it harder sums".

Then there's the story of the irrelevant sentences, and this one is my favourite. Someone re-ran the test on today's models, but first sat down and checked those sentences one by one. Out of 945, only 117 were genuinely irrelevant: all the others, looked at closely, really could have an effect on the answer. And on those 117 the drop is between zero and two points.

In other words: the model wasn't getting fooled by a random sentence. It was taking seriously a piece of information that might matter, which is exactly what you would do faced with a badly written problem.

(This second one is a replication published on a blog, not peer-reviewed work. I cite it because it is reproducible and the source is reliable.)

The reversal curse

This one, on the other hand, holds. And it's far more interesting.

It's called the reversal curse: a model that has learned "Tom Smith is the father of John Doe" may be unable to answer "who is Tom Smith's son?". (NB: again this is a hyper-simplified example just to show the concept, in reality far more complex relations were used)

It looks like the parrot's smoking gun. Then in 2025 two researchers went looking for why it happens, and did the one thing that actually mattered: they took the names away.

Instead of "Tom Smith" spelled out in letters, every entity gets a fixed identifier of its own, a sort of serial number, always the same, which the model doesn't try to look up in its memory because it recognises it as an identifier. Nothing else changes: same relations, same training, same transformer as before. The only thing that disappears is the step "I read this sequence of letters and work out who we're talking about".

And it learns the reversal.

So the relation isn't the problem: the model can flip that just fine. The problem is the bridge between the name and the thing. And more precisely: that bridge doesn't hold when the same entity changes role.

In cognitive science this phenomenon has had a name for decades: it's called the binding problem.

It's when you run into someone on the street and you know perfectly well who they are: where they work, who they're with, that you went to middle school together. The name just won't come. Nobody would say that in that moment you're missing a model of that person, you know everything about them. What you're missing is the label.

That's exactly where the difference sits. A parrot gets it wrong because there's nothing underneath. Here everything is underneath, and what broke is the thread tying it to the name.

The parrot isn't dead: it's an edge case. It describes a failure of the system. It doesn't describe the system.


To compress is to understand

So why should a machine trained to guess the next word build itself a model of the world?

The most elegant answer I know is an argument Ilya Sutskever has been repeating for years, and it rests entirely on a single word: compression.

Predicting the next token over a corpus as large as everything ever written, with a finite number of parameters at your disposal, is a compression problem. You have to fit that text into that space. And at that point you have two roads.

Pure memorisation is the worst compression possible: it takes up an enormous amount of room and does nothing for you on the next sentence. The regularities that produced that text, on the other hand, are compact and reusable.

Put another way: the world's text was written by people living inside a reality that has laws. To predict that text well, the cheapest route is to reconstruct those laws. Statistics is the means. Semantics is the shortcut the network finds on its own because it costs less.

Nice as an argument. The point is that it isn't just an argument any more.

My favourite piece of evidence

In 2023 a group of researchers from Harvard and Northeastern trained a small GPT on Othello to predict the next move in lists of games. The model has never seen a board (it doesn't even know one exists) and has no idea what Othello is. It sees sequences of codes, which we happen to read as moves.

Then they went rummaging around inside it with probes, and everything was there in the activations, including the board. The state of the game, square by square, reconstructed on its own with no input at all.

And here comes the part that settles the question. A correlation can always be dismissed as coincidence, so they did the one thing that really removes the doubt: they intervened. They surgically edited that internal representation, changing the contents of one square, and asked for the next move.

The model answered consistently with the fake board.

There's also a footnote I find rather funny: at first the board looked encoded in a strange, convoluted way. Then it turned out the model wasn't representing it as "black pieces and white pieces", but as "mine and theirs". The model of the world was there; it just wasn't in our coordinates.

Let's keep this one aside, we'll need it later.

And it isn't an isolated case

Bender and Koller's octopus has never seen a bear. But if it has read enough about bears, about woods, about fear and about levers, then it will know what to answer.


"Yeah but they're just algorithms"

This is the final objection, the one that shows up once everything else has been conceded: at the end of the day it's just maths, it's just matrix multiplications, it's just code.

And it's true. It is literally true. The problem is that as an argument it doesn't work, because it is reductionism.

By the exact same reasoning, The Starry Night is a bit of chemical pigment smeared on a piece of canvas. The description is 100% correct, and it explains nothing about what happens when you look at it.

Salvatore Sanfilippo, creator of Redis and Ds4 (as far as I'm concerned one of the very few genuinely lucid people in the field), explains this concept beautifully in a video you can find here. The parallel there is water and whirlpools; I'll try with starlings.

Starlings

At sunset there are clouds of starlings that bend and twist as if they were a single animal. Anyone who has ever seen them has had, for a moment, the impression that someone up there was directing traffic.

Nobody is directing. And the remarkable thing is that we know it for certain, because it has been measured: the group that studied those clouds (the ones above Roma Termini) is Giorgio Parisi's, Nobel laureate in physics in 2021 precisely for the work on disordered systems.

Two extremely cool things came out of those measurements:

So: none of the properties you see when you look up exists inside a single starling. It's not that each bird holds a little piece of them. They simply aren't there. They emerge from the interaction, and the level at which you describe them is not the level at which the rules are written.

A few hundred billion parameters interacting at every single token is, in every relevant respect, a complex system. Saying "they're just algorithms" to close the discussion is the same as saying a flock is just a bunch of birds. True, and completely useless.


The paper shouldn't be thrown out, it should be promoted

And now that we've looked at what's inside these systems, we can go back to 2021 as promised.

I said at the start that the parrot, in that paper, is one paragraph. The rest is four warnings, and it's the part nobody ever cites:

  1. Costs. Training ever-larger models costs energy and money in quantities that are hard to justify, and the environmental bill is paid mostly by the people who will never see any benefit from those models.
  2. Data nobody has read. A dataset of hundreds of billions of words isn't documented by anyone any more. And the internet is not humanity: it's the portion of humanity that has a connection, free time and the urge to write. The model inherits that point of view and hands it back with the authority of a machine.
  3. Opportunity cost. If scaling works, all the research crowds in there and the alternative roads stop being travelled.
  4. Manipulation. Plausible text at zero cost, in industrial quantities, is a problem even if the machine understands nothing. In fact: especially if it understands nothing.

Five years on, none of these four points has been disproved, if anything it looks like they had a crystal ball.

And the second one now weighs far more than it did back then. Because if these systems really do have a model of the world, the biases Bender was talking about stop being stains on the output and become structure. Which makes every single concern in that paper more urgent.

Why it became a flag

This has to be said too, because it explains a lot about how things went afterwards. That paper, while it was still a draft in internal review, cost Timnit Gebru their job at Google: December 2020, three months before it came out. The company maintains they resigned, they maintain they were fired. Margaret Mitchell, who co-led that same Ethical AI team and who would later sign under the pseudonym mentioned above, left two months after that.

From that moment on the stochastic parrot stopped being a scientific hypothesis and became a flag.

And the problem with flags is that nobody updates them.

That's why what should be abandoned isn't the paper: it's the use that has been made of it. A 2021 hypothesis used as a shield in 2026 isn't scientific rigour. It's the exact opposite, it's politics.


We're missing the word

The real problem, in the end, is that we have two boxes: it understands and it doesn't understand. And both of them are wrong.

"It doesn't understand" no longer holds in front of a system that plans a rhyme three words before writing it and that keeps its concepts in a single place regardless of the language.

But "it understands" doesn't hold either, at least not in the sense we mean when we say it of a human being: for us, understanding means living, interacting with the world, and in there there is no lived experience at all.

What there is is a third thing, and we don't have the term for it: a semantics without a referent. One that has a model of the world, built entirely out of the shadow the world cast onto text.

And here the earlier note comes back, the one about the "mine and theirs" board. Because the lesson isn't that the model resembles us. It's the opposite: the model of the world is there, but it's written in coordinates of its own. They aren't ours, and there was no reason they should have been.


Sources