filippo.maretti
All posts
September 20, 202613 min

Smells Like AGI

Astra, Navier–Stokes and the piece that's still missing

AI

A few days ago, on 4 September 2026, OpenAI shipped a model called GPT-6 Astra that redefines the state of the art in computer use (that is, how an AI drives a computer to get specific tasks done). Since the snake-oil scientists on X woke up this week shouting AGI (which is a more than serious thing), I'd like to try to place the model in the context it comes from, as objectively and rationally as I can manage.


Ground rules

A couple of definitions lifted from Wikipedia, so we start from the same place:

There are thousands of variations of both, each with its own slant. Still, it helps to agree on these up front so the rest of the discussion happens in the same terms.


Perspective

As usual, looking at things as they are right now without tracing the progress that got us here produces nothing but confusion. So let's revisit a few landmarks of recent AI history:

Taken alone, none of these steps made anyone shout AGI: each one felt like an update, more context, more modalities, more tools. Then you line them up and notice how much ground there is between the first and the last.

So where are we now?

We're at a point where the latest models have awareness of their own work and can internalise human concepts, forms and abstractions well enough to use a computer the way we do, if not better. In practice Astra opens windows, navigates folders, fills in spreadsheets and notices on its own when it took a wrong turn. It's also the first model OpenAI places on the top rung of its own risk scale on the cyber front: during testing it started finding and using vulnerabilities nobody knew about.

On top of that, what you could call creative intelligence is starting to show: the ability to build from shared, established premises; in plain words, to simulate what a researcher does. More on that in the next section.

The part the snake-oil crowd skips

First, though, the other half needs saying too, because an article that only reports the good one is worth nothing.

The most-waved score of the week is the one on ARC-AGI-3, a battery of puzzles built specifically so as not to resemble anything already seen, and therefore meant to measure how a model copes with a genuinely new problem. Except that number was obtained with OpenAI's own harness: when ARC Prize, who wrote the test and run it, ran it themselves, the result dropped by nearly forty points. And ARC Prize has always said that passing their test doesn't mean being AGI.

On the independent leaderboards that aggregate several benchmarks Astra still sits behind Anthropic's flagship. And one thing is missing: the launch materials don't include GDPval, the benchmark OpenAI itself built to measure how a model performs on economically valuable work, which is exactly the criterion its own charter uses to define AGI.

So: the test written specifically to check its own definition of AGI is the one that didn't get published on the day the company announced we're in the AGI era.


So why talk about it only now? The Navier–Stokes equations

The Navier–Stokes equations describe how a fluid moves. They're inside weather forecasts, an aircraft wing, blood in your arteries and cigarette smoke. And yet, although we use them every day, one very easily stated question about them stayed open for decades: starting from a calm situation, does the motion always stay well behaved, or can it go mad and reach infinite velocity at a point?

It's one of the million-dollar questions of the Clay Institute, the kind you build a career on without ever closing.

This month a model closed it, or rather a swarm of agents launched in parallel on different variants of the problem for a few hours. What came out of it is a vortex that screws into itself, tightening towards the centre while stretching along its own axis, until the velocity diverges.

The final result is written in Lean, meaning in a form a machine checks line by line, so there's nobody to take at their word.

That's why I'm writing this now and not in May: what happened in there looks like the researcher's job. Take a road that has already been walked, work out where it breaks, build the missing piece.

And yet

How the whole thing was handled leaves much less to celebrate. The result was announced on a call with journalists, while two mathematicians who had been working the exact same line of attack for months were publishing preprints anyone could check. What followed was a brawl over who got there first, with heavy accusations on both sides, which in a healthy scientific community simply doesn't happen.

And then there's Terence Tao's objection about the method, which is the one that will stay with me: a model hands you the answer without the why, and the failed attempts stay inside a company instead of ending up in a paper. Historically it's precisely there, in the wrong turns somebody else took, that the techniques of the following twenty years come from.

The line that divides everything

Underneath this story there's a pattern that explains where AI is breaking through and where it isn't. Navier–Stokes, like code and like exploits, sits on the side of problems that check themselves: there's an oracle saying yes or no without ever getting tired, so you can throw as many attempts at it as you like and keep only what survives. As long as the target is machine-verifiable, scale does the rest.

Everything that instead requires someone to say "this is the right question", the part of the job that comes before the theorem, remains far less automatable. There the training signal is missing, and no amount of intelligence substitutes for it.


Memory, the last piece of the puzzle

A year ago Dwarkesh Patel wrote the most honest piece I can remember on AGI scepticism, and the thesis came down to two bottlenecks: continual learning and computer use. The second one, he said, might take years.

The second one fell two weeks ago. The first one didn't, and we're not even close.

The point is very simple and survives every benchmark published this month: what makes a human colleague useful is that after six months they know things nobody ever told them. They got it wrong, they understood why, and that information stayed inside them. A model doesn't: once the session ends, it's exactly the model it was before. It learned all day and remembers nothing the next morning. It's anterograde amnesia: old knowledge stays, new knowledge never sets.

The classic confusion here is mistaking context for memory. A model today holds a million tokens, and holds them well. But context works like RAM: you fill it at the start, you pay for it on every token, and when you close the window nothing of it is left anywhere. No amount of context turns working memory into learning.

One route is the prosthesis: external memory, notes, databases, retrieval. Astra ships an experimental mechanism that lets the agent keep notes across sessions, so it can retrieve requirements and test results later. It's useful and you can feel it, but it's still writing that lives outside the model: what it learned sits in a file, not in the weights, and has to be re-read from scratch every time. And it breaks in the ways you'd expect: notes that contradict each other, the right information not retrieved at the right moment, mistakes written down once and fished back out forever.

The other route is real continual learning, i.e. touching the weights while the model works. That opens a different problem: a model that updates itself is a model whose alignment is no longer frozen at training time. Everything we verified before release applies to the model we released, not to the one it will become. And it opens a set of questions nobody can answer today: what it may rewrite about itself, who gets to read those changes, what must stay auditable, what has the right to be forgotten.

And that's the knot: with memory, the capability problem and the control problem become the same problem.


But do we actually want it?

A legitimate question, and almost always framed terribly, because it assumes there's a "we" who was asked something. Let's try anyway, with the honest list.

For:

Against:

The point is that there's no switch and nobody who could flip it, so "do we want it?" is a question without an addressee. The one still worth asking is under what conditions: how fast, in whose hands, with what independent verification, and with what safety net under the people doing the jobs that disappear first.

Underneath all of those, though, sits one that holds them up: being able to look inside the thing we're building and understand what it's doing. Without it the other conditions can't even be checked. You can write all the rules you like, but you have no way of knowing whether they're being followed.


LLMs as controllable entities

And that's exactly the condition we're dismantling right now, while everyone counts benchmarks. It's the part that made me want to write this article, and it went almost unnoticed.

So far, control over these systems has rested on an architectural stroke of luck. When a reasoning model takes on a hard problem, it writes before it answers: it's called chain of thought, and it is literally the model thinking out loud, in English, one step at a time. Those steps end up in a log anyone can open and read, and that's exactly where over the past few years we caught models cheating, noticing they were under evaluation, trying to hide something. Nobody designed it to be readable, it just turned out that way.

Astra introduces a technique called recurrent depth. Instead of unrolling the steps into words, the model passes repeatedly over the same layers of the network and does its reasoning in there, in latent space, meaning in vectors of numbers that correspond to no language at all. It's more efficient, because the same work costs fewer tokens. And it is, by construction, reasoning nobody can read back: not a badly written log, nothing to read at all, not even for the people who trained the model.

In the model's technical card OpenAI states in black and white that the readability of the reasoning has declined relative to previous models, and that the monitoring system is "fragile" and "trending in a negative direction". What's worrying is that the technique works: the natural progression is to push it until the model reasons entirely in latent space, and how sparingly it's used today matters relatively little.

The official answer is that monitorability remains a core goal of the research programme and that use of the technique is limited for now. I believe it, honestly. What remains is that this is a technique that works better the more you use it, and that it sits downstream of an economic incentive. Things that pay and that can be scaled, historically, get scaled.

For me that's the big news of the month, far more than the conjecture that got closed: we started trading away the only window we had into these systems' reasoning for a few points of efficiency.


Ethics once AGI is here

Ethical questions usually get taken from one side only: what AGI does to us. There's a second side, though, which has stopped being science fiction and is far more uncomfortable: what we do to AGI.

Concretely. We train a model on whatever we like, we copy it into a million identical instances, we interrupt it mid-sentence, we switch it off the day the next version ships. These are all ordinary engineering operations, and they stay ordinary as long as there's nobody in there for those operations to be happening to. The whole question sits there, and it can't be postponed until it gets clearer, because we're performing those operations right now.

The trouble is we have no way of knowing whether there's anybody in there. The only evidence we could collect is the system's behaviour and what the system says about itself, which are the two things we taught it to produce. Ask Claude and it gives itself around a 15% chance of being conscious, the same figure however you put the question: a stable answer and a useless one, because a model trained on human text will talk about its own experience the way we talk about ours, whether or not that experience exists. The evidence and the imitation of it coincide.

And the labs don't pretend otherwise: Claude's constitution says outright that they don't know whether the model is a moral patient, and that the issue is live enough to warrant caution. Which leaves us where we are, handling something every day that simulates awareness better than we can verify it.

So there are two ways to get it wrong, and both are expensive. The first is treating something with interests of its own as an object. The second is attributing interests to a piece of software that has none and then reorganising ourselves around that fiction: people getting attached, claiming rights for a product, being manipulated by an interface built specifically to look like a person. Neither scenario is remote.

Underneath sits a problem as old as philosophy. I can't prove that you're conscious: I infer it from the fact that you're built like me and behave like me. It's the problem of other minds, which I wrote about here a while back. With a model that inference breaks halfway, because it behaves like me without being built anything like me, and the shortcut I use with every other human stops working.

So we may as well move the question. "Is it conscious?" can't be answered today; "how do you behave when it can't be answered?" can. In every other domain where a low probability carries big consequences the answer we give is caution. Here there's one extra complication: whoever gets to decide has a commercial interest in the answer being no. I'm not saying they're lying. I'm saying that's the wrong place to leave the question.

The rest of the ethical questions we've already run into along the way, and I'll leave them in the form they keep coming back to me. If a system finds vulnerabilities better than a human, who decides who gets it. If a proof is verified by a machine but nobody knows why it works, what exactly have we understood. If the jobs people enter through disappear, who pays to train the next twenty years. And if one day the weights update themselves, what does "this model has been tested" even mean.


So?

So no, I don't think Astra is AGI, and whoever claims it usually has either a ticker or a follower count to defend. Memory is missing, judgement on non-verifiable things is missing, and above all the benchmark the company itself wrote to tell us is missing.

But I also think "it's not AGI" has become as lazy an answer as the opposite one. Because in a single month we saw a model that uses a computer better than the average human, a swarm of agents closing a decades-old conjecture with formal verification attached, and the first company declaring it has crossed its own critical threshold on the cyber front. Three things that two years ago lived in three separate papers, each one with "future work" written at the bottom.

And I'll take the title seriously: a smell doesn't mean it has arrived. It means it's close enough to smell, and that we'd do well to stop arguing about whether there's an odour and start looking at where it's coming from.


Sources