AI 4 min read

What If AI Didn't Out-Think Mathematicians — It Just Out-Remembered Them?

AI systems are clearing olympiad problems and poking at open conjectures, and the obvious question follows: is AI now better at math than mathematicians? Researcher Davide Piffer wants to flip the question entirely. His claim is that AI didn’t out-reason anyone — it just showed up with vastly more working memory than a human brain has. It sounds like a cheap deflection. Follow it far enough and it gets uncomfortable.

Human working memory is embarrassingly small

Start with the number. When cognitive psychologists talk about human working memory capacity, the figure they keep landing on is 4±1. George Miller’s famous “magic number seven” got revised downward for decades, and after Nelson Cowan’s work the consensus settled near four pure slots.

Four. That’s how many chunks of information you can juggle simultaneously right now. Humans route around this with chunking, of course. A chess grandmaster recalls an entire board not by memorizing 32 pieces individually but by collapsing the position into a handful of familiar patterns. Mathematicians do the same thing. Ten lines of proof that a beginner tracks step by step is one object to an expert.

That’s where Piffer’s argument starts. What if mathematical difficulty overlaps heavily with the number of conditions you have to hold simultaneously?

Context windows aren’t in the same order of magnitude

Now the other number. Frontier models are shipping with context windows around 1 million tokens — several books’ worth. And every token inside that window attends to every other token at every step. No flipping back three pages to check what condition three said.

The obvious objection is fair: a context window isn’t working memory. Context is closer to passive storage; human working memory is an active manipulation space. The lost-in-the-middle problem is well documented, and plenty of benchmarks show performance degrading as context grows. But even after discounting for all that degradation, the effective capacity isn’t four. It isn’t in the neighborhood of four.

A soccer analogy: the other side isn’t faster than you. They brought a hundred players onto the field instead of eleven. You lose, but not because anyone out-dribbled you.

So how much of “reasoning” is real?

What makes this argument interesting is that it pokes AI skeptics and AI optimists at the same time.

To the skeptics: if AI solves problems with a large working memory, why doesn’t that count as intelligence? Human intelligence also rides on neural hardware specs. Nobody argues that synapse count and brain volume are irrelevant to cognition. Discounting AI’s results because the capacity is large is a double standard.

To the optimists: if what AI currently excels at clusters around sweeping a wide search space, then the parts of math that actually matter — taste in choosing the right problem, the leap to a genuinely new concept — remain untested. No amount of elegant brute force produces Galois theory.

My take: the framing is the problem

Honestly, “did AI beat mathematicians” is not a productive frame. When the pocket calculator arrived, nobody said it beat the mental-arithmetic prodigies. We said the tools changed.

Something similar is happening here. If the scarce resource for a human mathematician was working memory, AI is a tool aimed precisely at that deficit. Look at how working mathematicians — Terence Tao among them — actually use these systems. It’s not “think for me.” It’s “sweep the cases I missed.” Not a replacement for thought, an extension of its reach.

But there’s a sharper point buried in Piffer’s claim. A benchmark score going up and reasoning ability going up may be two different events. If scores climb because the context window got bigger, we’re buying memory, not intelligence. And memory scales. Conceptual leaps might not.

Fair warning: this isn’t a hot fight yet

To set expectations: this hasn’t blown up into a real community argument. Searching the last 30 days, there’s no Reddit thread of meaningful size on it. It sits at the intersection of a cognitive-science claim and an AI-benchmark dispute, which splits the audience into two narrow slices that don’t overlap much.

So read this less as a live controversy and more as the early shape of a question that will keep coming back. Every time AI cracks another hard problem, this same argument will resurface. Having a reflex for asking “capacity or reasoning this time?” makes you a lot harder to sway.

The takeaway

Mathematics built by creatures with four working-memory slots and mathematics built by something holding a million tokens at once will probably look different. Rather than ranking them, it’s more accurate to say that different constraints produce different kinds of beauty. Which leaves the open question: is intelligence ultimately a capacity problem, or is there something left over that capacity never explains?

AI LLM cognitive science working memory mathematics

Comments

    Loading comments...