Davide Piffer argues that AI’s stronger performance in mathematics may not mainly reflect deeper reasoning. His explanation is simpler: models have access to something closer to a massive notebook, one that can hold hundreds of intermediate calculations, abandoned routes, assumptions, and unresolved branches at once.
Human math performance runs into memory limits early
Piffer starts from an observation about human mathematicians: the first constraint is often not logic itself, but how many unfamiliar items a person can keep active in mind at the same time. In mental calculation, the hard part is often carrying forward prior partial results while continuing the next step.
On that view, pen and paper do not make people smarter. They offload working memory. Professional mathematicians rely on notation, scratch work, and previously proved lemmas to spread reasoning onto the page and keep moving.
Working memory and IQ are not the same variable
Piffer cites several psychology studies to support the distinction. Alloway and Passolunghi (2011) found that working memory made an independent contribution to math achievement. A six-year longitudinal study by Alloway and Alloway (2010) reported that working memory performance at age 5 predicted later academic achievement better than IQ tests. Studies by Blankenship (2015) and Friso-van den Bos (2013) reached similar conclusions.
In practical terms, if two children have similar IQ scores, the one with stronger working memory will usually do better in mathematics. Piffer says that offers a useful way to think about AI: human mathematical performance is already bottlenecked by memory, so a machine with a much larger symbolic workspace is operating under different conditions from the start.
The context window works like an external ledger
Piffer describes the context window as a large external notebook. A model can keep the problem statement, definitions, intermediate expressions, and discarded attempts visible in one place and revisit them as it continues.
He also notes that this is not identical to human working memory. Human memory updates and replaces items actively. A language model, by contrast, cannot easily overwrite what it has already written; it mostly adds more text on top. For that reason, he says the better description may be an enhanced symbolic working memory rather than working memory in the ordinary human sense.
That distinction matters in mathematics because assumptions, definitions, known conditions, proof targets, and excluded cases can usually be written in unambiguous symbolic form. Once placed on the page, those meanings do not drift quietly with mood or context.
Piffer gives a simple example. If a problem requires keeping track of the facts that n is odd, p is prime, and x is not equal to 0, a human may drop one of those conditions and accidentally divide by zero later in the derivation. A model can restate the full premise at each step. The same applies to long reasoning chains: the challenge is not just whether any single step is hard, but how many steps must be coordinated at once. In that sense, when AI improves by "thinking longer," part of the gain may come from searching more broadly in a bigger notebook rather than reaching a deeper conceptual understanding.
Math also has another advantage. Results can be checked by substitution, run through software, or tested with formal proof systems. That means external memory can be corrected quickly instead of letting mistakes accumulate unchecked.
Bigger memory does not fix ambiguous causal questions
Piffer draws a sharp line here. If the task shifts to an uncertain question such as why someone suddenly stopped replying to messages, a larger context window does not solve much. The missing facts may never have been observed in the first place: the person could be angry, busy, dealing with a broken phone, or avoiding a different issue entirely.
The same limitation shows up with words such as fairness, success, and harm, whose meanings can shift with context. A model may retain the whole conversation and still miss what the conversation is really about. Piffer says politics, history, and business judgment face a similar problem. The core difficulty is not remembering more evidence, but finding the right causal explanation when the evidence is incomplete.
Piffer limits the claim to verifiable symbolic domains
He also acknowledges that human working memory includes attention, inhibition, and continuous updating. A model’s context is more passive, and written output is not easily revised once produced. So his claim is narrower than a broad statement that AI simply has better memory and is therefore better at everything.
Instead, he places the advantage in areas like mathematics, where external symbols matter and results can be checked. He also offers a testable prediction: AI should look strongest on problems with many conditions, long calculations, and multiple case splits. On problems that depend on a single conceptual leap, the gap should narrow. By the same logic, cutting available context or blocking intermediate steps should hurt performance on long-form math problems disproportionately.
Fast like von Neumann, not yet deep like Einstein
Piffer closes by invoking a comparison from physicist Eugene Wigner. Wigner knew both John von Neumann and Albert Einstein. He described von Neumann as the fastest and sharpest mind he had seen, able to absorb large amounts of information quickly and reason across fields. Einstein, in Wigner’s telling, had deeper and more original understanding, the kind that reframed the problem itself.
Piffer says current AI looks more like a machine-amplified von Neumann: fast and able to retain a great deal, but not yet Einstein-like in the sense of overturning existing frames and redefining the question.

