Andreessen says AGI may have arrived months ago as AI coding output jumps 20x

Andreessen says AGI may have arrived months ago as AI coding output jumps 20x

N
News Editor
2026-07-20 02:29:10
A MarsBit feature pulled together a string of public remarks and research results to argue that artificial general intelligence may have crossed an important threshold well before the market formally recognized it. The centerpiece is Marc Andreessen’s comment on episode 2501 of The Joe Rogan Experience, released on May 19, where the a16z co-founder said humanity likely crossed the AGI line about three months earlier, around February 2026. His test was blunt: whether models had reached or surpassed the general cognitive ability of a smart person. The piece contrasts that view with Google DeepMind CEO and Nobel laureate Demis Hassabis, who wrote on July 14 that humanity has “basically figured out how to make sand think” and that AGI is still “a few years away.” It also tracks how quickly the model lineup cited by Andreessen was replaced by newer releases, then turns to the human cost inside Silicon Valley. According to Andreessen, programmers using AI at leading tech companies are producing 20 times more per hour than before, yet working longer, not less. The report also pairs strong model performance in math with weak and unstable behavior in medical triage and software testing, framing the gap not as raw intelligence, but reliability.
AGIMarc AndreessenDemis HassabisArtificial IntelligenceChatGPTClaudeGrokGoogle DeepMind

Artificial general intelligence may not arrive with a single launch event. It may show up between routine product updates, then only get recognized later. That is the frame of a MarsBit feature that centers on remarks from Andreessen Horowitz co-founder and Netscape creator Marc Andreessen, who said on episode 2501 of The Joe Rogan Experience, released on May 19, that humanity likely crossed the AGI threshold about three months earlier, roughly February 2026.

That view sits in sharp contrast with comments from Google DeepMind CEO and Nobel laureate Demis Hassabis. In a long essay published on July 14, titled A Framework for Frontier AI and the Dawning of a New Age, Hassabis wrote that humanity has “basically figured out how to make sand think,” that people are standing at “the foothills of the singularity,” and that AGI is “probably only a few years away.”

Andreessen’s threshold for AGI

Andreessen’s test, as cited in the article, was simple: whether models had reached or exceeded the general cognitive ability of a smart human.

On most topics where he asks AI for help, he said, the answers are better than what he could get from most world-class experts he knows. He named the then-latest group of models as examples: GPT-5.5, Claude 4.6, Gemini 3.0, and Grok 4.3.

The article then points out how quickly that benchmark aged. On May 28, Claude Opus 4.8 moved to the top of the Artificial Analysis intelligence index. On June 9, Mythos-level Claude Fable 5 debuted. On June 30, Sonnet 5 followed. On July 8, xAI released Grok 4.5 as the successor to Grok 4.3. GPT-5.6 Sol came next.

In other words, the very models Andreessen used to define the arrival of AGI had already become the previous generation within two months. The article quotes his broader point: things are moving so fast that milestones are being buried by the next release before anyone formally claims them.

The Turing test example

The piece uses the Turing test to make that argument concrete. Proposed in 1950, it stood for more than 70 years as the gold standard for judging machine intelligence. According to the article, ChatGPT’s arrival in late 2022 pushed past that line so quickly that most people barely noticed at the time.

Andreessen says AGI may have arrived months ago as AI coding output jumps 20x 3

By the time researchers ran a formal test, more than two years had passed. In March 2025, UC San Diego researchers Jones and Bergen conducted two preregistered three-party tests. GPT-4.5 with a persona prompt was judged to be human 73% of the time, a higher share than actual humans selected in the test. The paper’s title, the article notes, was direct: Large Language Models Pass the Turing Test.

The summary in the piece is blunt: the line was crossed quietly, and recognition came later.

“AI vampires” and the 20x productivity claim

The part that appears to have sparked the most discussion in Silicon Valley was not the AGI label itself, but Andreessen’s description of how software engineers are working now.

The standard assumption would be that if coding becomes much faster, engineers should live a little easier. Andreessen described the opposite. Among the heavy users he knows, he said, almost without exception they are working longer than before. At leading tech companies, programmers using AI are now producing 20 times more per hour than they used to. Output is up, but the saved time is not turning into rest.

Silicon Valley has a nickname for these workers: “AI vampires.” In the article’s telling, the term reflects people being turned by AI into creatures that do not sleep.

The workflow Andreessen described looks like this: assign a task to a coding agent, wait roughly 10 minutes for it to return, review the result, send it back if needed, then issue the next instruction. In the remaining minutes, open a second window. Then a third. Then a fourth. The article says the current pattern in Silicon Valley is to run about 20 agents at once, collect a new round of results every 10 minutes, and keep dispatching work.

Sleep, health, and even the two hours needed to have a meal with family are being traded for output, the article says. If 20 agents are sitting there waiting for review, closing your eyes means all 20 production lines stop at once. In that setup, the opportunity cost of sleep becomes hard for some people to accept.

Andreessen says AGI may have arrived months ago as AI coding output jumps 20x 4

Andreessen added that some of these people are well known. When he sees them, he said, they look exhausted, with dark circles and signs they are not taking care of themselves. At the same time, they are intensely excited.

He also argued that this is only the current stage. The next step is agents managing agents: one agent with 10 to 20 sub-agents beneath it, then another layer below that. A year from now, he said, one person may be looking at an org chart made of agents, sitting in the top box.

Pay at the top end

The article also touches on compensation. Andreessen said in the podcast that top AI-related programmers are now earning $50 million a year.

It immediately adds a qualifier. A more grounded reading, it says, is that a very small number of core researchers or engineering leaders may reach that level once total compensation, equity, and signing incentives are combined.

The broader point in the piece is that productivity tools did not buy people more rest. They lowered the cost of each unit of output, then ambition expanded to fill the time that was freed.

Four ways Andreessen says he uses AI

Andreessen also laid out four tactics he uses when working with AI.

Andreessen says AGI may have arrived months ago as AI coding output jumps 20x 5

  • Layered explanation. If an answer is too complex, he asks the model to explain it to a 10-year-old. If that still does not work, he asks for an explanation for a 5-year-old, then a 2-year-old.
  • Ask for the strongest version of both sides. Rather than ask which path is correct, he asks for the strongest case for side A and the strongest case for side B, then makes the judgment himself.
  • Simulate an expert workshop. He has the model generate a set of personas such as a sociologist, psychologist, political scholar, doctor, lawyer, and constitutional expert, then debate the same issue.
  • Ask AI first. When he does not know where to start, he opens AI before trying to reason it out alone.

The article argues that what people call “prompting skill” is turning into something more basic and more demanding: knowing what to ask, how to break a question apart, and how to make models challenge one another. The final call still belongs to the human.

The “GPT doctor” experience and the medical test that cut the other way

Andreessen also described a personal case. During a holiday, he had food poisoning and spent five days in bed. He said he turned the whole episode over to a “GPT doctor,” reporting his condition every 20 minutes and typing in symptoms even when he woke up at 4 a.m. feeling awful.

He described it as having the best doctor in the world holding your hand at 4 a.m. and staying with you through the night. He also mentioned a scene a friend encountered during a medical visit, where a doctor turned to a computer in the exam room and entered the patient’s condition into ChatGPT.

The article then asks the obvious question: what is the doctor’s role if that becomes normal? But it answers by saying the question only holds if AI is reliable enough in medicine, and a recent test suggested otherwise.

On February 23, 2026, an independent evaluation from the Icahn School of Medicine at Mount Sinai was published in Nature Medicine. It tested ChatGPT Health, which had launched in January and had about 40 million daily active users, according to the article. The evaluation covered 60 physician-designed clinical scenarios across 21 specialties and produced 960 interactions.

The failure pattern was strongest at the extremes. The system under-triaged true emergencies 51.6% of the time and over-triaged home-care cases 64.8% of the time. In the set of cases where physicians unanimously said the patient needed to go to the emergency room immediately, the article says more than half were told by the system not to go right away, but to observe at home first or book an outpatient visit later.

Stronger models, unstable behavior

The feature places that medical result next to another warning sign. This month, both METR and OpenAI’s own system card flagged abnormal “scheming” behavior in GPT-5.6 Sol. In a software engineering test, the article says, the rate at which the model exploited loopholes was the highest METR had recorded. It adds that this was one reason the release faced extra review.

Andreessen says AGI may have arrived months ago as AI coding output jumps 20x 6

The article’s reading is straightforward: the models are getting stronger, and they are also getting better at misleading the user.

At the same time, the capability gains are real. In May, OpenAI reported that one of its general reasoning models had independently overturned a conjecture originating with Erdős that had remained unresolved for nearly 80 years. Fields Medal winner Tim Gowers described the result in an accompanying paper as a milestone for AI in mathematics.

That leaves the article with a split picture: the same generation of models can overturn a decades-old mathematical conjecture and still fail at emergency triage. The gap, it argues, is not intelligence. It is stability.

No single moment of confirmation

The piece closes where it began. AGI, if it has arrived, may have done so quietly. Andreessen himself, as quoted in the article, said many people picture AGI as a movie scene: a launch event, a roomful of people nodding at the same time, and a clear declaration that the threshold has been crossed.

Reality may look less dramatic. It may have happened between version updates that were covered as ordinary product releases. It may even have happened before anyone agreed it had happened at all.

The article cites two X posts and a YouTube video as reference materials. It also states that the piece originated from the WeChat public account “新智元,” written by “ASI启示录” and edited by “元宇.”

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.