GPT-6 Astra revives the compute-demand trade, but the fourth-wave case is still unproven

GPT-6 Astra revives the compute-demand trade, but the fourth-wave case is still unproven

N
News Editor
2026-09-09 01:07:05
GPT-6 Astra’s Sept. 3 release quickly fed into a rebound in semiconductor and memory names, with SOXX up 3.5% on Sept. 4, Micron gaining 6.1%, SanDisk rising 11.9%, and South Korea’s KOSPI adding 4.61% on Sept. 7 as Samsung Electronics climbed 5.7% and SK hynix rose about 8%. The move came after a long stretch of volatility driven by concerns that AI capital spending had run too hot and that demand for compute might be nearing a ceiling. MarketWatch described the rally as Astra having “reignited the memory-chip trade.” A more aggressive interpretation came from Tae Kim, author of The Nvidia Way and a former Barron’s technology reporter, who argued AI may be entering a fourth exponential wave of compute demand after chatbots, reasoning models, and coding agents. In his framing, Astra’s emphasis on Computer Use could expand persistent agent workloads beyond programmers into mainstream desktop work such as Excel, Blender, CAD, Power BI, browsers, and enterprise software. Still, OpenAI has not published post-Astra Computer Use task volumes, daily token figures, GPU utilization, or inference-throughput curves, leaving the thesis unverified. The article also points to OpenAI internal data disclosed on Sept. 6 showing that by mid-August, the median researcher’s daily Coding Agent usage exceeded $600 at public API-equivalent pricing, while the 90th percentile topped $7,000. By that point, each human workday in OpenAI research corresponded to 3.1 agent workdays. Markets have started to price in a fresh AI infrastructure leg, but reliability, cost, and the long-term shift toward direct software interfaces remain the real tests for this demand story.

GPT-6 Astra’s release on Sept. 3 was followed by a sharp rebound in semiconductor and memory shares. On Sept. 4, the iShares Semiconductor ETF (SOXX) rose 3.5% in a single session, while Micron gained 6.1% and SanDisk jumped 11.9%. On Sept. 7, South Korea’s KOSPI climbed 4.61%, with Samsung Electronics up 5.7% and SK hynix up about 8%.

The move came after an extended period of volatility in AI-linked names as investors questioned whether capital spending had become overheated and whether demand for compute was approaching a peak. MarketWatch went as far as describing the rally as Astra having “reignited the memory-chip trade.”

Tae Kim, author of The Nvidia Way and a former Barron’s technology reporter, offered a more forceful interpretation: AI may be entering a fourth exponential wave of compute demand over the past four years.

A fourth wave tied to Computer Use

In Kim’s framework, the first three waves came from chatbots, reasoning, and coding agents. The fourth, he argues, could come from Computer Use, which Astra put at the center of its product demonstration.

His logic is straightforward. Chatbots brought inference into the mass market and pushed hundreds of millions of users into consuming compute. Reasoning models increased the amount of internal computation spent on each answer. Coding agents shifted the model’s role from producing one response to working continuously for tens of minutes or even hours. Computer Use now extends that persistent-agent pattern from software developers into routine desktop work across Excel, Blender, CAD, Power BI, and browser-based tasks used by ordinary knowledge workers.

Kim has tracked Nvidia and semiconductor investing for years. His newsletter, Key Context, is centered on technology-investment calls, and his recent writing has remained broadly constructive on AI infrastructure.

That matters because the “fourth exponential wave” is, at this stage, an investment thesis advanced by a bull on AI compute, not an established industry rule. The timing, though, is notable.

Before Astra’s launch, the agent model had just gone through another round of skepticism over hitting a capability wall. Short tasks kept improving, but reliability still dropped in long-horizon, multi-step, real-software environments. That fed directly into doubts about whether the compute story could continue. If agents cannot break through that ceiling, the case for sustained growth in compute demand becomes much harder to defend.

The question now is whether Astra can move that ceiling.

One human workday now maps to 3.1 agent workdays

The earlier three “exponential” waves refer to a pattern seen over the past few years: each shift in how AI was used raised the amount of compute one user could consume.

ChatGPT opened a mass market for inference. Reasoning models began spending longer on internal computation for a single prompt. Then coding agents turned a simple instruction into a repeated loop of reading code, editing, testing, checking errors, and revising again.

At the same time, agents have started breaking through the physical limit that a person can only work so many hours in a day.

On Sept. 6, OpenAI disclosed a set of internal numbers. By mid-August, researchers at the 50th percentile of usage were consuming more than $600 per day in Coding Agent inference, converted at public API pricing. At the 90th percentile, that figure exceeded $7,000 a day. Business Insider, using OpenAI data, reported that the median was only about $162 in July, implying growth of nearly 3.7 times in just over a month.

These figures convert internal usage into API-equivalent pricing, so they are best read as a measure of per-user inference intensity.

OpenAI also said that before June, the combined runtime of all agents in its research organization had not exceeded researchers’ own work time. By mid-August, each human workday already corresponded to 3.1 agent workdays. More researchers were also running multiple agents at once.

That points to a structural shift. Under the agent model, one user may now be paired with several parallel digital processes, no longer constrained by the speed of human input. A single person can keep three or four agents running continuously, and agents may in turn create sub-agents.

That may change how compute demand is measured. The key variable is no longer just the number of users, but how many agents operate behind each user and how many hours those agents run each day.

Computer Use expands the addressable workspace

Coding agents have been growing quickly, but they still face a clear boundary. Most people running multiple agent tasks today are programmers, and programmers represent only a small share of the world’s knowledge workers.

Computer Use pushes that boundary much further out. After Astra launched, Kim ran a test of his own. He said he did not know how to use Blender beforehand, so he asked Astra on a Mac to research the space shuttle, open Blender, and build a 3D model. Roughly 10 minutes later, it had produced a rotatable model.

Standard chat roughly follows an input-inference-output loop. Coding agents changed that to read code, reason, modify, test, and reason again. Computer Use adds another layer: observe the screen, interpret the interface, decide on an action, execute it, wait for the result, inspect again, verify, and correct errors.

A task that takes a human 10 minutes may require dozens of rounds of visual understanding, reasoning, and tool calls on the model side.

Once that capability moves from browsers and code editors into Excel, Salesforce, SAP, Power BI, Photoshop, CAD, and internal enterprise systems, the potential user base broadens from programmers to nearly every knowledge worker sitting in front of a computer.

That is the core of Kim’s “fourth wave” argument: coding agents aim to capture programmers’ computer time; Computer Use aims to capture all white-collar computer time.

For now, though, it remains a possibility rather than a confirmed trend.

OpenAI has not published post-Astra Computer Use task volumes. It also has not shared daily curves for per-user tokens, GPU utilization, or inference throughput. Without that data, there is no firm basis to say Astra has already produced a fourth wave of compute growth.

Are there early signs of compute strain?

After Astra’s release, Tencent Technology noted a type of feedback appearing in some developer communities that is difficult to quantify but still hard to dismiss: users in different regions, and on different accounts, seemed to be having different experiences.

Some developers in Asia-Pacific said ChatGPT and Codex had recently become slower to respond, and that complex tasks felt less stable than on U.S.-based accounts. In community shorthand, that experience is often described as a model getting “dumber.”

At least part of that may not be purely subjective.

On Sept. 4, OpenAI’s official status page separately reported degraded performance in Asia-Pacific, affecting products including ChatGPT, Work, and Codex Cloud. Four days later, OpenAI announced a multi-year agreement with Firmus, a data-center operator backed by Nvidia support, to secure dedicated compute capacity from two data centers in Malaysia.

Even so, there is no evidence at this point that OpenAI responded to higher Astra-related loads by lowering inference budgets for Asia-Pacific users or by switching them to weaker models.

The perception of a weaker experience in Asia-Pacific may reflect at least three different factors mixed together: regional infrastructure and routing issues, account-level risk controls and rate limits, and dynamic resource allocation during peak periods. On the user side, all of them can show up as slower response times, more failures, interrupted tool calls, or worse final answers on complex tasks.

Still, this feedback offers another way to watch for stress in AI compute supply. As models become more capable of consuming compute continuously, tight supply may first appear as fluctuations in latency, failure rates, and service quality across regions and account types.

The problem is that one critical dataset is still missing: whether the same model, under the same plan, on the same task shows stable differences across regions in time to first token, total inference time, tool success rate, and final task-completion rate.

Until that data is available, claims that the model has become “dumber” can only be treated as clues, not proof of compute scarcity.

On Sept. 8, a user posted a screenshot showing GPT-6 Astra at “capacity full” and said none of the user’s four accounts could operate normally, with the system suggesting a switch to another model. Tibo reposted it and joked, “we are so back,” reading the full-capacity message as a sign that AI demand was heating up again.

Markets are already trading the no-peak-demand narrative

Even without public data on Astra’s actual post-launch compute curve, capital markets have started to trade the idea that AI demand has not topped out.

MarketWatch said the iShares Semiconductor ETF rose about 3.5% after Astra’s launch, while memory-linked companies such as Samsung, SK hynix, and Kioxia outperformed. Some foreign media reports described the move directly as Astra having “reignited the memory-chip trade.”

The standout in this rally was not Nvidia. It was memory again.

If Computer Use does become a persistent agent workload, the added demand would not be limited to GPU cycles. Longer contexts, KV cache, concurrent agents, virtual machines, browsers, and software environments could all push demand higher for HBM, DRAM, CPUs, networking, and storage.

That means the market is not simply trading on the idea that GPT-6 is stronger. The central question is whether a smarter model will persuade users to buy more compute.

This is also one of the main differences between the current AI infrastructure story and traditional software.

In conventional software, better optimization can reduce the server resources needed per user. In generative AI, the opposite may happen through a version of the Jevons paradox: the more efficient the model becomes, and the higher the task success rate gets, the more willing people are to hand over longer and more complex work, causing total compute consumption to keep rising.

That said, markets have their own incentives.

AI infrastructure had already been through a large run-up and a correction before this. Semiconductor stocks also saw notable pullbacks this year on concerns over overheated capital spending and weaker free cash flow at cloud providers. A few days of rebound after Astra only show that investors have started placing the bet again. They do not prove that a fourth wave of compute demand has already arrived.

Three practical questions behind the thesis

Whether Kim’s argument ultimately holds depends on three more practical questions.

The first is reliability.

Real enterprise work is not a demo. It is one thing to show an agent completing a task once. It is another for that agent to operate ERP systems, financial models, or engineering software for hours while staying correct through pop-up windows, permission changes, and data updates.

The second is cost.

How much will it cost for Computer Use to complete one hour of work equivalent to a human in front of a computer?

If AI can do $50 worth of employee work for $10, demand could expand quickly. If it costs $100, the addressable market looks very different. The internal OpenAI figure of $600 per day shows that agents can generate heavy demand, but it also suggests this mode of work may still be expensive today.

The third question is whether Computer Use could end up erasing part of its own demand.

Models need to look at screens and search for buttons today because most software was designed for humans. If Salesforce, Microsoft, Adobe, or internal enterprise systems begin exposing direct APIs, MCP, and agent interfaces, many tasks would no longer need to simulate a mouse and keyboard.

A more likely end state is a merger of Computer Use and Tool Use. Where structured interfaces exist, agents will call them directly. In legacy systems without interfaces, in hard-to-refactor software, and on open web pages, visual operation will remain useful.

That suggests the market opened by Computer Use may not be one where AI clicks through software forever like a human. Instead, it may matter because it gives agents a forceful way to cross into software environments that automation could not previously reach.

That, in turn, is the part of Kim’s “fourth exponential wave” that still needs to be tested: whether Computer Use is merely a wedge, or a genuine new shift in the way compute demand scales.

After GPT-6, the compute story that had become harder to tell because of doubts around agent limits is being told again.

Capital markets are already placing the bet. Whether this time is actually different remains unanswered.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
800

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.