Etched is betting that inference, not training, is where chip economics change
Specialized chips built for AI inference are starting to challenge the standing of general-purpose processors. In a report published by TechFlowPost, that shift is framed as one of the bigger changes now unfolding inside the AI hardware market. For years, AI workloads largely ran on broad-use chips. Nvidia’s GPUs could handle graphics, training, and inference, but that flexibility came with trade-offs.

Now a different approach is gaining ground. OpenAI is working on its own chips, Google already has TPU, Amazon has Trainium, and Broadcom’s work helping Google build this class of hardware has helped lift its stock this year.
At the center of the TechFlowPost report is Etched, a startup few outside the sector had heard of until recently. The company was founded by three young Harvard dropouts who are now around 25 years old. Four years after launch, Etched still has not sold a single product. Even so, it has raised about $800 million at a $5 billion valuation from backers that include Peter Thiel, Jane Street, and a venture fund linked to TSMC.
Public reports cited in the article say the company has also locked in more than $1 billion in orders for a chip that has yet to be built. The startup’s core view is simple: the biggest line item in AI is moving away from model training and toward inference, the process that happens each time systems such as ChatGPT or Claude answer a prompt.
A chip that hardwires Transformer into silicon
Etched’s design choice is unusually rigid. Rather than making a flexible chip that can be adapted to new model structures through software, the company is building a processor dedicated to inference and physically embedding the Transformer architecture used by large language models into the chip itself.
According to figures released by Etched, eight of its chips can do the same job as more than 100 Nvidia H100 GPUs. The article notes that this figure has not been independently verified. Still, it is the foundation of the company’s case that a narrowly built inference chip can outperform a general-purpose GPU by a wide margin on the task it was designed to handle.
The upside comes from giving up flexibility. Space on the chip that would otherwise support broader compatibility can instead be used for compute. That can deliver far more speed. The risk is obvious too. If leading AI systems shift away from Transformer-based designs, the chip could lose most of its value and leave little room for recovery through software updates.
Etched CEO Gavin Uberti acknowledged that point in a public interview during the company’s Series A financing in 2024. He said the company was making what he described as one of the biggest bets in AI. If the architecture disappeared, the company would fail. If it held, Etched could become the biggest company in history. At the time, Uberti was 23 and the company had $120 million on its balance sheet. Two years later, the article says, Etched has $800 million in funding and $1 billion in orders, while the chip still has not reached market.

Why inference has become the pressure point
TechFlowPost argues that Nvidia’s moat was built around training and the CUDA ecosystem, not necessarily around inference. Training happens in large bursts over months as models are built. Inference is different. It is the repeated, daily cost of serving answers to users at scale.
Multiple foreign media reports cited in the article say inference has replaced training as the largest ongoing expense for AI companies and has become a key bottleneck. The report also mentions claims that Anthropic could turn profitable this quarter on the back of inference margins.
Most inference work today still runs on Nvidia GPUs, but those processors were not built solely for that purpose. In interviews, Uberti has said GPU clusters running inference often achieve actual compute utilization of only 20% to 30%, with more than half of their capability going unused. In the article’s framing, if a buyer spends $50,000 to $150,000 on an Nvidia machine for inference, only about 30% of that capacity may be doing productive work, while the rest is consumed by heat and power draw.
That is the opening Etched and its investors are trying to exploit. Training remains where Nvidia’s position is strongest. Inference is more repetitive and more constrained, which makes it a more natural fit for a dedicated processor.
The article also names Bryan Johnson as one of the investors. Johnson has said publicly that Etched’s founders approached him years ago as dropouts and argued that faster AI chips could speed up longevity research. The faster a chip could generate tokens, the faster researchers could search for drugs and study disease.
Summer delivery will put the thesis in front of customers
The chip Etched has spent four years building is due to ship this summer. For a company that has never delivered a product, this will be the first real check on whether its hardware performs in customer data centers the way it does in presentations.
But the market it is entering has shifted while the chip was being built. The design Etched hardwired into silicon reflects its view of AI from three years ago. During that same period, some leading models changed structure. Instead of using one large model block for every query, a newer generation of open-source systems breaks the workload into many smaller parts and activates only some of them for each task. The report specifically names DeepSeek and Qwen as examples of Chinese open-source models at the front of that shift.

Software can be updated in weeks. Chips can take two years to move from blueprint to volume production. Hardwiring an architecture into silicon is, in effect, a bet that change will not come too fast.
TechFlowPost draws a comparison with the previous crypto cycle. Ethereum once had a hardware race built around specialized mining rigs that did little beyond mining Ethereum and far outperformed general-purpose graphics cards at that job. In September 2022, Ethereum changed its rules from hardware-based mining to a staking model. Hardware built for the old system, along with large numbers of GPUs, lost their purpose almost overnight, wiping out billions of dollars in value.
Etched’s founders know the parallel. Uberti’s answer, according to the article, is speed. In interviews he has repeated the phrase “too late,” arguing that Etched is at least 18 months ahead of larger rivals such as Nvidia and Google in specialized inference chips. By the time those firms respond with similar products, he believes Etched can already be on a second generation and have customers locked in.
That means the startup is making two bets at once. It is betting that Transformer remains central to mainstream AI inference. It is also betting that it can move fast enough to deliver, scale, and win customers before architecture changes outrun the chip.
A broader signal for tech investors
The article ends by arguing that Etched is not just a story about one startup. Investors in Nvidia, Broadcom, Cambricon, Cerebras, and other chip names are also making a judgment on how long a given technical path will remain useful. Etched simply pushes that logic to its most extreme form by removing nearly all room for adjustment.
If change slows, specialization looks powerful. If change speeds up, the most rigid designs face the hardest fall. This summer, Etched’s first shipments are set to offer an initial answer.


